Ke Liu 0004

dblp:32/2948-4 · DBLP profile ↗
← Back
31ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0003-4151-1416ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 14 · 6 first-author · 4 since 2021Systems, architecture and hardware · 11 · 7 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 RaidenSwap: A Multi-Swap Remote System for Multi-core Applications
abstract
Kernel-based remote memory systems are gaining traction in datacenters due to their significant improvement in memory utilization and their ability to transparently provide applications with unlimited memory capacity. However, the high degree of parallelism in contemporary applications leads to a significant demand for remote access throughput, which mismatches with the state-of-the-art kernel swap path due to its inherently limited parallelism. As a result, it cannot scale up the multi-core applications. We dive into the implementation of the swap path and identify the root cause behind it - significant lock contentions and inefficient swap tasks offloading.
Kefan Liu, Ke Liu 0004, Xu Zhang 0033, Ning Liu 0031, Sa Wang, Yungang Bao, Mingyu Chen 0001, Chenxi Wang 0005
EuroSys2
2025 A Unified Framework for DRL-Based Congestion Control to Optimize QoS Over Mobile Networks
abstract
Deep Reinforcement Learning-based Congestion Control Algorithms (DRL-based CCA) have shown their great potential to adapt to various environments automatically (e.g., Orca). However, it is difficult for existing DRL-based CCAs to achieve superior QoS consistently, particularly over mobile networks with rapid network fluctuations. The fundamental problem stems from the training of a single model that encompasses a wide range of network conditions. The resulting model can be overly generalized, rendering it less accurately or optimally tailored for a specific network condition. To tackle this challenge, we develop Network-segmented Model Specialization (NMS), a framework that automatically maximizes QoS for any DRLbased CCA under different network conditions. Specifically, NMS generates a set of models offline, each trained for a specific network segment. Online, it selects the model based on the current segment. We showed that NMS not only improves the QoSs of existing DRL-based CCAs consistently, but also opens a new way for the exploration of a clean-slate approach, as opposed to following the hybrid TCP/DRL approach. Hence, we designed Galaxy, a novel clean-slate DRL-based CCA that addresses the inherent limitations in clean-slate approaches by incorporating network segments' knowledge. Extensive evaluations show NMSoptimized Galaxy further explores NMS to achieve superior QoS.
Ke Liu 0004, Jack Y. B. Lee, Theophilus Benson, Yungang Bao, Mingyu Chen 0001
IWQoS2
2025 DRack: A CXL-Disaggregated Rack Architecture to Boost Inter-Rack Communication
Xu Zhang 0033, Ke Liu 0004, Yuan Hui 0001, Yisong Chang, Yizhou Shan, Ke Zhang 0017, Yungang Bao, Mingyu Chen 0001, Chenxi Wang 0005
USENIX ATC2
2025 SLVS: A Self-Learning Approach to Achieve Near-Second Low-Latency Video Streaming Under Highly Variable Networks
abstract
Fueled by the rapid advances in high-speed mobile networks, live video streaming has seen explosive growth in recent years and some DASH-based algorithms were specifically proposed for low-latency video delivery. We conducted a measurement study for the state-of-the-art algorithms with large-scale network traces. It reveals that these algorithms are susceptible to network condition changes due to the use of solo universal adaptation logics, resulting in the playback latency that has substantial variations across highly fluctuating networks. To tackle this challenge, this paper proposes Stateful Live Video Streaming (SLVS), which is a novel self-learning approach that learns the various network features and optimizes the adaptation logic separately for different network conditions, then dynamically tunes the logic at runtime, so that bitrate decision can better match the changing networks. Moreover, we further generalize SLVS to complement the streaming platform already in service to make it compatible with any live streaming services. Extensive evaluations based on real system prototypes show that SLVS can control playback latency down to 1 s while improving Quality-of-Experience (QoE) by 17.7% to 31.8%. Moreover, it has strong robustness to maintain near-second latency over highly fluctuating networks as well as long periods of video viewing.
Ke Liu 0004, Mengbai Xiao, Bingshu Wang, Dongxiao Yu, Xiuzhen Cheng
IEEE Trans. Mob. Comput.2
2023 MARB: Bridge the Semantic Gap between Operating System and Application Memory Access Behavior
abstract
The virtual memory subsystem (VMS) is a long-standing and integral part of an operating system (OS). It plays a vital role in enabling remote memory systems over fast data center networks and is promising in terms of transparency and generality. Specifically, these systems use three VMS mechanisms: demand paging, page swapping, and page prefetching. However, the VMS inherent data path is costly, which takes a huge toll on performance. Despite prior efforts to propose page swapping and prefetching algorithms to minimize the occurrences of the data path, they still fall short due to the semantic gap between the OS and applications - the VMS has limited knowledge of its running applications' memory access behaviors. In this paper, orthogonal to prior efforts, we take a fundamen-tally different approach by building an efficient framework to collect full memory access traces at the local bus, and make them available to the OS through CPU cache. Consequently, the page swapping and page prefetching can use this trace to make better decisions, thereby improving the overall performance of systems. We implement a proof-of-concept prototype on commodity x86 servers using a hardware-based memory tracking tool. To show-case our framework's benefits, we integrate it with a state-of-the-art remote memory system and the default kernel page eviction subsystem. Our evaluation shows promising improvements.
Ke Liu 0004, Ting Liang, Zuojun Li, Tianyue Lu, Yisong Chang, Yinben Xia, Yungang Bao, Mingyu Chen 0001, Yizhou Shan
DATE2
2023 HoPP: Hardware-Software Co-Designed Page Prefetching for Disaggregated Memory
abstract
Memory disaggregation is a promising direction to mitigate memory contention in datacenters. To make memory disaggregation practical, prior efforts expose remote memory to applications transparently via virtual memory subsystem’s swapping interface. However, due to the semantic gap between OS and applications – OS cannot know the memory accessing sequences of an application but via page faults. This approach has two limitations. First, it learns little from page faults’ access history, which leads to sub-optimal prefetching predictions. Second, a page fault can still occur even if there is a prefetch-hit which leads to a large kernel overhead.To address such limitations, our key insight is to decouple the address capturing from page faults by collecting full memory access traces in the memory controller. Using this idea, we buildHoPP– a hardware-software co-designed prefetching framework.HoPPadds hardware modules to the memory controller to feed sufficient hot pages to OS in real-time, which has three benefits inHoPP’s software design: 1) it improves existing prefetching algorithms with simple revamps, also offers more insights to build better policies; 2) the prefetch algorithm can run as a separate data path alongside the normal remote data path via page faults, potentially hiding the swap latency from applications, and enabling fine-grained control over prefetching behaviors; 3) the prefetch-hit overhead can be eliminated by early page table entry (PTE) injection, i.e., inject PTE for the prefetched page as soon as it returns. We implemented a proof-of-concept prototype using commodity servers along with a hardware-based memory tracking tool calledHMTTto emulate a modified memory controller. Results show that compared to Fastswap and Leap,HoPP-optimized prefetching algorithm achieves over 90% accuracy and coverage, which leads to up to 59% completion time improvement for various datacenter applications.
Ke Liu 0004, Ting Liang, Zuojun Li, Tianyue Lu, Yinben Xia, Yungang Bao, Mingyu Chen 0001, Yizhou Shan
HPCA2
2023 A Data-Driven Framework for TCP to Achieve Flexible QoS Control in Mobile Data Networks
abstract
Learning-based approaches have shown their great potential to adapt themselves to various environments (e.g., PCC and Sprout). Unfortunately, they do not consistently achieve superior QoS across different network conditions and configurations in mobile networks. Furthermore, although they can offer multiple application objectives by adjusting a preference weight vector, it is challenging for users to accurately express an application objective with a weight vector. In this work, we argue that, if configured correctly, the delay-based TCP scheme can outperform learned ones, and allow users to directly specify their objectives. To this end, we propose Post-QoS Analysis (PQSA), a data-driven framework that trains the key QoS-impacting parameters of the scheme to capture the statistical correlations between QoS objectives, network conditions, and configurations, thereby determining the optimal parameter-set that meets the user-defined QoS objective under different network conditions and configurations. To support this, we enhance conventional delay-based TCP design to develop a Generalized TCP-like Rate controller (GR) by exporting three key parameters. Extensive evaluations show that PQSA-optimized GR outperforms existing schemes in different scenarios consistently, and enables service providers to control the QoS flexibly.
Ke Liu 0004, Ting Liang, Theophilus Benson, Jack Y. B. Lee, Vaneet Aggarwal, Yungang Bao, Mingyu Chen 0001
IWQoS2
2023 An Intelligent Learning Approach to Achieve Near-Second Low-Latency Live Video Streaming under Highly Fluctuating Networks
abstract
Fueled by the rapid advances in high-speed mobile networks, live video streaming has seen explosive growth in recent years and many DASH-based bitrate adaptive streaming algorithms were specifically proposed for low-latency video delivery. However, our investigations revealed that these algorithms are susceptible to network condition changes due to the use of solo universal adaptation logics, resulting the playback latency that has substantial variations across highly-fluctuating network environments and fails to meet the service quality requirement all the time. To tackle this challenge, this paper proposes Stateful Live Video Streaming (SLVS), which is a novel learning approach that learns the various network features and optimizes the adaptation logic separately for different network conditions, then dynamically tunes the logic at runtime, so that bitrate decision can better match the changing networks. Extensive evaluations show that SLVS can control playback latency down to 1s while improving Quality-of-Experience (QoE) by 17.7% to 31.8%. Moreover, it has strong robustness to maintain near-second latency over highly-fluctuating networks as well as long-period of video viewing.
Ke Liu 0004, Mengbai Xiao, Bingshu Wang, Vaneet Aggarwal
ACM Multimedia2
2023 Post-Streaming Wastage Analysis - A Data Wastage Aware Framework in Mobile Video Streaming
abstract
Mobile video streaming is now ubiquitous among mobile users. This work investigates a less studied and yet significant problem in mobile video streaming – data wastage, i.e., some downloaded video data may not be played back but discarded by video players due to early departure or video skip, thus the bandwidth consumed in transferring them is wasted. Our measurements show that data wastage is significant in practice, e.g., 25.2 percent∼51.7 percent of video data downloaded are in fact wasted. Moreover, substantial data wastage exists not only in current commercial streaming platforms, but also in state-of-the-art adaptive streaming systems proposed in the literature. This work develops a new post-streaming wastage analysis (PSWA) framework to tackle this problem by converting existing adaptive streaming algorithms into data wastage aware versions. PSWA enables the streaming vendors to explicitly control the tradeoff between data wastage and quality-of-experience (QoE). Extensive evaluations show that PSWA can reduce data wastage significantly (e.g., 80 percent) without any adverse impact on QoE. Moreover, it has strong robustness to perform consistently across a wide range of networks. PSWA can be readily implemented into current streaming platforms, and thus offers a practical solution to data wastage for mobile streaming services.
Ke Liu 0004, Haibo Hu 0001, Vaneet Aggarwal, Jack Y. B. Lee
IEEE Trans. Mob. Comput.2
2023 DUASVS: A Mobile Data Saving Strategy in Short-Form Video Streaming
abstract
Fueled by the emerging short video applications (e.g., TikTok), streaming short-form videos nowadays is ubiquitous among mobile users. During the viewing, one common action is to scroll the screen to switch videos, which is a handy operation for the viewers to quickly search for content of interest. However, our empirical measurements reveal that frequent video switching can result in nearly half of the mobile data quota being used for transferring the video data that is never watched. This problem is called data loss in this work. Given the immense cost of the network infrastructure, such a high proportion of data loss is financially tremendous to both mobile users and streaming vendors. To tackle the problem, this study proposes a novel system called Data Usage Aware Short Video Streaming (DUASVS), where a new Integrated Learning is used to capture the characters of past network conditions and then trains intelligent adaptation models to reduce data loss and save data usage. Extensive evaluations show that DUASVS is able to save 70.7%∼83.2% of mobile data usage without incurring any QoE degradation. Moreover, the system exhibits strong robustness, performing consistently over a wide range of network environments as well as video streaming sessions.
Jie Zhang 0042, Ke Liu 0004, Jack Y. B. Lee, Haibo Hu 0001, Vaneet Aggarwal
IEEE Trans. Serv. Comput.3
2022 GraFF: A Multi-FPGA System with Memory Semantic Fabric for Scalable Graph Processing
abstract
FPGA has been a promising solution for graph processing in many scenarios. With a rapid growth in graph size, the on/off-chip memory capacity of a single FPGA is insufficient to hold large-scale graphs. To tackle such problem, in this position paper, we introduce GraFF, a Graph processing system with multiple FPGAs interconnected via a custom memory semantic Fabric. In order to efficiently exploit system parallelism, we first split the traversal of graph data into a series of independent fine-grained flits that are concurrently delivered among FPGAs as sheer memory semantic transactions. Then we relax FPGAs' synchronization from strict barrier boundaries between adjacent supersteps to fully parallelize graph traversing and computing. We build a prototype of GraFF with four custom FPGA nodes. Preliminary evaluation result based on the Breadth First Search (BFS) algorithm shows that the peak performance of GraFF reaches up to 6.23 GTEPS. Moreover, GraFF exhibits linear scalability when the number of FPGAs rises from one to four.
Xu Zhang 0033, Yisong Chang, Tianyue Lu, Ke Liu 0004, Ke Zhang 0017, Mingyu Chen 0001
FPT4
2022 HCMonitor: An accurate measurement system for high concurrent network services
abstract
Abstract This article aims to enhance the monitoring accuracy of high concurrent network services. As modern network services grow rapidly in data centers, tail latency has become one of the most crucial deciding factors on user experience. Latency measurement and anomaly detection are essential in evaluating service performance. Existing monitoring tools can be divided into two categories according to estimation methods. First, approaches based on sample traffic sample network packets to unburden the measurement. Second, approaches based on full traffic like wrk, analyze all of the packets from the kernel network stack and load the client‐side overhead into response delay. Therefore, we propose a high‐performance monitor system named HCMonitor, which computes the server‐side response latency and the round‐trip time of per‐request. It can afford full traffic monitoring on the basis of userspace, “zero copy” and pipeline. By switch mirroring, the measured latency eliminates the kernel network stack overhead and the queuing delay of the client‐side. Such measurement results in improved accuracy, online analysis, anomaly detection, real‐time display and transparent to network services. Our evaluations show HCMonitor obtains a higher throughput compared with tcpdump by over 200 times. Compared with wrk, the tail latency accuracy shows an increase by up to 72%–76% in high concurrent networks.
Ke Liu 0004, Yifan Shen 0002, Mingyu Chen 0001
Concurr. Comput. Pract. Exp.3
2021 Short Video Streaming With Data Wastage Awareness
abstract
Fueled by emerging short video applications (e.g., TikTok), streaming short videos is now ubiquitous among mobile users. A major problem in short video streaming is data wastage, i.e., the downloaded video data is not watched but discarded due to video switching, so the bandwidth consumed in video transferring is wasted. Our measurements show that 45.2% of video data is wasted in practice, which is a significant proportion that can lead to tremendous financial loss. To tackle the problem, this work develops a novel algorithm called Wastage Aware Streaming (WAS) which learns viewing behaviors and network conditions to reduce data wastage while keeping Quality-of-Experience (QoE) intact. Extensive evaluations show that WAS can substantially reduce data wastage, e.g., 70%, without any adverse impact on QoE, thus it offers a practical solution for data wastage in short video services.
Ke Liu 0004, Haibo Hu 0001
ICME2
2021 A Unified Framework for Flexible Playback Latency Control in Live Video Streaming
abstract
Live video streaming has seen tremendous growth in the past decade. An important fact in live streaming is that the demand for low playback-latency inherently conflicts with the desire for high QoE. This requires different types of live services to seek different latency-QoE tradeoffs according to their service-requirements. However, our investigations revealed that it is fundamentally difficult for existing streaming algorithms to keep consistent latency in changing network conditions, let alone achieve the service-desired latency-QoE tradeoff. To tackle the challenge, this article develops a novel framework called Flexible Latency Aware Streaming (FLAS) that not only can achieve consistent low latency, but also control the latency-QoE tradeoff flexibly. Specifically, FLAS generates a set of adaptation logics offline, each optimized for a candidate tradeoff point, then selects the most appropriate one to run online. We first show how FLAS can be applied to optimizing the existing algorithms, then developed a novel Genetic Programming approach to fully exploit FLAS's potential. Extensive evaluations show that FLAS can precisely control latency all the way down to 1s and achieve substantially higher QoE than state-of-the-arts. FLAS can be readily implemented into real streaming platforms, offering a practical and reliable solution for live-streaming services.
Jack Y. B. Lee, Ke Liu 0004, Haibo Hu 0001, Vaneet Aggarwal
IEEE Trans. Parallel Distributed Syst.3
2020 Freeway: an order-less user-space framework for non-real-time applications
abstract
The demand for high network capacity has been rapidly increasing, such as inter-DC WANs. But the transport over the network with high bandwidth and delay cannot fully utilize the bandwidth because of the inevitable packet loss and the flow control bottleneck caused by the out-of-order data blocked in the receive buffer. We further found that a lot of applications over inter-DC WANs are non-real-time, which are insensitive to data arriving sequence. Thus, we design and implement Freeway, a user-space bulk-data network transfer framework, to improve the bandwidth of these non-real-time applications over inter-DC WANs. Experimental results show that Freeway achieves 100% more bandwidth utilization than Linux TCP stack, and significantly reduces memory cost.
Yifan Shen 0002, Ke Liu 0004, Ziting Guo, Vaneet Aggarwal, Mingyu Chen 0001
CF2
2020 Labeled Network Stack: A High-Concurrency and Low-Tail Latency Cloud Server Framework for Massive IoT Devices
Ke Liu 0004, Yifan Shen 0002, Yazhu Lan, Mingyu Chen 0001, Yuan-Fei Chen
J. Comput. Sci. Technol.2
2020 Optimized Preference-Aware Multi-Path Video Streaming with Scalable Video Coding
abstract
Most client hosts are equipped with multiple network interfaces (e.g., WiFi and cellular networks). Simultaneous access of multiple interfaces can significantly improve the users' quality of experience (QoE) in video streaming. An intuitive approach to achieve it is to use Multi-path TCP (MPTCP). However, the deployment of MPTCP, especially with link preference, requires OS kernel update at both the client and server side, and a vast amount of commercial content providers do not support MPTCP. Thus, in this paper, we realize a multi-path video streaming algorithm in the application layer instead, by considering Scalable Video Coding (SVC), where each layer of every chunk can be fetched from only one of the orthogonal paths. We formulate the quality decisions of video chunks subject to the available bandwidth of the different paths, chunk deadlines, and link preferences as an optimization problem. The objective is to to optimize a QoE metric that maintains a tradeoff between maximizing the playback rate of every chunk and ensuring fairness among chunks. The proposed metric prefers to use bandwidth of the links to optimize a concave utility function of the chunk quality. Even though the formulation is a non-convex discrete optimization, we provide a quadratic complexity algorithm which is shown to be optimal in some special cases. We further propose an online algorithm where several challenges including bandwidth prediction errors, are addressed. Extensive emulated experiments in a real testbed with real traces of public dataset reveal the robustness of our scheme and demonstrate its significant performance improvement compared to other multi-path algorithms.
Anis Elgabli, Ke Liu 0004, Vaneet Aggarwal
IEEE Trans. Mob. Comput.2
2020 Optimizing TCP Loss Recovery Performance Over Mobile Data Networks
abstract
Recent advances in high-speed mobile networks have revealed new bottlenecks in ubiquitous TCP protocol deployed in the Internet. In addition to differentiating non-congestive loss from congestive loss, our experiments revealed two significant performance bottlenecks during the loss recovery phase: flow control bottleneck and application stall, resulting in degradation in QoS performance. To tackle these two problems, we first develop a novel opportunistic retransmission algorithm to eliminate the flow control bottleneck, which enables TCP sender to transmit new packets even if receiver's receiving window is exhausted. Second, application stall can be significantly alleviated by carefully monitoring and tuning the TCP sending buffer growth mechanism. We implemented and modularized the proposed algorithms in the Linux kernel thus they can plug-and-play with the existing TCP loss recovery algorithms easily. We evaluated our proposed algorithms over emulated and real experiments and showed that, compared to existing TCP loss recovery algorithms, the proposed optimization algorithms improve the bandwidth efficiency by up to 133 percent and completely mitigate RTT spikes, i.e., over 50 percent RTT reduction, over the loss recovery phase.
Ke Liu 0004, Zhongbin Zha, Wenkai Wan, Vaneet Aggarwal, Binzhang Fu, Mingyu Chen 0001
IEEE Trans. Mob. Comput.1
2019 HCMonitor: An Accurate Measurement System for High Concurrent Network Services
abstract
As user-interactive services grow explosively in datacenters, latency has become one of the most deciding factors on user experience. Therefore, estimating the latency and detecting anomalies from the expected latency is essential to evaluate services' performance. Although many existing tools have been used widely, their estimation methods can be divided into two categories. First, the traffic-sample-based approaches sample the network traffic for accelerating the estimation rather than measure every response time. Second, the full-traffic-based approaches, such as tcpdump and wrk, analyze data from kernel and leave the latency computation to the client-side. In this paper, we attempt to compute the applications' server-side latency for every request in real-time, and eliminate kernel processing delay. We propose a system named HCMonitor. It monitors all the traffic by switch mirroring, which results in high throughput and more accuracy in server-side latency estimation. The latency measurement is transparent to network services and can be displayed in real time. Our evaluations show HCMonitor obtains higher throughput than tcpdump by over 1000 times. Compared to wrk, the tail latency accuracy estimated by HCMonitor shows a promotion by up to 72%~76% in high concurrent network, by eliminating delay produced by packet transfer, kernel network stack and packets queuing on client side.
Ke Liu 0004, Yifan Shen 0002, Mingyu Chen 0001
NAS3
2018 Labeled Network Stack: A Co-designed Stack for Low Tail-Latency and High Concurrency in Datacenter Services
Ke Liu 0004, Lan Yu, Mingyu Chen 0001
NPC2
2017 Joint Upload-Download TCP Acceleration over Mobile Data Networks
abstract
Upload and download traffic often coexist in mobile networks. However, TCP download throughput could be substantially degraded by upload traffic even if the downlink is not the bottleneck. Previous works such as RSFC and TCP-RRE can substantially improve TCP download throughput in the presence of concurrent TCP upload flows, albeit at the expense of significantly degraded upload throughput performance. This work addresses this limitation by developing a novel Aggregate Transmission Rate Controller with Upload and Download flows aggregations (ATRC-UD) to jointly accelerate concurrent TCP upload and download flows from/to the same mobile device. The insight is that existing TCP as well as other flow-based approaches all suffer from ACK packets delayed by data packets from TCP flows in the opposite direction, resulting in significant errors in bandwidth estimation. By contrast, ATRC-UD exploits data packets of the opposite direction to enable continuously estimation of the downlink bandwidth and queueing delay even when ACK packets are significantly delayed. This allows ATRC-UD to track the bandwidth and delay variations more closely to maintain a shorter queue length at the downlink, thus jointly improve the download-upload throughput. Extensive emulated and real-world experiments showed that ATRC-UD enables TCP to achieve 96% downlink bandwidth utilization while improving uplink bandwidth utilization by over 115% compared to existing approaches, such as TCP-RRE and RSFC.
Ke Liu 0004, Vaneet Aggarwal, Ziyu Shao, Mingyu Chen 0001
SECON1
2016 Intra-host Rate Control with Centralized Approach
abstract
Today's datacenter is shared among various applications with different QoS requirements, which poses a great challenge to deliver low delay transport with high throughput. Most of works address this challenge by reducing the in-network delay, but assumes a negligible local delay. However, we show that this assumption does not hold for a multi-tenant datacenter that a physical machine is shared by multiple tenants with virtual machines running different applications. As measured, we found that VMs in a PM competing for bandwidth resources introduce delays as high as 13 ms, resulted from the packet queueing at QDisc layer of that PM, because current VMs' rate control still operates in a distributed manner without exploiting knowledge of the QoS requirements of applications running in VMs. This work addresses this problem by proposing a centralized rate adaptation (CERA) that operates in the host PM, dynamically schedules the flows from all VMs in a centralized manner. We implemented a CERA prototype and evaluated CERA through testbed experiments. Our results show that CERA reduces the local delay significantly thus reduces the average request latency of delay sensitive applications, e.g., memcached, by a factor of 6.3, without sacrificing the throughput performance of throughput intensive applications, e.g., iperf.
Ke Liu 0004, Yifan Shen 0002, Jack Y. B. Lee, Mingyu Chen 0001, Lixin Zhang 0002
CLUSTER2
2016 Isolating bandwidth guarantees from work conservation in the cloud
abstract
To predict lower bounds on the performance of applications, the cloud should provide guarantees on bandwidth that each virtual machine can obtain. By competing for spare network bandwidth, current solutions that provide bandwidth guarantees can achieve work conservation as well. However, they usually fail to provide accurate bandwidth guarantees, for the interference between traffic for achieving the two objectives respectively. In order to eliminate the interference, they reserve sufficient bandwidth headroom for every link, which cannot be allocated to tenants as guarantees, incurring a decrease in the total of guarantees that each link can offer and thus a decline in the revenue of the cloud provider. To address the problem, this paper proposes a new mechanism, namely DFlow, which achieves bandwidth guarantees and work conservation simultaneously. Specifically, DFlow isolates its solutions for achieving bandwidth guarantees and work conservation from each other by splitting every flow into two subflows with distinct priorities. They are then used to achieve the two objectives respectively. Our evaluations show that DFlow can provide accurate bandwidth guarantees without reserving any bandwidth headroom while achieving work conservation to effectively utilize spare network bandwidth.
Ke Liu 0004, Binzhang Fu, Mingyu Chen 0001, Lixin Zhang 0002
ISCC2
2016 Adaptive rate control over mobile data networks with heuristic rate compensations
abstract
Mobile data networks exhibit highly variable data rates and stochastic non-congestion-related packet loss. These challenges result in key performance bottlenecks in current Transmission Control Protocol (TCP) implementations: bandwidth inefficiency and large end-to-end delay. This work addresses these challenges by first developing a Sliding Interval based Rate Adaptation (SIRA) that tracks bandwidths with a fixed time interval and applies them to its transmission rate periodically. Extensive experiments confirmed that SIRA achieves 96.3% bandwidth utilization and reduces the average queueing delay by a factor of 1.37, compared to TCP CUBIC, the preferred variant for Internet servers. However, the resultant end-to-end delay is still much larger for interactive applications, thus we complement SIRA with two heuristic rate compensation algorithms (SIRA-H) given that the bandwidth does not vary significantly in long time scales. Specifically, SIRA-H first reduces the transmission rate of SIRA if the estimated RTT is above a prefigured threshold. Meanwhile, it computes the amount of unsent data that would be transmitted if SIRA were used, and compensates the rate reduction with those unsent data as if their ACKs were received, when the queue is detected to be empty. We evaluated SIRA-H through a combination of trace-driven emulations and real-world experiments, and showed that it reduces the 95thpercentile queueing delay by a factor of over 3.9, while maintains a similar throughput compared to the original SIRA. In comparison to state of the art protocols such as Sprout and Verus, SIRA-H also reduces the 95thpercentile queueing delay by a factor of over 0.8.
Ke Liu 0004, Jack Y. B. Lee, Mingyu Chen 0001, Lixin Zhang 0002
IWQoS1
2016 On Improving TCP Performance over Mobile Data Networks
abstract
Mobile data networks such as 3G and LTE exhibit properties that are fundamentally different from those of fixed networks. These differences result in severe performance bottlenecks in current Transmission Control Protocol (TCP) implementations, which are the foundation for most of today's mobile applications. This work addresses this challenge by developing a transparent protocol optimization device to perform on-the-fly protocol optimization to improve TCP's throughput performance while maintaining full compatibility with current end-host TCP implementations. The proposed protocol optimizations can achieve near optimal bandwidth utilization. This was verified and confirmed in production 3G and LTE networks using a prototype implementation. Compared to current TCP implementations, the proposed protocol optimization device can raise throughput by 48 to 163 percent. In contrast to inventing a new transport protocol or modifying an existing TCP implementation, the proposed approach does not require any modification to the existing TCP implementation at the client/server hosts, does not require any reconfiguration of the server or client, and hence can be readily deployed in today's 3G and 4G mobile networks, raising the throughput performance of all existing network applications running atop TCP.
Ke Liu 0004, Jack Y. B. Lee
IEEE Trans. Mob. Comput.1
2015 Optimizing TCP loss recovery performance over mobile data networks
abstract
Recent advances in high-speed mobile networks have revealed new bottlenecks in ubiquitous TCP protocol deployed in the Internet. In addition to differentiating non-congestive loss from congestive loss, our experiments revealed two significant performance bottlenecks during loss recovery phase: flow control bottleneck and application stall, resulting in degradation in QoS performance. To tackle these two problems we firstly develop a novel opportunistic retransmission algorithm to eliminate the flow control bottleneck, which enables TCP sender to transmit new packets even if receiver's receiving window is exhausted. Secondly, the application stall can be significantly alleviated by carefully monitoring and tuning the TCP send buffer growth mechanism. We implemented and modularized the proposed algorithms in the Linux kernel thus they can plug-and-play with the existing TCP loss recovery algorithms easily. Using emulated experiments we showed that with the proposed optimization techniques the existing loss recovery algorithms can at most achieve 98.3% bandwidth utilization during loss recovery phase, and reduce RTT by at most 80% after loss recovery phase.
Zhongbin Zha, Ke Liu 0004, Binzhang Fu, Mingyu Chen 0001
SECON2
2014 Uplink delay variation compensation in queue length estimation over mobile data networks
abstract
Knowledge of the queue length and link buffer size of the bottleneck link has many potential applications such as congestion control, traffic engineering, traffic policing, content adaptation, QoS monitoring and provisioning, etc. A recent work proposed a new Sum-of-Delay (SoD) algorithm which can accurately estimate such link properties in mobile networks. This work presents a delay variation compensation algorithm - Sum-of-Delay with Timestamp (SoD-TS), to further improve SoD's estimation accuracy in networks with large uplink delay variations. The proposed algorithm exploits the existing TCP Timestamp option which is widely deployed and thus can be implemented without modification to the receiver's TCP implementation. By making use of SoD-TS we developed a novel transport layer application - queue-length-based congestion control algorithm (QCC), on top of TCP that tackles the problem of bufferbloat to provide better QoS for applications requiring low end-to-end delay and/or high bandwidth utilization. Trace-driven simulation shows that SoD-TS can effectively eliminate estimation errors caused by uplink delay variations. Moreover, by applying QCC to the widely deployed TCP CUBIC it can reduce the RTT by a factor of over 3.6 while still achieving over 87% bandwidth utilization.
Ke Liu 0004, Jack Y. B. Lee
ICCCN1
2014 Achieving high throughput and low delay by accurately regulating link queue length over mobile data network
abstract
Knowledge of the queue length of the bottleneck link has many potential applications such as congestion control, traffic engineering, traffic policing, content adaptation, QoS monitoring and provisioning, etc. A recent work proposed a new Sum-of-Delay with Timestamp (SoD-TS) algorithm which can accurately estimate queue length in mobile networks with both bandwidth variations and uplink delay variations by exploiting the existing TCP Timestamp option. By making use of SoD-TS we developed a novel transport protocol - queue-length-aware TCP (TCP-QLA), that tackles the problem of bufferbloat to provide better QoS for applications requiring low end-to-end delay and/or high bandwidth utilization. Trace-driven simulation shows that compared to TCP CUBIC TCP QLA can reduce the RTT by a factor of 2.7 while still achieving over 97% bandwidth utilization. Moreover, TCP-QLA further reduces RTT by 50% compared to delay-based TCP such as FAST TCP and TCP Vegas.
Ke Liu 0004, Jack Y. B. Lee
WiMob1
2014 On Queue Length and Link Buffer Size Estimation in 3G/4G Mobile Data Networks
abstract
The emerging mobile data networks fueled by the world-wide deployment of 3G, HSPA, and LTE networks created new challenges for the development of Internet applications. Unlike their wired counterpart, mobile data networks are known to exhibit highly variable bandwidth. Moreover, base stations are often equipped with large buffers to absorb bandwidth fluctuations to prevent unnecessary packet losses. Consequently to optimize protocol performance in mobile data networks it is essential to be able to accurately characterize two key network properties: queue length and buffer size of the bottleneck link. This work tackles the challenge in estimating these two network properties in modern mobile data networks. Using extensive trace-driven simulations based on actual bandwidth trace data measured from production mobile data networks, we show that existing queue-length and link buffer size estimation algorithms no longer work well in bandwidth-varying networks. We develop a novel sum-of-delays algorithm which incorporates the effect of bandwidth variations into its estimation. Extensive trace-driven simulation results show that it can accurately estimate the queue length and link buffer size under both fixed and varying bandwidth conditions, outperforming existing algorithms by up to two orders of magnitude.
Stanley C. F. Chan, K. M. Chan, Ke Liu 0004, Jack Y. B. Lee
IEEE Trans. Mob. Comput.3
2013 Improving TCP performance over mobile data networks with opportunistic retransmission
abstract
Recent advances in high-speed mobile networks have revealed new bottlenecks in the ubiquitous TCP protocol deployed in the Internet. In addition to differentiating random loss from congestion loss, our experiments revealed that TCP's flow control mechanism can become a significant bottleneck during TCP's loss recovery phase, resulting in bandwidth efficiency as low as 0.05. To tackle this problem we develop a novel opportunistic retransmission algorithm to enable the TCP sender to transmit new packets even before the loss recovery phase is completed, resulting in significant improvement in TCP's bandwidth efficiency. We applied the proposed mechanism to three TCP variants and developed system models to analyze their performance. Experimental results showed that TCP's throughput can be improved by up to 56% in real-world settings.
Ke Liu 0004, Jack Y. B. Lee
WCNC1
2011 Mobile accelerator: A new approach to improve TCP performance in mobile data networks
abstract
This paper investigates the performance of TCP in mobile data networks and proposes a novel approach to address the problem of optimizing transport protocol for these networks - a network-centric approach. We propose to realize protocol optimizations within the network by means of deploying a network-layer device - called a mobile accelerator, which performs protocol optimizations on-the-fly to improve transport protocol performance, without requiring any modification to the server or the client protocol implementations in the OS. Our extensive experiments conducted in production 3G/HSPA and LTE networks show that the proposed mobile accelerator can increase the throughput performance of TCP by over 200%. This mobile accelerator can be readily deployed in existing mobile data networks and can improve the network performance of all existing network applications running atop TCP.
Ke Liu 0004, Jack Y. B. Lee
IWCMC1