EDBT 2026 Demo / reviewers in the wild / expert
Weihe Li
dblp:229/8830
· DBLP profile ↗
37ranked-venue papers
20as first author
31since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 19 · 10 first-author · 15 since 2021Systems, architecture and hardware · 7 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 7 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pluto: Fast and Accurate Persistent Flow Detection in High-Speed Networks via Adaptive Protection
Weihe Li, Jiawei Huang 0001, Tianyue Chu, Qichen Su, Jianxin Wang 0001 |
IWQoS | 1 |
| 2026 | Charon: Stratified Priority Sampling for Differentiated Per-Flow Measurement in High-Speed NetworksabstractPer-flow measurement of priority-heterogeneous traffic underpins cloud service-level agreement (SLA) enforcement, anomaly detection, and distributed AI training in high-speed networks, yet remains challenging in the fast L1/L2-cache memory regime where high-priority flows are vastly outnumbered. We propose Charon, a priority-aware sketch that replaces the structural separation used by prior methods with stratified admission sampling: a single, online-adaptive, parameter-free rule decides whether each packet is admitted to the sketch. Across multiple real-world traces, Charon achieves more than 2× higher detection accuracy for high-priority flows than the best baseline and up to four orders of magnitude lower average error than state-of-the-art priority-aware sketches, with the gap widening as memory tightens, at high processing throughput. The implementation on the industry-grade Tofino switch further demonstrates low resource utilization. Weihe Li, Xicheng Li, Dimitrios P. Pezaros, Paul Patras |
SIGCOMM | 1 |
| 2026 | Cardinality is Not Enough: Super Host Detection via Segmented Cardinality EstimationabstractAccurately detecting super host that establishes connections to a large number of distinct peers is significant for mitigating web attacks and ensuring high quality of web service. Existing sketch-based approaches estimate the number of distinct connections called flow cardinality according to full IP addresses, while ignoring the fact that a malicious or victim super host often communicates with hosts within the same subnet, resulting in high false positive rates and low accuracy. Though hierarchical-structure based approaches could capture flow cardinality in subnet, they inherently suffer from high memory usage. To address these limitations, we propose SegSketch, a segmented cardinality estimation approach that employs a lightweight halved-segment hashing strategy to infer common prefix lengths of IP addresses, and estimates cardinality within subnet to enhance detection accuracy under constrained memory size. Experiments driven by real-world traces demonstrate that, SegSketch improves F1-Score by up to 8.04× compared to state-of-the-art solutions, particularly under small memory budgets. Jiawei Huang 0001, Xianshi Su, Weihe Li, Qichen Su, Jin Ye 0003, Wanchun Jiang, Jianxin Wang 0001 |
WWW | 4 |
| 2026 | Toward QoE-Fairness for Video Streaming Over Heterogeneous Networks: An Innovative Bandwidth Allocation MechanismabstractWith the growing ubiquity of video streaming, ensuring a fair and high quality of experience (QoE) for users has emerged as a shared concern among video content providers. State-of-the-art video delivery systems achieve QoE fairness through bottleneck bandwidth allocation across multiple video streams, all based on the assumption of a unified congestion control (CC) protocol. However, the widespread use of heterogeneous CC protocols on the Internet not only disrupts QoE fairness among video streams but also poses challenges in achieving fast convergence under dynamic bandwidth. To address these issues, we propose a QoE-Fairnessawarebandwidthallocationmechanism called Fabam, which establishes a unified QoE control plane across heterogeneous CC protocols. Fabam constructs independent virtual targets based on the real-time QoE of each video stream to achieve QoE fairness, and offers rapid convergence for the underlying CC protocols to improve efficiency. In addition, we propose a Deep Neural Network (DNN)-based multi-step mapping model aimed at balancing the performance and overhead of Fabam, thereby enhancing its deployment potential in practical applications. We implement Fabam on QUIC and integrate it with Dash.js. The evaluation results demonstrate the significant superiority of Fabam over the state-of-the-art approaches, including an enhancement of 44.01% in QoE fairness and an improvement of 36.39% in QoE efficiency. Meanwhile, Fabam-DNN maintains satisfactory QoE fairness while supporting multiple users at a low cost. Qichen Su, Jiawei Huang 0001, Weihe Li, Tao Zhang 0019, Wanchun Jiang, Jianxin Wang 0001 |
IEEE Trans. Netw. | 3 |
| 2025 | Harmonia: A Swift and Accurate Approximate Data Structure for Real-Time Heavy Flow Detection in High-Speed Networks
Weihe Li, Tianyue Chu, Christos-Savvas Bouganis, Paul Patras |
ADMA (3) | 1 |
| 2025 | ECHO: Effective Coreset-Driven Learning via Hierarchical OptimizationsabstractDespite driving record performance, the increasing reliance of deep learning on ever-larger datasets has led to prohibitively high storage and management costs that threaten continued progress. While coreset selection offers a promising solution to this challenge, existing methods often rely on expensive iterative optimization procedures or fail to select samples that allow strong generalization across tasks. In this work, we introduce ECHO, a coreset construction and augmentation strategy that leverages the relational properties inherent to a dataset to find its most representative samples. Unlike prior methods, our approach constructs a structured graph that encodes intrinsic dataset patterns, based on which influential samples are identified and augmented to maximize generalization performance. Extensive experiments across five benchmark datasets and against eighteen different coreset selection baselines show that ECHO achieves up to 60% accuracy gains under extreme compression, while being orders of magnitude faster than state-of-the-art alternatives. These results establish a new benchmark for data-efficient learning, particularly under tight coreset budgets, and showcase the benefits of structured coreset selection for effective generalization. Alec F. Diallo, Weihe Li, Paul Patras |
ICDM | 2 |
| 2025 | Pallas: A Data-Plane-Only Approach to Accurate Persistent Flow Detection on Programmable Switches in High-Speed NetworksabstractIn high-speed data center networks, persistent flows are repeatedly observed over extended periods, potentially signaling threats such as stealthy DDoS or botnet attacks. Monitoring every flow in production-grade hardware switches that feature limited memory, however, is challenging under typical high flow rates and data volumes. To tackle this, approximate data structures, like sketches, are often employed. Yet many existing methods rely on per-time-window flag resets, which require frequent control-plane interventions that make them unsuitable for high-speed traffic. This paper introduces Pallas, a fully data-plane-implementable sketch for detecting persistent flows in high-speed networks with high accuracy, obviating the need for time-window-based resets. We further propose Opt-Pallas, an enhanced variant of Pallas that improves detection accuracy by incorporating flow arrival patterns. We present a rigorous error bound analysis for both Pallas and Opt-Pallas, along with extensive performance evaluations using a P4-based prototype on an Intel Tofino switch. Pallas scales persistent flow detection to line-rate capacity, while state-of-the-art solutions fail to operate beyond a few Mbps. Our results show that Pallas and Opt-Pallas can accurately detect persistent flows in traffic volumes over 60× higher than those handled by the best existing approach. Additionally, even under low-speed traffic, Pallas and Opt-Pallas achieve 4.21% and 7.85% higher lookup accuracy while consuming only 8.5% and 9.7% of switch resources, respectively. Extensive trace-driven results on a CPU platform further validate the high detection accuracy of Opt-Pallas compared to existing methods. Weihe Li, Beyza Bütün, Tianyue Chu, Marco Fiore 0001, Paul Patras |
ICNP | 1 |
| 2025 | DACC: Data Augmentation for Learning-based Congestion Control
Jiawei Huang 0001, Yijun Li 0002, Shengwen Zhou, Hui Li 0120, Weihe Li, Jingling Liu, Wanchun Jiang |
INFOCOM | 8 |
| 2025 | Pontus: A Memory-Efficient and High-Accuracy Approach for Persistence-Based Item Lookup in High-Velocity Data StreamsabstractIn today's web-scale, data-driven environments, real-time detection of persistent items that consistently recur over time is essential for maintaining system integrity, reliability, and security. Persistent items often signal critical anomalies, such as stealthy DDoS and botnet attacks in web infrastructures. Although various methods exist for identifying such items as well as for determining their frequency, they require recording every item for processing, which is impractical at very high data rates achieved by modern data streams. In this paper, we introduce Pontus, a novel approach that uses an approximate data structure (sketch) specifically designed for the efficient and accurate detection of persistent items. Our method not only achieves fast and precise lookup but is also flexible, allowing for minor modifications to accommodate other types of persistence-based item detection tasks, such as detecting persistent items with low frequency. We rigorously validate our approach through formal methods, offering detailed proofs of time/space complexity and error bounds to demonstrate its theoretical soundness. Our extensive trace-driven evaluations across various persistence-based tasks further demonstrate Pontus's effectiveness in significantly improving detection accuracy and enhancing processing speed compared to existing approaches. We implement Pontus in an experimental platform with industry-grade Intel Tofino switches and demonstrate the practical feasibility of our approach in a real-world memory-constrained environment. Weihe Li, Zukai Li, Beyza Bütün, Alec F. Diallo, Marco Fiore 0001, Paul Patras |
WWW | 1 |
| 2025 | Tile-size aware bitrate allocation for adaptive 360$^{\circ }$ video streaming
Jiawei Huang 0001, Jingling Liu, Feng Gao 0001, Weihe Li, Jianxin Wang 0001 |
Multim. Tools Appl. | 5 |
| 2025 | Pandora: An Efficient and Rapid Solution for Persistence-Based Tasks in High-Speed Data StreamsabstractIn data streams, persistence characterizes items that appear repeatedly across multiple non-overlapping time windows. Addressing persistence-based tasks, such as detecting highly persistent items and estimating persistence, is crucial for applications like recommendation systems and anomaly detection in high-velocity data streams. However, these tasks are challenging due to stringent requirements for rapid processing and limited memory resources. Existing methods often struggle with accuracy, especially given highly skewed data distributions and tight fastest memory budgets, where hash collisions are severe. In this paper, we introduce Pandora, a novel approximate data structure designed to tackle these challenges efficiently. Our approach incorporates the insight that items absent for extended periods are likely non-persistent, increasing their probability of eviction to accommodate potential persistent items more effectively. We validate this insight empirically and integrate it into our update strategy, providing better protection for persistent items. We formally analyze Pandora's error bounds to validate its theoretical soundness. Through extensive trace-driven tests, we demonstrate that Pandora achieves superior accuracy and processing speed compared to state-of-the-art methods across various persistence-based tasks. Additionally, we further accelerate Pandora's update speed using Single Instruction Multiple Data (SIMD) instructions, enhancing its efficiency in high-speed data stream environments. The code for our method is open-sourced. Weihe Li |
Proc. ACM Manag. Data | 1 |
| 2025 | Efficient Sketching for Heavy Item-Oriented Data Stream Mining With Memory ConstraintsabstractAccurate and fast data stream mining is critical to many tasks, including real-time series analysis for mobile sensor data, big data management and machine learning. Various heavy-oriented item detection tasks, such as identifying heavy hitters, heavy changers, persistent items, and significant items, have garnered considerable attention from both industry and academia. Unfortunately, as data stream speeds continue to increase and the available memory, particularly in L1 cache, remains limited for real-time processing, existing schemes face challenges in simultaneously achieving high detection accuracy, memory efficiency, and fast update throughput, as we reveal. To tackle this conundrum, we propose a versatile and elegant sketch framework named Tight-Sketch, which supports a spectrum of heavy-based detection tasks. Recognizing that, in practice, most items are cold (non-heavy/persistent/significant), we implement distinct eviction strategies for different item types. This approach allows us to swiftly discard potentially cold items while offering enhanced protection to hot ones (heavy/persistent/significant). Additionally, we introduce an eviction method based on stochastic decay, ensuring that Tight-Sketch incurs only small one-sided errors without overestimation. To further enhance detection accuracy under extremely constrained memory allocations, we introduce Tight-Opt, a variant incorporating two optimization strategies. We conduct extensive experiments across various detection tasks to demonstrate that Tight-Sketch significantly outperforms existing methods in terms of both accuracy and update speed. Furthermore, by utilizing Single Instruction Multiple Data (SIMD) instructions, we enhance Tight-Sketch's update throughput by up to 36%. We also implement Tight-Sketch on FPGA to validate its practicality and low resource overhead in hardware deployments. Weihe Li, Paul Patras |
IEEE Trans. Computers | 1 |
| 2024 | Achieving QoE Fairness in Video Streaming over Heterogeneous Congestion Control ProtocolsabstractWith the growing ubiquity of video streaming, ensuring a fair and high quality of experience (QoE) for users has emerged as a shared concern among video content providers. State-of-the-art video delivery systems achieve QoE fairness through bottleneck bandwidth allocation across multiple video streaming, all based on the assumption of a unified congestion control (CC) protocol. However, the widespread use of heterogeneous CC protocols on the Internet not only disrupts QoE fairness among video streaming but also poses challenges in achieving fast convergence under dynamic bandwidth. To address these issues, we propose a QoE-Fairness aware bandwidth allocation mechanism called Fabam, which establishes a unified QoE control plane across heterogeneous CC protocols. Fabam constructs independent virtual targets based on the real-time QoE of each video streaming to achieve QoE fairness, and offers rapid convergence for the underlying CC protocols to improve efficiency. We implement Fabam on QUIC and integrate it with Dash.js. The evaluation results demonstrate the significant superiority of Fabam over the state-of-the-art approaches, including an enhancement of 24.48% in QoE fairness and an improvement of 16.63% in QoE efficiency. Qichen Su, Jiawei Huang 0001, Weihe Li, Tao Zhang 0019, Wanchun Jiang, Jianxin Wang 0001 |
IWQoS | 3 |
| 2024 | Stable-Sketch: A Versatile Sketch for Accurate, Fast, Web-Scale Data Stream ProcessingabstractData stream processing plays a pivotal role in various web-related applications, including click fraud detection, anomaly identification, and recommendation systems. Accurate and fast detection of items relevant to such tasks within data streams, e.g., heavy hitters, heavy changers, and persistent items, is however non-trivial. This is due to growing streaming speeds, limited fast memory (L1 cache) available in current systems, and highly skewed item distributions encountered in practice. In effect, items of interest that are tracked only based on their features (e.g., item frequency or persistence value) are susceptible to replacement by non-relevant ones, leading to modest detection accuracy, as we reveal. In this work, we introduce the notion of bucket stability, which quantifies the degree of recorded item variation, and show that this is a powerful metric for identifying distinct item types. We propose Stable-Sketch, an elegant and versatile sketch that exploits multidimensional information, including item statistics and bucket stability, and adopts a stochastic approach to drive replacement decisions. We present a theoretical analysis of the error bounds of Stable-Sketch, and conduct extensive experiments to demonstrate that our solution achieves substantially higher accuracy and faster processing speeds than state-of-the-art sketches in a range of item detection tasks, even with tight memories. We further enhance Stable-Sketch's update throughput with Single Instruction Multiple Data (SIMD) instructions and implement our solution with P4, demonstrating real world deployment viability. Weihe Li, Paul Patras |
WWW | 1 |
| 2024 | A learning-based approach for video streaming over fluctuating networks with limited playback buffers
Weihe Li, Jiawei Huang 0001, Qichen Su, Wanchun Jiang, Jianxin Wang 0001 |
Comput. Commun. | 1 |
| 2024 | Learning Audio and Video Bitrate Selection Strategies via Explicit RequirementsabstractMobile video streaming dominates today's network traffic, and adaptive bitrate (ABR) algorithms have been routinely adopted for transmitting media content across dynamic mobile networks. State-of-the-art ABR algorithms mainly alter video bitrate without considering audio bitrate as they consider the impact on the video negligible due to their small size. However, to bring users an immersive experience, recent content providers have applied high-quality audio with large sizes, like stereophonic sound. Therefore, improper audio bitrate selection will adversely affect video bitrate selection, leading to undesirable audio/video combinations (the highest video quality with the lowest audio quality, and vice versa) and frequent playback interruptions. To address these inefficiencies, we propose a Self-Play reinforcement learning-based Audio-aware ABR algorithm named SPA to learn strategies for audio and video bitrate selections. By learning from explicit goals, SPA can match the actual requirements and attain good performance. By conducting trace-driven and testbed-based experiments, we observe SPA's considerable superiority compared to existing approaches, including reducing the undesirable combinations by up to 34.17× and achieving zero stall time across 88.57% of traces. We also invite 35 volunteers to join a subjective test, and the result shows that 33/35 people consider SPA provides them with a satisfactory viewing experience. Weihe Li, Jiawei Huang 0001, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | Optimizing Video Streaming in Dynamic Networks: An Intelligent Adaptive Bitrate Solution Considering Scene Intricacy and Data BudgetabstractAdaptive Bitrate (ABR) algorithms have become increasingly important for delivering high-quality video content over fluctuating networks. Considering the complexity of video scenes, video chunks can be separated into two categories: those with intricate scenes and those with simple scenes. In practice, it has been observed that improving the quality of intricate chunks yields more substantial improvements in Quality of Experience (QoE) compared with focusing solely on simple chunks. However, the current ABR schemes either treat all chunks equally or rely on fixed linear-based reward functions, which limits their ability to meet real-world requirements. To tackle these limitations, this paper introduces a novel ABR approach called CAST (Complex-scene Aware bitrate algorithm via Self-play reinforcemenT learning), which considers the scene complexity and formulates the bitrate adaptation task as an explicit objective. Leveraging the power of parallel computing with multiple agents, CAST trains a neural network to achieve superior video playback quality for intricate scenes while minimizing playback freezing time. Moreover, we also introduce a new variant of our proposed approach called CAST-DU, to address the critical issue of efficiently managing users' limited cellular data budgets while ensuring a satisfactory viewing experience. Furthermore, we present CAST-Live, tailored for live streaming scenarios with constrained playback buffers and considerations for energy costs. Extensive trace-driven evaluations and subjective tests demonstrate that CAST, CAST-DU, and CAST-Live outperform existing off-the-shelf schemes, delivering a superior video streaming experience over fluctuating networks while efficiently utilizing data resources. Moreover, CAST-Live demonstrates effectiveness even under limited buffer size constraints while incurring minimal energy costs. Weihe Li, Jiawei Huang 0001, Qichen Su, Jingling Liu, Wenjun Lyu, Jianxin Wang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | VASE: Enhancing Adaptive Bitrate Selection for VBR-Encoded Audio and Video Content With Deep Reinforcement LearningabstractAdaptive BitRate (ABR) algorithms have become increasingly prevalent in modern streaming platforms, offering users significant improvements in the Quality of Experience (QoE). With streaming providers like YouTube and Netflix shifting to high-fidelity audio formats such as stereophonic sound and Dolby Atoms, ensuring proper audio and video adaptation has become a critical aspect of modern streaming platforms. Additionally, Variable Bitrate (VBR) encoding has gained great popularity in encoding audio and video content, given its higher quality-to-bits ratio. However, the considerable variability in network bandwidth, in combination with VBR features such as significantly fluctuating audio/video chunk sizes and diverse content complexity, makes existing ABR schemes formidable to make optimal bitrate selection due to their overlook of audio adaptation or oblivious to VBR features. In this paper, we introduce a new ABR approach forVBR-basedAudio-aware videoStrEaming named VASE, which harnesses deep reinforcement learning (DRL) and exploits parallel computing with multiple agents to swiftly and adeptly manage fluctuations in video/audio chunk sizes, network bandwidth, and varying content complexity, all while operating without any assumptions. Besides, two variants are proposed to mitigate the download energy cost and handle audio and video content in finer granularity. Extensive trace-driven, testbed, and subjective evaluations show that our scheme surpasses existing advanced adaptation schemes regarding the overall QoE, effectively demonstrating its superiority. Weihe Li, Jiawei Huang 0001, Qichen Su, Wanchun Jiang, Jianxin Wang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | P-Sketch: A Fast and Accurate Sketch for Persistent Item LookupabstractIn large data streams consisting of sequences of data items, those appearing over a long period of time are regarded as persistent. Compared with frequent items, persistent items do not necessarily hold large amounts of data and thus may hamper the effectiveness of vanilla volume-based detectors. Identifying persistent items plays a crucial role in a range of areas such as fraud detection and network management. Fast detection of persistent items in massive streams is however challenging due to the inherently high data rates, while state-of-the-art persistent item lookup solutions routinely require large enough memory to attain high accuracy, which questions the feasibility of deploying them in practice. In this paper, we introduce P-Sketch, a novel approach to persistent item lookup that achieves high accuracy even with small memory (L1 Cache) budgets and maintains high update speed across different settings. Specifically, we introduce the concept of arrival continuity(hotness)that counts the number of consecutive windows in which an item appears, to effectively protect persistent items from being wrongly replaced by non-persistent ones. Through meticulous data analysis, we also reveal that items with higher persistence tend to possess a stronger hotness than non-persistent ones. Thus, we harness the information of persistence and hotness, and employ a probability-based replacement strategy to achieve a good balance between memory efficiency, lookup accuracy, and update speed. We also present a theoretical analysis of the performance of the proposed P-Sketch. Through trace-driven emulations, we demonstrate that our P-Sketch yields average F1 score and update throughput gains of up to 10.32$\times$and respectively 2.9$\times$, over existing schemes. Lastly, we show how to further boost the P-Sketch’s update speed with Single Instruction Multiple Data (SIMD) instructions. Weihe Li, Paul Patras |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | Tight-Sketch: A High-Performance Sketch for Heavy Item-Oriented Data Stream Mining with Limited Memory SizeabstractAccurate and fast data stream mining is critical and fundamental to many tasks, including time series database handling, big data management and machine learning. Different heavy-based detection tasks, such as heavy hitter, heavy changer, persistent item and significant item detection, have drawn much attention from both the industry and academia. Unfortunately, due to the growing data stream speeds and limited memory (L1 cache) available for real-time processing, existing schemes face challenges in simultaneously achieving high detection accuracy, high memory efficiency, and fast update throughput, as we reveal. To tackle this conundrum, we propose a versatile and elegant sketch framework named Tight-Sketch, which supports a spectrum of heavy-based detection tasks. Considering that most items are cold (non-heavy/persistent/significant) in practice, we employ different eviction treatments for different types of items to discard these potentially cold ones as soon as possible, and offer more protection to those that are hot (heavy/persistent/significant). In addition, we propose an eviction method that follows a stochastic decay strategy, enabling Tight-Sketch to only bear small one-sided errors (no overestimation). We present a theoretical analysis of the error bounds and conduct extensive experiments on diverse detection tasks to demonstrate that Tight-Sketch significantly outperforms existing methods in terms of accuracy and update speed. Lastly, we accelerate Tight-Sketch's update throughput by up to 36% with Single Instruction Multiple Data (SIMD) instructions. Weihe Li, Paul Patras |
CIKM | 1 |
| 2023 | CAST: An Intricate-Scene Aware Adaptive Bitrate Approach for Video Streaming via Parallel Training
Weihe Li, Jiawei Huang 0001, Jingling Liu, Wenlu Zhang, Wenjun Lyu, Jianxin Wang 0001 |
ICA3PP (4) | 1 |
| 2023 | REN: Receiver-Driven Congestion Control Using Explicit Notification for Data CenterabstractIn recent years, receiver-driven transport protocols have been proposed to use proactive congestion control to meet the stringent latency requirements of large-scale applications in data center. However, the receiver-driven proposals face the challenges brought by network dynamic. First, when the bursty flows start, the aggressive and blind line-rate transmission in the first RTT easily leads to persistent queue backlog. Second, when some flows finish transmissions, the remaining ones cannot increase their sending rates to seize the available bandwidth. To address these problems, this article presents a new receiver-driven congestion control design, called REN, which uses the under- and over-utilization notifications from switch to handle the dynamic traffic. With the aid of explicit feedback, REN alleviates the traffic burstiness due to aggressive start, mitigates the conservativeness in utilizing available bandwidth, and still retains the receiver-driven feature to achieve ultra-low latency. We implement the prototype of REN using DPDK. The experimental results of real testbed and large-scale NS2 simulation show that REN effectively reduces the average flow completion time (AFCT) by up to 68% over the state-of-the-art receiver-driven transmission schemes. Jiawei Huang 0001, Jinbin Hu 0001, Weihe Li, Tao Zhang 0019, Jingling Liu, Jianxin Wang 0001, Tian He 0001 |
IEEE Trans. Cloud Comput. | 4 |
| 2023 | RAV: Learning-Based Adaptive Streaming to Coordinate the Audio and Video Bitrate SelectionsabstractMost commercial players adopt adaptive bitrate (ABR) algorithms to dynamically decide each chunk's bitrate based on the perceived network bandwidth and buffer occupancy. However, current ABR algorithms are agnostic of audio bitrate selection since they deem it has negligible influence on video bitrate selection due to small size of audio chunks. Nevertheless, with the development of audio technologies, the bitrate of audio content increases dramatically in recent years. Thus, inappropriate audio selection can significantly affect video selection and deteriorate the viewing experience. To tackle these inefficiencies, we propose a deepReinforcement learning-based ABR algorithm that takesAudio andVideo quality into account (RAV) to circumvent a series of suboptimal performances, like low playback quality, frequent playback interruptions, poor playback smoothness, and undesirable combinations of video and audio chunks. Furthermore, RAV trains a neural network model that automatically outputs the bitrates for future audio and video chunks without relying on any presumptions about the environment, achieving good robustness to a broad spectrum of conditions. By conducting trace-driven and real-world experiments, we demonstrate that RAV significantly ameliorates the average overall viewing quality by 37.96%-118.20% over the state-of-the-art ABR algorithms. In addition, we also conduct subjective experiments by inviting 32 volunteers, and 27/32 users strongly agree that RAV provides them a better viewing experience than existing ABR solutions. Weihe Li, Jiawei Huang 0001, Wenjun Lyu, Baoshen Guo, Wanchun Jiang, Jianxin Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | An Apprenticeship Learning Approach for Adaptive Video Streaming Based on Chunk Quality and User PreferenceabstractVideo traffic has experienced an exponential increase in current years due to the growing ubiquity of mobile equipment and the constant network improvement. Most commercial players employ adaptive bitrate (ABR) algorithms to dynamically choose bitrate for each chunk based on perceived network capacity and buffer occupancy. Unluckily, even though improving the quality of chunks with dynamic scenes can achieve more QoE gain than static scenes, current ABR algorithms usually strive to maximize the average bitrate instead of perceptual quality, leading to the QoE degradation. To overcome this obstacle, we introduce a dynamic-chunk quality-aware adaptive bitrate algorithm through apprenticeship learning called DAVS (Dynamic-chunk qualityAwareVideoStreaming), where higher quality is selected for the dynamic chunks without reducing the quality of static chunks extravagantly. Furthermore, we take the user’s viewing preference into account to make DAVS adapt to the QoE diversity. The experimental results demonstrate that DAVS ameliorates the quality of dynamic chunks and significantly enhances the QoE compared with several representative ABR algorithms. Weihe Li, Jiawei Huang 0001, Shiqi Wang 0012, Chuliang Wu, Sen Liu 0002, Jianxin Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Asymmetry-Aware Load Balancing With Adaptive Switching Granularity in Data CenterabstractDatacenter networks provide large bisection bandwidth by load balancing traffic over rich parallel paths in multi-rooted tree topologies. Nevertheless, production datacenters operate under various path diversities caused by traffic dynamics, hardware failures and heterogeneous switching equipment. Therefore, the load balancing schemes in data center should be resilient to network asymmetry. Prior fine-grained schemes such as RPS and Presto are prone to experience packet reordering problem under asymmetric topology since they split flows into small units which are spread across all parallel paths. The coarse-grained solutions such as ECMP and LetFlow effectively avoid packet reordering, but easily leading to under-utilization of multiple paths. To solve these problems, we propose a load balancing mechanism called AG, which adaptively adjusts switching granularity according to the asymmetric degree of multiple paths. AG increases switching granularity to alleviate packet reordering under large degrees of topology asymmetry, while reducing switching granularity to obtain high link utilization under small degrees of topology asymmetry. Moreover, we design a switch-based scheme which measures the difference of one-way delay of multiple paths to obtain accurate state of topology asymmetry with low overhead. AG is a practical switch-based solution without modification at end hosts. The experimental results of NS2 simulations and real implementation show that AG reduces the average and$99^{th}$flow completion time by up to 54% and 65% compared with the state-of-the-art load balancing schemes, respectively. Jingling Liu, Jiawei Huang 0001, Weihe Li, Jianxin Wang 0001, Tian He 0001 |
IEEE/ACM Trans. Netw. | 3 |
| 2022 | Synthesizing Audio and Video Bitrate Selections via Learning from Actual RequirementsabstractAdaptive bitrate (ABR) algorithms are routinely adopted for transmitting media contents across dynamic networks. State-of-the-art ABR algorithms only adapt to video bitrate without considering audio bitrate adaption as they consider the im-pact on the video to be negligible due to the small size of the audio. However, to bring users an immersive experience, more and more content providers have applied high-quality audio with large sizes, like stereophonic and surround (Dolby Atmos). Therefore, improper audio bitrate selection will ad-versely affect video bitrate selection, leading to undesirable audio/video combinations (the highest video quality with the lowest audio quality, vice versa) and frequent playback inter-ruptions. To address these inefficiencies, we propose a Self-Play reinforcement learning-based Audio-aware ABR algorithm named SPA to learn strategies for audio and video bi-trate selections. Experimental results demonstrate SPA's con-siderable superiority as compared with existing approaches. Weihe Li, Jiawei Huang 0001, Jingling Liu, Feng Gao 0001 |
ICME | 1 |
| 2022 | UA-Sketch: An Accurate Approach to Detect Heavy Flow based on Uninterrupted ArrivalabstractHeavy flow detection in enormous network traffic is a critical task for network measurement. Due to the limited memory size and high link capacity, accurate detection of heavy flows becomes challenging in large-scale networks. Almost all existing approaches of detecting heavy flows use single-dimension statistics of flow size to make flow-replacement decisions. However, under the mass number of small flows, the heavy flows are prone to be frequently and mistakenly replaced, resulting in unsatisfactory accuracy. To solve this problem, we reveal that the number of uninterrupted arrival packets is a useful metric in identifying flow types. We further propose UA-Sketch that expels small flows and protects heavy ones according to the multiple-dimension statistics including both estimated flow size and number of uninterrupted arrival packets. The test results of trace-driven simulations and OVS experiments show that, even under small memory, UA-Sketch achieves higher accuracy than the existing works, with the F1 Score by up to 2.1 ×. Jin Ye 0003, Wenlu Zhang, Guihao Chen, Yuanchao Shan, Yijun Li 0002, Weihe Li, Jiawei Huang 0001 |
ICPP | 7 |
| 2022 | Opportunistic Transmission for Video Streaming over Wild InternetabstractThe video streaming system employs adaptive bitrate (ABR) algorithms to optimize a user’s quality of experience. However, it is hard for ABR algorithms to choose the right bitrate consistently under highly dynamic bandwidth fluctuations in wild Internet. In this article, we propose a building block on the client side named Opportunistic Chunk Replacement Mechanism (OCRM) to help existing ABR algorithms make full use of the available bandwidth to improve the network utilization and viewing experience of users. Specifically, the servers take advantages of the spare bandwidth to opportunistically transmit high-quality chunks (called opportunistic chunks ) with low priority to the client, without incurring any extra delay. Then, the client player replaces the low-quality chunks with the opportunistic ones that have high quality. We compare OCRM with state-of-the-art ABR algorithms by using trace-driven experiments spanning a wide variety of quality of experience metrics and network conditions. The test results show that OCRM effectively achieves high network utilization and improves the user’s viewing experience by up to 35%. Jiawei Huang 0001, Qichen Su, Weihe Li, Zhuoran Liu 0003, Tao Zhang 0019, Sen Liu 0002, Ping Zhong 0002, Wanchun Jiang, Jianxin Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Mitigating Port Starvation for Shallow-buffered Switches in Datacenter NetworksabstractExplicit Congestion Notification (ECN) is widely utilized in modern data centers to achieve low latency and high throughput for various applications. In recent years, however, even with the sustainable growth of link bandwidth in data centers, the switch buffer size does not increase remarkably. Consequently, the standard per-port ECN scheme suffers from excessive packet loss. Though the shared-buffer ECN scheme alleviates the packet loss, we observe that it leads to severe unfairness, which we term as the Port Starvation problem. When flows destined for some ports have aggressively occupied the shared buffer, the later-arrival flows destined for other ports will be ECN-marked unfairly and obtain significantly lower throughput. To address the port starvation problem, we design a buffer-aware fair ECN-marking (BFEM) scheme for shallow-buffered switch. BFEM leverages the shared buffer to reduce packet loss and meanwhile punishes aggressive flows by ECN marking. We evaluate BFEM with both 40Gbps P4 testbed implementation and large-scale NS2 simulation. The test results show that, by improving fairness between egress ports, BFEM increases total link utilization and reduces the average flow completion time by up to 40% compared with the state-of-the-art per-port and shared-buffer ECN marking schemes. Wenjun Lyu, Jiawei Huang 0001, Jingling Liu, Shaojun Zou, Weihe Li, Jianxin Wang 0001, Desheng Zhang 0002 |
ICDCS | 6 |
| 2021 | Adjusting Switching Granularity of Load Balancing for Heterogeneous Datacenter TrafficabstractThe state-of-the-art datacenter load balancing designs commonly optimize bisection bandwidth with homogeneous switching granularity. Their performances surprisingly degrade under mixed traffic containing both short and long flows. Specifically, the short flows suffer from long-tailed delay, while the throughputs of long flows also degrade dramatically due to low link utilization and packet reordering. To solve these problems, we design a traffic-aware load balancing (TLB) scheme to adaptively adjust the switching granularity of long flows according to the load strength of short ones. Under the heavy load of short flows, the long flows use large switching granularity to help short ones obtain more opportunities in choosing short queues to complete quickly. On the contrary, the long flows reroute flexibly with small switching granularity to achieve high throughput. Furthermore, under extremely bursty scenario, we utilize the packet slicing scheme for long flows to release bandwidth for short ones. The experimental results of NS2 simulation and testbed implementation show that TLB significantly reduces the average flow completion time of short flows by 16%-67% over the state-of-the-art load balancers and achieves the high throughput for long flows. Moreover, for extreme bursty case, at the acceptable throughput degradation of long flows, TLB with packet slicing reduces the deadline missing ratio of bursty short flows by up to 80%. Jinbin Hu 0001, Jiawei Huang 0001, Wenjun Lyu, Weihe Li, Wenchao Jiang, Jianxin Wang 0001, Tian He 0001 |
IEEE/ACM Trans. Netw. | 4 |
| 2021 | Mitigating Packet Reordering for Random Packet Spraying in Data Center NetworksabstractModern data center networks are usually constructed in multi-rooted tree topologies, which require the highly efficient multi-path load balancing to achieve high link utilization. Recent packet-level load balancer obtains high throughput by spraying packets to all paths, but it easily leads to the packet reordering under network asymmetry. The flow-level or flowlet-level load balancer avoids the packet reordering, while reducing the link utilization due to their inflexibility. To solve these problems, we design a Queueing Delay Aware Packet Spraying (QDAPS), that effectively mitigates the packet reordering for packet-level load balancer. QDAPS selects paths for packets according to the queueing delay of output buffer, and lets the packet arriving earlier be forwarded before the later packets to avoid packet reordering. Moreover, we adopt the “power-of- n-choices” paradigm on QDAPS to alleviate the impact of herd behavior under multiple forwarding engines. We compare QDAPS with ECMP, LetFlow and RPS through NS2 simulation and Mininet implementation. The test results show that QDAPS reduces flow completion time (FCT) by ~30%-50% over the state-of-the-art load balancing mechanism. Jiawei Huang 0001, Wenjun Lyu, Weihe Li, Jianxin Wang 0001, Tian He 0001 |
IEEE/ACM Trans. Netw. | 3 |
| 2020 | DAVS: Dynamic-Chunk Quality Aware Adaptive Video Streaming using Apprenticeship LearningabstractTo deliver video in a high quality across various network conditions, adaptive bitrate (ABR) algorithms dynamically select bitrate for each chunk according to perceived network rate and buffer occupancy. Unfortunately, though ameliorating the quality of chunks with dynamic scenes can obtain more QoE gain than the ones with static scenes, current ABR algorithms generally aim to maximize the average bitrate rather than perceptual quality, resulting in the QoE degradation. To address this issue, we propose a dynamic-chunk quality aware adaptive bitrate scheme via apprenticeship learning named DAVS, in which higher quality is chosen for the dynamic chunks without decreasing the quality of static chunks excessively. The experimental results show that DAVS enhances the quality of dynamic chunks and greatly improves the overall QoE compared with the state-of-the-art ABR algorithms. Weihe Li, Jiawei Huang 0001, Shiqi Wang 0012, Sen Liu 0002, Jianxin Wang 0001 |
GLOBECOM | 1 |
| 2020 | Pipeline-Based Chunk Scheduling to Improve ABR Performance in DASH SystemabstractTo deliver high quality video across different network conditions, the video chunks are explicitly fetched by client or proactively pushed by server in Dynamic Adaptive Streaming over HTTP (DASH) system. Unfortunately, on the one hand, the client fetch mechanism suffers from bandwidth wastage due to its stop-and-wait fashion when the network delay becomes large. On the other hand, the server push mechanism performs poorly because of its inflexibility in bitrate switching under fluctuating bandwidth. To address these inefficiencies, we propose a pipeline-based chunk scheduling scheme called PCS to auto-turn the sending time of each chunk. For a given ABR algorithm, PCS dynamically pre-schedules the chunk delivery according to the real-time network conditions. Using the pipelined-based chunk delivery, PCS flexibly adjusts the bitrate of each chunk and meanwhile avoids the unnecessary waiting time in the stop-and-wait transmission. The experimental results of testbed implementations show that PCS greatly improves the average bitrate of the state-of-the-art ABR algorithms by up to 26%, and reduces the rebuffer rate by up to 31%. Weihe Li, Jiawei Huang 0001, Shaojun Zou, Zhuoran Liu 0003, Qichen Su, Xuxing Chen, Jianxin Wang 0001 |
ICCCN | 1 |
| 2020 | Achieving High Utilization for Approximate Fair Queueing in Data CenterabstractModern data centers often host multiple applications with diverse network demands. To provide fair bandwidth allocation to several thousand traversing flows, Approximate Fair Queueing (AFQ) utilizes multiple priority queues in switch to approximate ideal fair queueing. However, due to limited number of queues in commodity switches, AFQ easily experiences high packet loss and low link utilization. In this paper, we propose Elastic Fair Queueing (EFQ), which leverages limited priority queues to flexibly achieve both high network utilization and fair bandwidth allocation. EFQ dynamically assigns the free buffer space in priority queues for each packet to obtain high utilization without sacrificing flow-level fairness. The results of simulation experiments and real implementations show that EFQ reduces the average flow completion time by up to 82% over the state-of-the-art fair bandwidth allocation mechanisms. Jingling Liu, Jiawei Huang 0001, Weihe Li, Jianxin Wang 0001 |
ICDCS | 4 |
| 2019 | AG: Adaptive Switching Granularity for Load Balancing with Asymmetric Topology in Data Center NetworkabstractModern data center topologies often take the form of a multi-rooted tree with rich parallel paths to provide high bandwidth. However, various path diversities caused by traffic dynamics, link failures and heterogeneous switching equipments widely exist in production datacenter network. Therefore, the multi-path load balancer in data center should be robust to these diversities. Although prior fine-grained schemes such as RPS and Presto make full use of available paths, they are prone to experience packet reordering problem under asymmetric topology. The coarse-grained solutions such as ECMP and LetFlow effectively avoid packet reordering, but easily lead to under-utilization of multiple paths. To cope with these inefficiencies, we propose a load balancing mechanism called AG, which adaptively adjusts switching granularity according to the asymmetric degree of multiple paths. AG increases switching granularity to alleviate packet reordering under large degrees of topology asymmetry, while reducing switching granularity to obtain high link utilization under small degrees of topology asymmetry. AG is deployed on the switches with negligible overhead, while making no modification on end-hosts. We evaluate AG through both Mininet testbed and large-scale NS2 simulations. The experimental results show that AG reduces the average and 99thflow completion time by up to 51% and 56% over the state-of-the-art load balancing schemes, respectively. Jingling Liu, Jiawei Huang 0001, Weihe Li, Jianxin Wang 0001 |
ICNP | 3 |
| 2019 | TLB: Traffic-aware Load Balancing with Adaptive Granularity in Data Center NetworksabstractModern datacenter topologies typically are multi-rooted trees consisting of multiple paths between any given pair of hosts. Recent load balancing designs focus on making full use of available parallel paths to provide high bisection bandwidth. However, they are agnostic to the mixed traffic generated by diverse applications in data centers and respectively use the same granularity in rerouting flows regardless of the flow type. Therefore, the short flows suffer the long-tailed queueing delay and reordering problems, while the throughputs of long flows are also degraded dramatically due to low link utilization and packet reordering under the non-adaptive granularity. To solve these problems, we design a traffic-aware load balancing (TLB) scheme to adopt different rerouting granularities for two kinds of flows. Specifically, TLB adaptively adjusts the switching granularity of long flows according to the load strength of short ones. Under the heavy load of short flows, the long flows use large switching granularity to help short ones obtain more opportunities in choosing short queues to complete quickly. When the load strength of short flows is low, the long flows switch paths more flexibly with small switching granularity to achieve high throughput. TLB is deployed at the switch, without any modifications on the end-hosts. The experimental results of NS2 simulations and Mininet implementation show that TLB significantly reduces the average flow completion time (AFCT) of short flows by ~15%-40% over the state-of-the-art load balancing schemes and achieves the high throughput for long flows. Jinbin Hu 0001, Jiawei Huang 0001, Wenjun Lv, Weihe Li, Jianxin Wang 0001, Tian He 0001 |
ICPP | 4 |
| 2018 | QDAPS: Queueing Delay Aware Packet Spraying for Load Balancing in Data CenterabstractModern data center networks are usually constructed in multi-rooted tree topologies, which require the highly efficient multi-path load balancing to achieve high link utilization. Recent packet-level load balancer obtains high throughput by spraying packets to all paths, but it easily leads to the packet reordering under network asymmetry. The flow-level or flowlet-level load balancer avoids the packet reordering, while reducing the link utilization due to their inflexibility. To solve these problems, we design a Queueing Delay Aware Packet Spraying (QDAPS), that effectively mitigates the packet reordering for packet-level load balancer. QDAPS selects paths for packets according to the queueing delay of output buffer, and lets the packet arriving earlier be forwarded before the later packets to avoid packet reordering. We compare QDAPS with ECMP, LetFlow and RPS through NS2 simulation and Mininet implementation. The test results show that QDAPS reduces flow completion time (FCT) by ~30%-50% over the state-of-the-art load balancing mechanism. Jiawei Huang 0001, Wenjun Lv, Weihe Li, Jianxin Wang 0001, Tian He 0001 |
ICNP | 3 |