VLDB 2026 Research / reviewers in the wild / expert
Qinghua Wu 0004
dblp:48/86-4
· DBLP profile ↗
38ranked-venue papers
2as first author
29since 2021 · last 2026
0000-0001-5526-4984ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 26 · 2 first-author · 19 since 2021Systems, architecture and hardware · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Moirai: Dependency-Impact-Based Communication Scheduling for Multi-job Distributed Deep Learning Cluster
Jingbin Yang, Jinglei Pei, Qinghua Wu 0004, Hongtao Guan |
Euro-Par (2) | 4 |
| 2026 | RLive: Robust Delivery System for Scaling Live Streaming ServicesabstractAs the demand for streaming services surges, content delivery network (CDN) operators face increasing pressure to scale live video delivery without proportionally increasing infrastructure costs. While best-effort edge resources offer a cost-effective extension to traditional CDN capacity, their limited bandwidth and unstable performance pose significant challenges. Our operational experience shows that naively layering such resources onto existing CDN infrastructure falls short in meeting performance and scalability demands. This paper presents RLive, a robust delivery system that scales CDN capacity by integrating best-effort edge resources. RLive features a redundancy-free multi-source data plane to support reliable and cost-efficient live streaming, along with a multi-layer collaborative control plane that combines the global view with local adaptability for scalable user-to-node mapping. Deployed in ByteDance CDN to support large-scale live streaming services with hundreds of millions of daily viewers, RLive has tripled delivery capacity while reducing rebuffering events by 14.9–20.1%. Yu Tian 0014, Gerui Lv, Qinghua Wu 0004, Ruili Fang, Yajie Peng, Zhichen Xue, Chuanqing Lin, Xiaofei Pang, Ri Lu, Zhenyu Li 0001 |
EuroSys | 3 |
| 2026 | GeoOrchestra: Orchestrating Heterogeneous Geo-Distributed Training with Network-Aware Scheduling
Qinghua Wu 0004, Jingbin Yang, Jinglei Pei, Zhenyu Li 0001 |
SIGCOMM | 2 |
| 2026 | Breath: Adaptive Protection Boundary in FEC Encoding for Mobile Real-Time Video StreamingabstractMobile real-time video streaming (RTVS) demands ultra-low latency to preserve content timeliness. Packet loss in mobile networks significantly inflates frame latency and thus degrades the quality of experience (QoE). As a promising solution, Forward Error Correction (FEC) encoding has been widely deployed in RTVS systems to recover from packet loss by introducing redundancy. However, existing schemes focus on per-frame FEC protection, failing to optimize QoE because they cannot precisely allocate redundancy to handle burst loss events. These events typically occur at the single-frame level, but can be smoothed out at the multi-frame level. We propose Breath, an adaptive FEC scheme that dynamically adjusts the protection boundary based on network and video dynamics. We have implemented Breath in a RTVS system and evaluated it in emulated mobile networks using network traces collected from the production system. Results show that, compared to state-of-the-art FEC schemes, Breath reduces deadline missing rate by 17.2%-22.5% while improving the average video bitrate by 10.6%-14.2%. Shiyang Huang, Gerui Lv, Yuankang Zhao, Qingyue Tan, Congkai An, Xinyi Zhang 0004, Qinghua Wu 0004, Zhenyu Li 0001 |
WWW | 9 |
| 2026 | ADePT: Latency prediction for edge CDN traffic scheduling via causal inferenceabstractTo satisfy the unprecedented Quality of Experience (QoE) and stringent latency Service Level Agreements (SLAs) of emerging interactive applications, modern Content Delivery Networks (CDNs) are deploying massively decentralized edge nodes. However, this paradigm shift poses a significant challenge: optimal traffic scheduling fundamentally depends on acquiring real-time, full-coverage end-to-end path latency data to prevent SLA violations. Current measurement methods cannot scale to monitor every possible user-to-node path, and traditional prediction approaches (e.g., relying on additional segmented measurements or low-rank matrix decomposition) fail to achieve satisfactory accuracy on the resulting extremely sparse datasets. In this work, we present ADePT (Application Delay PredicTion), a novel data-driven causal inference framework that provides comprehensive and precise latency predictions without requiring additional measurements. By explicitly decoupling user-side temporal variations (e.g., last-mile congestion) and node-side spatial variations (e.g., core propagation delays), ADePT successfully extracts high-dimensional latent embeddings from limited measurement data to infer the end-to-end path latency for any potential scheduling decision. Evaluated on a massive real-world dataset from a leading edge CDN, ADePT reduces prediction errors by 19.6% and achieves a median absolute error of 4.3 ms. Consequently, integrating ADePT’s accurate predictions into CDN traffic scheduling significantly improves scheduling decisions, increasing the ratio of traffic meeting strict applications’ latency requirements by 1.66 × . Chuanqing Lin, Gerui Lv, Yangguang Liang, Fuhua Zeng, Qinghua Wu 0004, Zhenyu Li 0001, Gaogang Xie |
Comput. Networks | 6 |
| 2026 | Understanding and Taming the Inflated Latency in Mobile Cloud RenderingabstractLow-latency cloud rendering enables mobile users to experience high-quality, real-time 3D graphics but achieving low Motion-to-Photon (MTP) latency while maintaining smooth playback is a significant challenge. Our real-world measurement study identifies Receive-to-Composition (R2C) latency, caused by ineffective jitter buffer management, as the primary factor contributing to increased MTP latency. To address this, we introduce JitBright, a client-side optimization strategy that dynamically reduces MTP latency through adaptive jitter buffer management. By adjusting buffer levels based on smoothing playback probability and implementing proactive keyframe requests to mitigate frame dependency, JitBright minimizes both active and passive waiting times. Our large-scale evaluation, conducted over 591,000 sessions across diverse network conditions (WiFi, 4G, 5G) and device types, demonstrates significant improvements in user experience. JitBright reduces median R2C latency by up to 87.5%, increases the proportion of sessions meeting strict MTP latency requirements by 6%–27%, and decreases the video freeze rate from 2.4%–2.8% to 0.4%–1.0%. Yuankang Zhao, Qinghua Wu 0004, Gerui Lv, Furong Yang, Jiuhai Zhang, Yanmei Liu, Zhenyu Li 0001, Ying Chen 0011, Gaogang Xie |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Bridge the Gap Between QoS and QoE in Mobile Short Video Service: A CDN Perspective
Chuanqing Lin, Yangguang Liang, Fuhua Zeng, Zhipeng Huang 0026, Yu Tian 0014, Gerui Lv, Qinghua Wu 0004, Zhenyu Li 0001, Gaogang Xie |
NPC (1) | 9 |
| 2025 | Predictable Real-Time Video Latency Control with Frame-Level CollaborationabstractReal-time video (RTV) systems place high demands on ultra low-latency (i.e., less than 100 ms). However, our large-scale measurements reveal that a significant portion of users still experience high video frame latency due to bandwidth jitters. Existing solutions attempt to mitigate this issue by lowering the sender's future video frame encoding bitrate. Nevertheless, as shown in our controlled experiments, they fail to drain existing packets queued on the bottleneck node (i.e., the 5G base station and Wi-Fi access point), still suffering from high tail latency as bandwidth decreases. In this paper, we propose Co-RTV, a collaborative RTV system that achieves predictable latency control. Specifically, Co-RTV enables endpoint-network collaboration between the bottleneck node and the sender. The collaboration speeds up the release of packets queued at the bottleneck node and facilitates accurate latency control at the RTV sender through scalable QoE-driven flow control. Extensive experiments in emulated networks and on a 5G testbed demonstrate the superior performance of Co-RTV, with tail latency reductions of 69.1% and 70.5%, respectively. Qinghua Wu 0004, Gerui Lv, Wenji Du, Qingyue Tan, Wanghong Yang, Yuankang Zhao, Yongmao Ren, Zhenyu Li 0001, Gaogang Xie |
RTSS | 2 |
| 2025 | MARC: Motion-Aware Rate Control for Mobile E-commerce Cloud Rendering
Yuankang Zhao, Furong Yang, Gerui Lv, Qinghua Wu 0004, Yanmei Liu, Jiuhai Zhang, Yutang Peng, Ying Chen 0011, Zhenyu Li 0001, Gaogang Xie |
USENIX ATC | 4 |
| 2024 | Deadline-oriented Flow Control for Real-time UHD Videos in 5G Edge NetworksabstractAccess networks, even with advanced 5G technology, often face bottlenecks when supporting concurrent real-time Ultra High Definition (UHD) video streams with high bandwidth and low latency (e.g., under 10 ms of one-way delay) requirements. Traditionally, end systems employ a combination of flow and congestion control mechanisms to control the sending rate to avoid overwhelming the receiver and the network. However, such control efforts induce prolonged tail delays, thereby sharply reducing the number of UHD video streams meeting delivery deadlines, and sometimes even zero. These outcomes are largely due to the inaccurate network status estimation associated with the control mechanisms. To address this challenge, we propose CFC, a deadline-oriented flow control mechanism that employs cross-layer status estimation to maximize user satisfaction with deadlines. CFC accurately assesses cross-layer information, including flow status and 5G access network status at minimal expense, thus ensuring the deadlines through effective concurrent flow control. Our experiments, conducted in both simulation and testbed settings, demonstrate significant improvements in delay and load-balancing for both reliable and unreliable transmissions. Wanghong Yang, Wenji Du, Baosen Zhao, Tingting Yuan 0001, Yongmao Ren, Qinghua Wu 0004, Xiaoming Fu 0001 |
ICCCN | 7 |
| 2024 | VAKY: Scheduling In-network Aggregation for Distributed Deep Training AccelerationabstractDistributed machine learning (DML) has recently experienced widespread application. A major performance bottleneck is the costly communication for gradients synchronization. Recently, researchers have explored the use of programmable switches for in-network synchronous aggregation of gradients to mitigate the communication overhead. Nevertheless, the performance of in-network synchronous aggregation is significantly impacted by the stragglers. Unfortunately, the schedulers in existing DML systems are no longer effective in dealing with stragglers because of the ignorance of the aggregation progress that is offloaded from the parameter servers to the programmable switches. To address this gap, this paper presents VAKY, an adaptive scheduler specifically designed for in-network aggregation. At the heart of VAKY is the variable K-block sync method, where the aggregators stop waiting for updates from more workers once having received updates from the fastest K workers for each block of gradients. We propose an efficient solution that can dynamically choose the optimal values of K during the training process, in order to minimize the expected training completion time. We have integrated VAKY into PyTorch, and our experiments show that compared to the state-of-the-art in-network aggregation systems, VAKY improves the aggregation throughput by up to $40 \%$ and reduces the training time by $25 \%$. Penglai Cui, Jianer Zhou, Qinghua Wu 0004, Zhaohua Wang, Zhenyu Li 0001 |
ICPADS | 4 |
| 2024 | Accurate Bandwidth Prediction for Real-Time Media Streaming with Offline Reinforcement LearningabstractIn real-time communication (RTC) systems, accurate bandwidth prediction is crucial for encoding and transmission strategies to optimize users' quality of experience (QoE) in various network environments. In this paper, we propose an offline reinforcement learning (RL) method to predict bandwidth for RTC video streaming. We use a representative algorithm, named Implicit Q-Learning (IQL), to train the model. To improve the performance, we carefully preprocess the given dataset and redesign the neural network structure and the reward function. Ablation studies are performed to verify our design choices. Furthermore, compared to a baseline method and six behavior policies, our method reduces the mean squared error (MSE) by 18%-22%, demonstrating high prediction accuracy. Our proposed method won the first prize in ACM MMSys 2024 Grand Challenge on Offline Reinforcement Learning for Bandwidth Estimation in Real Time Communications. The source code is available at https://github.com/n13eho/Schaferct. Qingyue Tan, Gerui Lv, Zejun Yang, Qinghua Wu 0004 |
MMSys | 7 |
| 2024 | Chorus: Coordinating Mobile Multipath Scheduling and Adaptive Video StreamingabstractIncreasing bandwidth demands of mobile video streaming pose a challenge in optimizing the Quality of Experience (QoE) for better user engagement. Multipath transmission promises to extend network capacity by utilizing multiple wireless links simultaneously. Previous studies mainly tune the packet scheduler in multipath transmission, expecting higher QoE by accelerating transmission. However, since Adaptive BitRate (ABR) algorithms overlook the impact of multipath scheduling on throughput prediction, multipath adaptive streaming can even experience lower QoE than single-path. This paper proposes Chorus, a cross-layer framework that coordinates multipath scheduling with adaptive streaming to optimize QoE jointly. Chorus establishes two-way feedback control loops between the server and the client. Furthermore, Chorus introduces Coarse-grained Decisions, which assist appropriate bitrate selection by considering the scheduling decision in throughput prediction, and Finegrained Corrections, which meet the predicted throughput by QoE-oriented multipath scheduling. Extensive emulation and real-world mobile Internet evaluations show that Chorus outperforms the state-of-the-art MPQUIC scheduler, improving average QoE by 23.5% and 65.7%, respectively. Gerui Lv, Qinghua Wu 0004, Yanmei Liu, Zhenyu Li 0001, Qingyue Tan, Furong Yang, Ying Chen 0011, Gaogang Xie |
MobiCom | 2 |
| 2024 | JitBright: towards Low-Latency Mobile Cloud Rendering through Jitter Buffer OptimizationabstractLow-latency cloud rendering services use high-performance servers to provide mobile device users with exquisite graphics and convenient access experiences. Due to the complexity of the system and the diversity of impacting factors, identifying system bottlenecks has become a significant challenge. To demystify system performance, we build an online cloud rendering system to measure the latency distribution of its key components. Yuankang Zhao, Qinghua Wu 0004, Gerui Lv, Furong Yang, Jiuhai Zhang, Yanmei Liu, Zhenyu Li 0001, Ying Chen 0011, Gaogang Xie |
NOSSDAV | 2 |
| 2024 | TECC: Towards Efficient QUIC Tunneling via Collaborative Transmission Control
Furong Yang, Qinghua Wu 0004, Yuanbo Zhang, Yanmei Liu, Zhenyu Li 0001 |
NSDI | 4 |
| 2024 | A multipath scheduler based on cross-layer information for low-delay applications in 5G edge networks
Baosen Zhao, Wanghong Yang, Wenji Du, Yongmao Ren, Jianan Sun, Qinghua Wu 0004 |
Comput. Networks | 6 |
| 2024 | DRTP: A generic Differentiated Reliable Transport Protocol
Yongmao Ren, Anmin Xu, Yifang Qin, Qinghua Wu 0004, Mohamed Ali Kâafar, Gaogang Xie |
Comput. Commun. | 8 |
| 2024 | Accurate Throughput Prediction for Improving QoE in Mobile Adaptive StreamingabstractVideo streaming is the most important mobile application today. To improve users’ quality of experience (QoE), the client player runs adaptive bitrate (ABR) algorithms that dynamically select the bitrate for video chunks based on throughput or delivery time predictions. This paper aims to design an accurate predictor for mobile adaptive streaming by investigating all its components, including input features, output target, and mapping function. We construct the first theoretical framework that reveals potential factors affecting chunk throughput and delivery time. To verify this framework, we provide formulation analysis and measurement observations based on 2500+ video sessions collected in real-world mobile networks. We find that previous works have failed to achieve accurate prediction due to overlooking the impact of the transport mechanism and application behavior on throughput. Furthermore, we show that throughput is a better target for data-driven predictors than delivery time, due to the long-tailed distribution of delivery time. Based on the above, we propose Lumos, a decision-tree-based throughput predictor that can be integrated into various ABR algorithms. Extensive experiments in real-world mobile Internet show that Lumos achieves high prediction accuracy and improves the QoE of MPC by 6.3%, and MPC+Lumos outperforms Pensieve by 19.2%. Gerui Lv, Qinghua Wu 0004, Qingyue Tan, Weiran Wang 0005, Zhenyu Li 0001, Gaogang Xie |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | Disco: A Framework for Dynamic Selection of Multipath Congestion Control AlgorithmsabstractMany mobile devices are usually equipped with multiple interfaces, providing the opportunity of using multipath transport protocols such as Multipath TCP (MPTCP) to boost performance. The multipath congestion control algorithm (CCA) in MPTCP plays a vital role in achieving high performance and multipath fairness in mobile environments where the paths are often heterogeneous and dynamic. Such environments are very challenging for existing one-size-fits-all CCAs to achieve high performance while ensuring multipath fairness. In this paper, we present a novel framework, Disco, to dynamically select the most appropriate CCAs for MPTCP subflows at runtime according to the perceived network condition. Extensive experiments show that compared with existing multipath CCAs, the proposed solution can improve the average throughput by 19% – 25% and reduce the average queuing delay by up to 21 % while it barely does harm to multipath fairness. Furong Yang, Zhenyu Li 0001, Jianer Zhou, Xinyi Zhang 0004, Qinghua Wu 0004, Giovanni Pau 0001, Gaogang Xie |
ICNP | 5 |
| 2023 | A Large-Scale Measurement and Optimization of Mobile Live Streaming ServicesabstractMobile Live Streaming (MLS) services are one of the most popular types of mobile apps. They involve a (often amateur) user broadcasting content to a potentially large online audience via unreliable networks. Nevertheless, we still lack a deep understanding of MLS user behavior that is critical for optimizing MLS systems, despite some active measurements on viewer-side behavior. Using detailed logs obtained from a major MLS provider, this paper first conducts an in-depth measurement study of both viewer-side and broadcaster-side behavior. Key findings include large wasteful uploads, strong viewing locality, and traffic dominance of loyal viewers. Specifically, 33.3% of uploads go unwatched, and the viewership of broadcasters tends to be localized. Inspired by our findings, we propose EDGEOPT– a centralized control center for MLS services for optimizing both the first-mile and the last-mile transmission in MLS. Specifically, EDGEOPT reduces wasteful uploading by 71% through adaptive uploading and enhances the replay quality of popular video segments by 10% via highlights retransmission. EDGEOPT also uses a learning-based content pre-fetching scheme that boosts the viewing startup by 29.5% and offloads at most 80% of the viewing workload from the edge servers with peer-assisted delivery. Zhenyu Li 0001, Jinyang Li 0009, Qinghua Wu 0004, Gareth Tyson, Gaogang Xie |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | A Machine Learning-Based Framework for Dynamic Selection of Congestion Control AlgorithmsabstractMost congestion control algorithms (CCAs) are designed for specific network environments. As such, there is no known algorithm that achieves uniformly good performance in all scenarios for all flows. Rather than devising a one-size-fits-all algorithm (which is a likely impossible task), we propose a system to dynamically switch between the most suitable CCAs for specific flows in specific environments. This raises a number of challenges, which we address through the design and implementation of Antelope, a system that can dynamically reconfigure the stack to use the most suitable CCA for individual flows. We build a machine learning model to learn which algorithm works best for individual conditions and implement kernel-level support for dynamically switching between CCAs. The framework also takes application requirements of performance into consideration to fine-tune the selection based on application-layer needs. Moreover, to reduce the overhead introduced by machine learning on individual front-end servers, we (optionally) implement the CCA selection process in the cloud, which allows the share of models and the selection among front-end servers. We have implemented Antelope in Linux, and evaluated it in both emulated and production networks. The results demonstrate the effectiveness of Antelope via dynamic adjusting the CCAs for individual flows. Specifically, Antelope achieves an average 16% improvement in throughput compared with BBR, and an average 19% improvement in throughput and 10% reduction in delay compared with CUBIC. Jianer Zhou, Xinyi Qiu, Zhenyu Li 0001, Qing Li 0006, Gareth Tyson, Jingpu Duan, Yi Wang 0004, Qinghua Wu 0004 |
IEEE/ACM Trans. Netw. | 9 |
| 2023 | Large-Scale Measurements and Prediction of DC-WAN TrafficabstractLarge cloud service providers have built an increasing number of geo-distributed data centers (DCs) connected by Wide Area Networks (WANs). These DC-WANs carry both high-priority traffic from interactive services and low-priority traffic from bulk transfers. Given that a DC-WAN is an expensive resource, providers often manage it via traffic engineering algorithms that rely on accurate predictions of inter-DC high-priority (delay-sensitive) traffic. In this article, we perform a large-scale measurement study of high-priority inter-DC traffic from Baidu. We measure how inter-DC traffic varies across their global DC-WAN and show that most existing traffic prediction methods either cannot capture the complex traffic dynamics or overlook traffic interrelations among DCs. Building on our measurements, we propose theInterrelated-TemporalGraph ConvolutionalNetwork(IntegNet) model for inter-DC traffic prediction. In contrast to prior efforts, our model exploits both temporal traffic patterns and inferred co-dependencies between DC pairs. IntegNet forecasts the capacity needed for high-priority traffic demands by accounting for the balance between resource provisioning (i.e., allocating resources exceeding actual demand) and QoS losses (i.e., allocating fewer resources than actual demand). Our experiments show that IntegNet can keep a very limited QoS loss, while also reducing overprovisioning by up to 42.1% compared to the state-of-the-art and up to 66.2% compared to the traditional method used in DC-WAN traffic engineering. Zhaohua Wang, Zhenyu Li 0001, Yunfei Chen 0011, Qinghua Wu 0004, Gareth Tyson |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2022 | On Uploading Behavior and Optimizations of a Mobile Live Streaming ServiceabstractMobile Live Streaming (MLS) services are now one of the most popular types of mobile apps. They involve a (often amateur) user broadcasting content to a potentially large online audience via unreliable networks (e.g., LTE). Although prior work has focused on viewer-side behavior, it is equally important to study and improve broadcaster operations. Using detailed logs obtained from a major MLS provider, we first conduct an in-depth measurement study of uploading behavior. Our key findings include large wasteful uploads, strong viewing locality, and traffic dominance of loyal viewers. Specifically, 33.3% of uploads go unwatched, and the viewership of broadcasters tends to be localized to a small set of broadcaster-specific network regions. Inspired by our findings, we propose two system innovations to streamline MLS systems: adaptive uploading and edge server pre-fetching. These optimizations leverage machine learning for reduced waste and improved QoE. Trace-driven experiments show that the adaptive uploading reduces the resources wastage by 63%, and the pre-fetching boosts the startup by 29.5%. Jinyang Li 0009, Zhenyu Li 0001, Qinghua Wu 0004, Gareth Tyson |
INFOCOM | 3 |
| 2022 | Lumos: towards Better Video Streaming QoE through Accurate Throughput PredictionabstractABR algorithms dynamically select the bitrate of chunks based on the network capacity. To estimate the network capacity, most ABR algorithms use throughput prediction while recent works start to leverage delivery time prediction. We in this paper examine all components of the predictor for ABR algorithms, i.e., input features, mapping function and output target. We build an automated video streaming measurement platform, and collect extensive dataset under various network environments, containing 2500+ video sessions. Through analysis, we find that most of previous works failed to achieve accurate prediction due to ignoring how application behavior influences application throughput, e.g., the strong correlation between chunk size and throughput. Then we identify underlying factors affecting this correlation, and consider them as features for more accurate prediction. Moreover, we show that throughput is a better target for data-driven predictors than delivery time in terms of prediction error, due to the long tail distribution of delivery time. Based on those above, we propose a decision-tree-based throughput predictor, named Lumos, which acts as a plug-in for ABR algorithms. Extensive experiments in real-world Internet demonstrate that Lumos achieves high prediction accuracy and improves the QoE of ABR algorithms when integrated into them. Gerui Lv, Qinghua Wu 0004, Weiran Wang 0005, Zhenyu Li 0001, Gaogang Xie |
INFOCOM | 2 |
| 2022 | MD-Roofline: A Training Performance Analysis Model for Distributed Deep LearningabstractDue to the bulkiness and sophistication of the Distributed Deep Learning (DDL) systems, it leaves an enormous challenge for AI researchers and operation engineers to analyze, diagnose and locate the performance bottleneck during the training stage. Existing performance models and frameworks gain little insight on the performance reduction that a performance straggler induces. In this paper, we introduce MD-Roofline, a training performance analysis model, which extends the traditional rooftine model with communication dimension. The model considers the layer-wise attributes at application level, and a series of achievable peak performance metrics at hardware level. With the assistance of our MD-Roofline, the AI researchers and DDL operation engineers could locate the system bottleneck, which contains three dimensions: intra-GPU computation capacity, intra-GPU memory access bandwidth and inter-GPU communication bandwidth. We demonstrate that our performance analysis model provides great insights in bottleneck analysis when training 12 classic CNNs. Tianhao Miao, Qinghua Wu 0004, Penglai Cui, Zhenyu Li 0001, Gaogang Xie |
ISCC | 2 |
| 2022 | LiveNet: a low-latency video transport network for large-scale live streamingabstractLow-latency live streaming has imposed stringent latency requirements on video transport networks. In this paper we report on the design and operation of the Alibaba low-latency video transport network, LiveNet. LiveNet builds on a flat CDN overlay with a centralized controller for global optimization. As part of this, we present our design of the global routing computation and path assignment, as well as our fast data transmission architecture with fine-grained control of video frames. The performance results obtained from three years of operation demonstrate the effectiveness of LiveNet in improving CDN performance and QoE metrics. Compared with our prior state-of-the-art hierarchical CDN deployment, LiveNet halves the CDN delay and ensures 98% of views do not experience stalls and that 95% can start playback within 1 second. We further report our experiences of running LiveNet over the last 3 years. Jinyang Li 0009, Zhenyu Li 0001, Ri Lu, Jufeng Chen, Chunli Zong, Aiyun Chen, Qinghua Wu 0004, Gareth Tyson, Hongqiang Harry Liu |
SIGCOMM | 10 |
| 2022 | Modeling and Optimizing the Scaling Performance in Distributed Deep Learning TrainingabstractDistributed Deep Learning (DDL) is widely used to accelerate deep neural network training for various Web applications. In each iteration of DDL training, each worker synchronizes neural network gradients with other workers. This introduces communication overhead and degrades the scaling performance. In this paper, we propose a recursive model, OSF (Scaling Factor considering Overlap), for estimating the scaling performance of DDL training of neural network models, given the settings of the DDL system. OSF captures two main characteristics of DDL training: the overlap between computation and communication, and the tensor fusion for batching updates. Measurements on a real-world DDL system show that OSF obtains a low estimation error (ranging from 0.5% to 8.4% for different models). Using OSF, we identify the factors that degrade the scaling performance, and propose solutions to effectively mitigate their impacts. Specifically, the proposed adaptive tensor fusion improves the scaling performance by 32.2%∼ 150% compared to the constant tensor fusion buffer size. Tianhao Miao, Qinghua Wu 0004, Zhenyu Li 0001, Guangxin He, Jiaoren Wu, Shengzhuo Zhang, Xingwu Yang, Gareth Tyson, Gaogang Xie |
WWW | 3 |
| 2022 | BBRv2+: Towards balancing aggressiveness and fairness with delay-based bandwidth probing
Furong Yang, Qinghua Wu 0004, Zhenyu Li 0001, Yanmei Liu, Giovanni Pau 0001, Gaogang Xie |
Comput. Networks | 2 |
| 2021 | Examination of WAN traffic characteristics in a large-scale data center networkabstractLarge cloud service providers have built an increasing number of geo-distributed data centers (DCs) connected by WAN to host their diverse services. While we have seen a large body of work on traffic engineering of WAN, the WAN traffic characteristics of production DC networks remain not well understood. In this paper, we report on the network traffic observed in Baidu's DC network (DCN) that consists of tens of geo-distributed DCs. Baidu hosts both traditional services like Web and Computing, as well as emerging services, such as Analytics, AI, and Map. We analyze WAN traffic characteristics in Baidu's DCN from the perspectives of traffic demands, traffic communication among DCs, and traffic characteristics of diverse services. Specifically, we focus on the disparity that might exist among different types of services. We also discuss the implications of our findings for WAN traffic engineering, fabric design, and service deployment. Zhaohua Wang, Zhenyu Li 0001, Yunfei Chen 0011, Qinghua Wu 0004 |
Internet Measurement Conference | 5 |
| 2019 | A Data-Driven Approach to Client-Transparent Access Selection of Dual-Band WiFiabstractDual-band WiFi which supports both 2.4 GHz and 5 GHz has been widely deployed, aiming to expand wireless capacity, and eliminate serious interference in 2.4 GHz. As the proportion of dual-band APs and clients increase enormously, how to select which band to access to achieve considerable user experience is becoming essential in wireless network. Clients' native decisions that tend to prefer 5 GHz will consequently cause serious interference in 5 GHz and leave 2.4 GHz notably idle. Through analyzing a unique dataset in the wild, we quantitatively study the impact of various WiFi factors on the wireless delay. We propose a decision tree approach to intelligent access selection that decides which band to access dynamically according to prior learned schemes. A prototype of the access selection system named LazyAS, which only requires modification at AP side, is realized and deployed in a production WiFi network. Evaluation results demonstrate that LazyAS reduces the 90th percentile of wireless delay in the production WiFi network from 32 ms to 12 ms. Jun Zhang 0033, Guangxing Zhang, Qinghua Wu 0004, Binbin Liao, Gaogang Xie |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2019 | TCP Stalls at the Server Side: Measurement and MitigationabstractTCP is an important factor affecting user-perceived performance of Internet applications. Diagnosing the causes behind TCP performance issues in the wild is essential for better understanding the current shortcomings in TCP. This paper presents a TCP flow performance analysis framework that classifies causes of TCP stalls. The framework forms the basis of a tool that we use to analyze packet-level traces of three services (cloud storage, software download, and web search) deployed by a popular service provider. We find that as many as 20% of the flows are stalled for half of their lifetime. Network-related causes, especially timeout retransmissions, dominate the stalls. A breakdown of the causes for timeout retransmission stalls reveals that double retransmission and tail retransmission are among the top contributors. The importance of these causes depends however on the specific service. Based on these observations, we propose smart-retransmission time out (S-RTO), a mechanism that mitigates timeout retransmission stalls through careful and gentle aggression for retransmission. S-RTO is evaluated in a controlled network and also in a production network. The results consistently show that it is effective at improving TCP performance, especially for short flows. Jianer Zhou, Zhenyu Li 0001, Qinghua Wu 0004, Peter Steenkiste, Steve Uhlig, Jun Li 0002, Gaogang Xie |
IEEE/ACM Trans. Netw. | 3 |
| 2018 | Access Types Effect on Internet Video Services and Its Implications on CDN CachingabstractVideo providers heavily rely on geographically distributed content distribution networks (CDNs) to place video content as close to users as possible, with an aim of improving video quality and avoiding single point of failure at the server side. The effectiveness of CDNs is mostly dependent on the content consumption patterns. Currently, video providers are offering access to content from different platforms (e.g., mobile devices and PC clients), which might result in distinct video content consumption patterns and finally affect the efficiency of CDN caching. Nevertheless, the access type effect on Internet videos is not well understood. In this paper, using a data set consisting of 26 million video requests of a large-scale commercial video-on-demand system, we study the effect of three main access types, i.e., proprietary software on PC clients, Web browser, and mobile apps. Several observations suggest that access types should be considered carefully in CDN design. In particular, the user engagement, user interests in content, and video popularity dynamics patterns, three important factors for video caching, vary remarkably in the three access types. Leveraging off our findings, we propose an access type-aware CDN caching system that associates a cache for each access type and also several optimizations, including partial caching of videos based on chunk-level caching, cross-platform read-only cache access, and prefiltering of the least popular videos. Trace-driven simulations demonstrate that the access type-aware CDN caching achieves high cache hit rate and, more importantly, greatly reduces the disk load that is measured by the number of cache replacement operations. Gaogang Xie, Zhenyu Li 0001, Mohamed Ali Kâafar, Qinghua Wu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | LazyAS: Client-Transparent Access Selection in Dual-Band WiFiabstractDual-band WiFi which supports both 2.4GHz and 5GHz has been widely deployed, aiming to expand wireless capacity and eliminate serious interference in 2.4GHz. Thus, how to select which band to access to achieve considerable user experience is becoming essential in wireless network. Clients' native decision that always prefers 5GHz will consequently cause serious interference in 5GHz and leave 2.4GHz notably idle. Through analyzing a unique dataset in the wild, we quantitatively study the impact of various WiFi factors on the wireless delay. We propose a decision tree approach for intelligent access selection that decides which band to access dynamically according to prior learned schemes. A prototype of the access selection system named LazyAS, which only requires modification at AP side, is realized and deployed in production WiFi network. Evaluation results demonstrate that our proposed LazyAS reduces the 90th percentile of wireless delay in the production WiFi network from 32ms to 12ms. Jun Zhang 0033, Guangxing Zhang, Qinghua Wu 0004, Gaogang Xie |
ICCCN | 3 |
| 2016 | Privacy-Aware Multipath Video Caching for Content-Centric NetworksabstractThe prevalence of Internet video streaming challenges the design and operation of modern networks. Content centric networking (CCN) has been proposed to address the challenges through ubiquitous in-network caching. While the expected benefits include higher performance and lower bandwidth consumption, CCN introduces new privacy issues at layer 3. This is because adversaries could infer the content consumed by others by checking cached data in routers. In this paper, we first analyze the design space to improve both caching performance and cache privacy for video delivery in CCN. In light of the observation that these two metrics need to be balanced, we propose CodingCache. It adopts network coding and random forwarding to exploit the potentials of multipath routing in CCN to improve both the diversity of cached content along different paths and the anonymity set for consumers. We evaluate CodingCache through extensive experiments based on a real-world topology and a unique data set of video access logs from a large-scale commercial video service. Our results demonstrate that, compared with the existing CCN strategies, CodingCache is able to increase the cache hit rate while also improve the use of caches across the network, together with reasonable cache privacy. Qinghua Wu 0004, Zhenyu Li 0001, Gareth Tyson, Steve Uhlig, Mohamed Ali Kâafar, Gaogang Xie |
IEEE J. Sel. Areas Commun. | 1 |
| 2015 | Demystifying and mitigating TCP stalls at the server sideabstractTCP is an important factor affecting user-perceived performance of Internet applications. Diagnosing the causes behind TCP performance issues in the wild is essential for better understanding the current shortcomings in TCP. This paper presents a TCP flow performance analysis framework that classifies causes of TCP stalls. The framework forms the basis of a tool that is publicly available to the research community. We use our tool to analyze packet-level traces of three services (cloud storage, software download and web search) deployed by a popular Chinese service provider. We find that as many as 20% of the flows are stalled for half of their lifetime. Network-related causes, especially timeout retransmission, dominate the stalls. A breakdown of the causes for timeout retransmission stalls reveals that double retransmission and tail retransmission are among the top contributors. The importance of these causes depends however on the specific service. We also propose S-RTO, a mechanism that mitigates timeout retransmission stalls. S-RTO has been deployed on production front-end servers and results show that it is effective at improving TCP performance, especially for short flows. Jianer Zhou, Qinghua Wu 0004, Zhenyu Li 0001, Steve Uhlig, Peter Steenkiste, Gaogang Xie |
CoNEXT | 2 |
| 2015 | A proactive transport mechanism with Explicit Congestion Notification for NDNabstractNamed Data Networking (NDN) shifts the communication paradigm from the quest of where the content is to what content is to be consumed. In such a new Internet architecture, transmission control mechanisms are of particular importance and have to be carefully designed to enable efficient data transmission. Existing work advocates the use of TCP-like reactive mechanisms for NDN transmission control. In this paper, we show that the statefull and adaptive forwarding properties of NDN makes proactive and efficient mechanisms for transmission control possible. We achieve this by using Explicit Congestion Notifications (ECN), which explicitly notify content consumers about network conditions through the communication path. Specifically, we propose an ECN-based proactive interest-sending rate control mechanism, which aims to achieve a high link utilisation for fast data transmission as well as a low packet dropping rate. To have a globally optimal data transmission, we further propose a smart forwarding mechanism, which locally utilises network-wide information to select the forwarding paths for individual flows. Extensive packet-level simulations in ndnSIM demonstrate that the ECN-based approach, coupled with smart forwarding, outperforms TCP-like reactive mechanisms in terms of link utilisation, packet dropping rate and flow completion time. Jianer Zhou, Qinghua Wu 0004, Zhenyu Li 0001, Mohamed Ali Kâafar, Gaogang Xie |
ICC | 2 |
| 2015 | Video Delivery Performance of a Large-Scale VoD System and the Implications on Content DeliveryabstractVideo delivery performance is the main factor that affects Internet video quality. Characterizing the video delivery performance, especially the delivery throughput, can help content providers as well as Internet service providers (ISPs) in system optimization and network planning. Based on a unique dataset consisting of 20 million video download speed measurements , this paper comprehensively studies the video delivery throughput of a large-scale commercial video-on- demand (VoD) system. We observe that user speed exhibits a large variation over time of day as well as across provincial locations. In particular, the worst performance of day is 30% lower than the peak performance . The analysis also reveals that video download speed has a notable impact on Internet video quality, which in turn influences user engagement . The impact, however, becomes limited when the speed increases beyond a certain threshold, which is mostly dependent on the video encoded bitrates. We further examine the interaction between Internet infrastructure and video delivery throughput using the linear regression model and find that crossing the ISP or regional network border yields 15-20% speed loss. Based on these observations , we finally evaluate the potential of edge caching and hybrid CDN-P2P in the improvement of video download performance and video quality. Zhenyu Li 0001, Qinghua Wu 0004, Kavé Salamatian, Gaogang Xie |
IEEE Trans. Multim. | 2 |
| 2012 | Efficient traffic flow measurement for ISP networksabstractTraffic flow measurement is of great importance to ISPs for various network engineering tasks. An interesting problem is that how to determine the minimum number of links by monitoring which one can obtain the traffic flows of the whole ISP network. Previous works view the problem as Vertex Cover problem. They suffer from high time complexity and redundant monitoring. Different from these works, we study the problem from the perspective of edges and propose two models. The first model, Extended Edge Cover model, can determine the minimum set of monitored links, which are 30% less than that of previous works. The second model, shared-path model, is more suitable when the monitoring resources are limited but one still wants to measure a large part of the networks. Using this method, one can measure 85% of the network by monitoring 5% of links. Finally, we evaluate the performance of the two models through extensive simulations. The experimental results show the effectiveness and robustness of the two models. Qinghua Wu 0004, Zhenyu Li 0001, Gaogang Xie, Kavé Salamatian |
LCN | 1 |