VLDB 2026 Research / reviewers in the wild / expert
Haiping Wang 0002
dblp:68/7989-2
· DBLP profile ↗
12ranked-venue papers
4as first author
6since 2021 · last 2026
0009-0006-2124-5746ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 3 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hermit: A Flow-Collaborative Transport Scheme for Multi-Source Video On-Demand StreamingabstractToday's fast-growing Video-on-Demand (VoD) service needs efficient content delivery to guarantee the user experience. To reduce costs, the industry has been exploring the adoption of unstable, heterogeneous, low-performance edge nodes as cost-efficient alternatives to expensive CDN servers. To compensate for the resulting degradation in user experience, Multi-source Parallel Downloading (MPD) is becoming a new VoD transport paradigm. However, existing transport optimization solutions face performance obstacles when applied to the MPD scenarios. They cannot handle the contention between MPD flows of the same download task, which is likely to occur at the shared last-hop, and lack the ability to quickly adapt to the unstable network environments brought by dynamic, heterogeneous, and low-performance edge nodes. To fill this gap, we propose Hermit, a VoD-oriented MPD transport algorithm. Hermit (1) continuously monitors the state of the flows and makes timely scheduling decisions, and (2) efficiently coordinates across the flows to mitigate self-contention at the shared last hop. As a client-driven scheme, Hermit does not require cumbersome coordination among edge nodes, nor does it increase server complexity. Through extensive experiments on real-world large-scale testbed and locally emulated network conditions, we demonstrate that Hermit can improve the consistent downloading rate by 9.2% to 21.3%. Shaorui Ren, Enhuan Dong, Haiping Wang 0002, Jia Zhang 0010, Zili Meng, Mingwei Xu 0001, Shu Shi, Hebin Yu, Zhichen Xue, Yajie Peng, Xiaofei Pang |
ICC | 4 |
| 2026 | Medley: Optimizing Midgress Bandwidth for Commercial Live Streaming CDNs
Haiping Wang 0002, Wanxin Shi, Sandesh Dhawaskar Sathyanarayana, Shu Shi, Yinghao Yu, La Zuo, Hebin Yu, Ruoshi Sun, Yajie Peng, Xiaofei Pang, Ruili Fang, Zhenpeng Zhu, Yang Xu 0010 |
NSDI | 1 |
| 2025 | ACE: Sending Burstiness Control for High-Quality Real-time CommunicationabstractModern real-time communication (RTC) demands both ultra-low latency and consistently high visual quality. Yet, as content becomes more dynamic and RTTs shrink, we reveal a previously overlooked problem: long-tail queuing latency in the sender's pacing queue between encoder and network. This phenomenon is rooted in a mismatch between the bursty frame stream produced by the encoder and the smooth traffic expected by the network. Existing approaches trying to smoothen the bitrate inevitably force an undesirable trade-off between latency and video quality. To address this, we propose a dual-control approach that manages both the encoding and transmission burstiness. At the sender, we dynamically adjust the bucket size of a token-based pacer to control burstiness at the granularity of frame level. Within the encoder, we introduce an adaptive complexity mechanism that smoothens frame sizes without sacrificing quality. Trace-driven emulation and real-world experiments show our solution ACE reduces end-to-end 95th percentile latency by up to 43% while maintaining superior visual quality versus the state of the art. Xiangjie Huang, Haiping Wang 0002, Hebin Yu, Sandesh Dhawaskar Sathyanarayana, Shu Shi, Zili Meng |
SIGCOMM | 3 |
| 2024 | Magpie: Improving the Efficiency of A/B Tests for Large Scale Video-on-Demand SystemsabstractWith the exponential rise in video traffic, researchers and developers require more effective tools to validate the efficacy of designed algorithms for Video-on-Demand (VoD) system. However, traditional experimental platforms face two main challenges: a lack of realistic testing and the need for longer and significant effort. To overcome these limitations, we propose Magpie, an efficient experimental platform tailored for VoD systems. Magpie leverages a realistic operational setting, rapid testing, and high reproducibility to closely simulate online user environments without impacting production systems. Compared to conventional simulations, our evaluation demonstrates that Magpie reduces the disparity with online experiments by 85.6%. Deployed within our company-a leading video content provider in China-Magpie has efficiently validated over tens of algorithms, with 80% demonstrating enhanced performance in subsequent online tests. Hebin Yu, Haiping Wang 0002, Chenfei Tian, Sandesh Dhawaskar Sathyanarayana, Shu Shi, Zhichen Xue, Shuaixin Yu, Yajie Peng, Xiaofei Pang |
IMC | 2 |
| 2024 | Enhancing Resource Management of the World's Largest PCDN System for On-Demand Video Streaming
Haiping Wang 0002, Shu Shi, Xiaofei Pang, Yajie Peng, Zhichen Xue, Jiangchuan Liu |
USENIX ATC | 2 |
| 2023 | TwinStar: A Practical Multi-path Transmission Framework for Ultra-Low Latency Video DeliveryabstractUltra-low latency video streaming has received explosive growth in the past few years. However, existing methods all focus on single-path transmission, which is ineffective in dealing with really poor network conditions. To tackle their problems, we propose TwinStar, a novel multi-path framework to improve the experience quality of ultra-low latency video. The core idea of TwinStar is to concurrently leverage multiple paths to mitigate the negative impacts of network jitter on a single path. In particular, by carefully designing the video encoding, data allocation and loss recovery, TwinStar is very robust to handle network dynamics and deliver high-quality video services. We have deployed TwinStar in a commercial cloud gaming platform and evaluated it with real-world networks. The extensive experiments demonstrate that TwinStar significantly outperforms the single-path transmission methods, with 91% reduction in stall ratio and 11% improvement in PSNR across all regions. Haiping Wang 0002, Siping Tao, Hebin Yu, Shu Shi |
ACM Multimedia | 1 |
| 2019 | Bubble: Lightweight Core Sharing in NFVabstractMany researches have revealed the requirement of enabling multiple network functions (NFs) to share a CPU core in Network Function Virtualization (NFV) to support fine-grained NF models, efficient resource utilization, and chain consolidation. However, these works usually enable core sharing via kernel-level threads, which incurs significant performance degradation. In this paper, we present Bubble to enable lightweight core sharing in NFV. Bubble leverages user-level threads to eliminate the performance overhead introduced by kernel-level thread scheduling. Bubble is designed to satisfy unique requirements in NFV by providing accurate and low-overhead scheduling, in support of on- demand resource allocation, and accurate NF load measurement. Evaluations over a Bubble prototype implementation demonstrate that Bubble can improve the performance by 1.6Ã- to 6.2Ã- for co-located NFs and by 3.7Ã- to 68.8Ã- for a consolidated Service Function Chain (SFC) in a core against two state- of-the-art solutions. Haiping Wang 0002, Zhilong Zheng, Chen Sun 0005, Jun Bi |
GLOBECOM | 1 |
| 2019 | Octans: Optimal Placement of Service Function Chains in Many-Core SystemsabstractNetwork Function Virtualization (NFV) has the potential to offer service delivery flexibility and reduce overall costs by running service function chains (SFCs) on commodity servers with many cores. Existing solutions for placing SFCs in one server treat all CPU cores as equal and allocate isolated CPU cores to different network functions (NFs). However, advanced servers often adopt Non-Uniform Memory Access (NUMA) architecture to improve the scalability of many-core systems. CPU cores are grouped into nodes, incurring performance bottleneck due to cross-node memory access and intra-node resource contention. Our evaluation shows that randomly selecting cores to place NFs in an SFC could suffer from 39.2% lower throughput comparing to an optimal placement solution. In this paper, we propose Octans, an NFV orchestrator to achieve maximum aggregate throughput of all SFCs in many-core systems. Octans first formulates the optimization problem as a Non-Linear Integer Programming (NLIP) model. Then we identify the key factor for problem solving as evaluating the throughput drop of an NF caused by other NFs in the same SFC or different SFCs, i.e. performance drop index, and propose a formal and precise prediction model based on system level performance metrics. Finally, we propose an efficient heuristic algorithm to quickly find near-optimal placement solutions. We have implemented a prototype of Octans. Extensive evaluation shows that Octans significantly improves the aggregate throughput comparing to two state-of the-art placement mechanisms by 26.7%~51.8%, with very low prediction errors of SFC performance (an average deviation of 2.6%). Moreover, Octans could quickly find a near-optimal placement solution with tiny optimality gap (1.2%~3.5%). Zhilong Zheng, Jun Bi, Heng Yu 0005, Haiping Wang 0002, Chen Sun 0005, Hongxin Hu |
INFOCOM | 4 |
| 2019 | MicroNF: An Efficient Framework for Enabling Modularized Service Chains in NFVabstractThe modularization of service function chains (SFCs) in network function virtualization (NFV) could introduce significant performance overhead and resource efficiency degradation due to introducing frequent packet transfer and consuming much more hardware resources. In response, we exploit the reusability, lightweightness, and individual scalability features of elements in modularized SFCs (MSFCs) and propose MicroNF, an efficient framework for MSFC in NFV. MicroNF addresses the performance overhead and resource efficiency problems in three ways. First, MicroNF graph constructor reuses the processing results of elements from different NFs and reconstructs the MSFC after modularization to shorten the chain latency. Second, optimized placer pays attention to the problem of which elements to consolidate and provides a performance-aware placement algorithm to place MSFCs compactly and optimize the global packet transfer cost. Third, MicroNF individual scaler innovatively introduces a push-aside scaling up strategy to avoid degrading performance and taking up new CPU cores. To support MSFC reusing and consolidation, MicroNF also designs a high-performance infrastructure to efficiently forwarding packets with consistency ensured and to automatically scheduling elements with fairness ensured when the elements are consolidated on the CPU core. Our evaluation results show that MicroNF achieves significant performance improvement and efficient resource utilization on several metrics. Zili Meng, Jun Bi, Haiping Wang 0002, Chen Sun 0005, Hongxin Hu |
IEEE J. Sel. Areas Commun. | 3 |
| 2018 | CoCo: Compact and Optimized Consolidation of Modularized Service Function Chains in NFVabstractThe modularization of Service Function Chains (SFCs) in Network Function Virtualization (NFV) could introduce significant performance overhead and resource efficiency degradation due to introducing frequent packet transfer and consuming much more hardware resources. In response, we exploit the lightweight and individually scalable features of elements in Modularized SFCs (MSFCs) and propose CoCo, a compact and optimized consolidation framework for MSFC in NFV. CoCo addresses the above problems in two ways. First, CoCo Optimized Placer pays attention to the problem of which elements to consolidate and provides a performance-aware placement algorithm to place MSFCs compactly and optimize the global packet transfer cost. Second, CoCo Individual Scaler innovatively introduces a push-aside scaling up strategy to avoid degrading performance and taking up new CPU cores. To support MSFC consolidation, CoCo also provides an automatic runtime scheduler to ensure fairness when elements are consolidated on CPU core. Our evaluation results show that CoCo achieves significant performance improvement and efficient resource utilization. Zili Meng, Jun Bi, Haiping Wang 0002, Chen Sun 0005, Hongxin Hu |
ICC | 3 |
| 2018 | Grus: Enabling Latency SLOs for GPU-Accelerated NFV SystemsabstractGraphics Processing Unit (GPU) has been recently exploited as a hardware accelerator to improve the performance of Network Function Virtualization (NFV). However, GPU-accelerated NFV systems suffer from significant latency variation when multiple network functions (NFs) are co-located in the same machine, which prevents operators from supporting latency Service Level Objectives (SLOs). Existing research efforts to address this problem can only guarantee a limited number of SLOs with very low resource utilization efficiency. In this paper, we present the Grus framework to support latency SLOs in GPU-accelerated NFV systems. Grus thoroughly analyzes the sources of latency variation and proposes three design principles: (1) dynamic batch size setting is needed to bound packet batching latency in CPU; (2) a reordering mechanism for data transfer over PCI-E is required to guarantee the stalling time; and (3) maximizing concurrency in GPU is necessary to avoid NF execution waiting time. Guided by the principles, Grus consists of two logical layers including an infrastructure layer and a scheduling layer. The infrastructure layer is equipped with an in-CPU Reorder-able Worker Pool that could adjust batching size and packet transfer order, and in-GPU Controllable Concurrent Executors to provide maximized concurrency. The scheduling layer runs a heuristic algorithm to perform accurate and fast scheduling to guarantee SLOs based on our prediction models. We have implemented a prototype of Grus. Extensive evaluations demonstrate that Grus can significantly reduce latency variation and satisfy 4.5 × more SLO terms than state-of-the-art solutions. Zhilong Zheng, Jun Bi, Haiping Wang 0002, Chen Sun 0005, Heng Yu 0005, Hongxin Hu, Kai Gao 0001 |
ICNP | 3 |
| 2017 | Joint Optimization in Software Defined Wireless Networks with Network Coded Opportunistic RoutingabstractApplying network coding and opportunistic routing can significantly improve the throughput performance of wireless multi-hop networks, but a mismatch problem still exists among the upper flow rate, routing and lower transmission resource scheduling. This paper models the throughput optimization as a network utility function maximization in wireless multi-hop networks. By applying Lagrangian dual decomposition theory and the sub-gradient method, the total utility maximization problem is decomposed into source rate control, routing and scheduling problems. These three subproblems are solved independently and are linked by queue length to achieve coordination and joint optimization of network throughput. Based on OpenFlow, this paper implements the joint optimization algorithm in a software-defined network and verifies its performance through experiments. The results show that the joint optimization algorithm has better performance in terms of the total network throughput, network transmission efficiency and inter-flow fairness compared with the existing network coded opportunistic routing method, which only considers the optimization of routing. Haiping Wang 0002, Sanfeng Zhang 0002 |
MASS | 1 |