Haipeng Du

dblp:150/8534 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-1120-7096ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Concern-Decoupled Architecture for Scenario-Optimized Congestion Control in Large-Scale Live Video CDNs
abstract
Optimizing quality of experience (QoE) for live video streaming (LVS) remains a long-standing challenge for content delivery network (CDN) providers. Today’s CDNs predominantly employ static congestion control (CC) configurations, yet the heterogeneity of LVS scenarios undermines the efficacy of a uniform CC solution applicable to all users, as evidenced by our production measurements. While learning-based CC approaches show promise, they often suffer from limited generalizability, non-transparent design, and high computational overhead, struggling in large-scale CDNs. In this paper, we propose BIFROST, a new CC architecture grounded in a concerndecoupling paradigm. BIFROST separates fixed control logic from scenario-specific parameter optimization, enabling CDN systems to adapt dynamically to diverse LVS scenarios. To materialize it at CDN scale, BIFROST introduces a bilateral collaboration mechanism that identifies the LVS scenarios of each session by extracting client-side characteristics from viewing requests. It further employs offline model-based optimization to derive more effective control parameters for each scenario. We have deployed BIFROST on Alibaba Cloud’s production CDN for nearly a year, serving a commercial LVS application. BIFROST has markedly improved QoE metrics by 2.7% to 32.1%, reinforcing the competitive edge in the CDN market. We also share our experiences and lessons learned from its large-scale deployment.
Danfu Yuan, Yubing Qiu, Weizhan Zhang, Haipeng Du
IEEE Trans. Netw.5
2026 CollabVisAdapt: Spatio-Temporal Context-Aware Adaptation of Shared Object Visualization for MR Telecollaboration
abstract
Mixed Reality (MR) telecollaboration aims to enable users to share local objects as real-time synchronized virtual replicas to remote partners and collaborate on physical tasks as if they were co-located. However, in everyday scenarios with mobile and easy-to-setup MR devices, visualizing shared objects in a single modality, ranging from 2D images to 3D reconstruction, struggles to simultaneously optimize all the aspects of Spatiality, Fidelity, and Real-time performance. To overcome this issue, existing methods explore integrating multiple visualization modalities to leverage their respective advantages in subsets of the three aspects. However, they focus on fixed modality combinations without considering user-centered task contexts and workflow, where users may prioritize different aspects of the visualization across task phases. Moreover, they lack support for switching or require manual switching across modalities, which could become disruptive and tiring. In this paper, we propose adapting object visualization based on spatiotemporal contexts in telecollaboration. Specifically, we first couple task type with the user's relative viewing distance as the spatial context, and examine its impact on users' prioritized visualization aspects, and the corresponding switching thresholds. With differing generation speeds of modalities, we then explore temporal switching schemes when the preferred modality is not immediately available. With the obtained design choices, we implement CollabVisAdapt, a proof-of-concept prototype that supports automatic adaptation of object visualization based on spatiotemporal contexts in MR telecollaboration. A user study in remote maintenance verifies the effectiveness of the proposed workflow with adaptive visualization and the usability of the system.
Xuanyu Wang 0001, Weizhan Zhang, Shuaichen Guo, Caixia Yan, Shuming Yang, Haipeng Du, Wangdu Chen, Qi Wang 0180
IEEE Trans. Vis. Comput. Graph.7
2025 Efficient Real-Time On-Mobile Video Super-Resolution with Automatic Evolutionary Neural Architecture Search
Xuncheng Liu, Weizhan Zhang, Caixia Yan, Haipeng Du
ICANN (2)5
2025 CFNet: An Efficient Convolutional Neural Network With Channel-Focused Convolution and Cross-Channel Mixing
Haipeng Du, Yudeng Xin, Jiageng Zhang
PRCV (4)3
2025 Lightweight Configuration Adaptation With Multi-Teacher Reinforcement Learning for Live Video Analytics
abstract
The proliferation of video data and advancements in Deep Neural Networks (DNNs) have greatly boosted live video analytics, driven by the growing video capture capabilities of mobile devices. However, resource limitations necessitate the transmission of endpoint-collected videos to servers for inference. To meet real-time requirements and ensure accurate inference, it is essential to adjust video configurations at the endpoint. Traditional methods rely on deterministic strategies, posing difficulties in adapting to dynamic networks and video content. Meanwhile, emerging learning-based schemes suffer from trial-and-error exploration mechanisms, resulting in a concerning long-tail effect on upload latency. In this paper, we propose a novel lightweight and robust configuration adaptation policy (LCA), which fuses heuristic and RL-based agents using multi-teacher knowledge distillation (MKD) theory. Firstly, we propose a content-sensitive and bandwidth-adaptive RL agent and introduce a Lyapunov-based optimization agent for ensuring latency robustness. To leverage both agents' strengths, we design a feature-guided multi-teacher distillation network to transfer their advantages to the student. The experimental results across two vision tasks (pose estimation and semantic segmentation) demonstrate that LCA significantly reduces transmission latency compared to prior work (average reduction of 47.11%-89.55%, 95-percentile reduction of 27.63%-88.78%) and computational overhead while maintaining comparable inference accuracy.
Yuanhong Zhang, Weizhan Zhang, Muyao Yuan, Caixia Yan, Tieliang Gong, Haipeng Du
IEEE Trans. Mob. Comput.7
2024 DualConvNet: Enhancing CNN Inference Efficiency Through Compressed Convolutions and Reparameterization
abstract
Convolutional Neural Networks (CNNs) have significantly advanced computer vision tasks, but their increasing complexity poses challenges for efficient inference, particularly on resource-constrained devices. We present DualConvNet, a novel CNN architecture that enhances inference efficiency through two key innovations: compressed convolutions and reparameterization. Compressed convolutions reduce computational complexity by selectively processing input channel subsets during both training and inference. For inference, we introduce a reparameterization technique that merges the multi-branch structure into a single, efficient operation, significantly improving speed. Experiments on CIFAR-10, CIFAR-100, and ImageNet-1k demonstrate DualConvNet’s effectiveness, consistently outperforming state-of-the-art models in both accuracy and inference speed. On the COCO dataset, DualConvNet shows competitive accuracy in object detection and instance segmentation tasks while substantially reducing GPU latency. Ablation studies validate the impact of our dual-strategy approach, revealing significant improvements in both accuracy and computational efficiency compared to alternative designs. These results demonstrate DualConvNet’s effectiveness in improving inference efficiency while maintaining high accuracy across various tasks and datasets, making it particularly suitable for real-time applications in resource-constrained scenarios.
Haipeng Du, Muyan Jiao, Jiageng Zhang
TrustCom1
2024 FHVAC: Feature-Level Hybrid Video Adaptive Configuration for Machine-Centric Live Streaming
abstract
With the widespread deployment of edge computing, the focus has shifted to machine-centric live video streaming, where endpoint-collected videos are transmitted over networks to edge servers for analysis. Unlike maximizing user's Quality of Experience (QoE), machine-centric video streaming optimizes the machine's Quality of Inference (QoI) by balancing the inference accuracy, inference delay, and transmission latency with video adaptive configuration. Traditional heuristic configuration adaption methods are reliable but unable to respond to erratic network fluctuations. Reinforcement learning (RL) based algorithms exhibit superior flexibility but suffer from exploration mechanisms, resulting in long-tail effects on upload latency. In this paper, we propose FHVAC, which dynamically selects video encoding parameters for live streaming by coherently fusing rule-based and RL-based agent at the feature level. We initially develop a robust rule-based approach for ensuring the low latency in transmission, and employ imitation learning to convert it into a neural network equivalently. Subsequently, we design a novel module to combine the two approaches and assess various fusion mechanisms. Our evaluation of FHVAC across two vision tasks (pose estimation and semantic segmentation) in two scenarios (trace-driven simulation and testbed-based experiment) shows that FHVAC enhances the average QoI, and reduces 10.61%-65.27% latency tail performance compared to prior work.
Yuanhong Zhang, Weizhan Zhang, Haipeng Du, Caixia Yan, Li Liu 0036
IEEE Trans. Parallel Distributed Syst.3
2022 Deep Reinforcement Learning Based Adaptive 360-degree Video Streaming with Field of View Joint Prediction
abstract
With the development of 360-degree video and HTTP adaptive streaming (HAS), tile-based adaptive 360-degree video streaming has become a promising paradigm for reducing the bandwidth consumption of delivering the panoramic video content. However, there are two main challenges for the adaptive 360-degree video streaming, accurate long-term prediction of the future field of view (Fo V) and optimal adaptive bitrate (ABR) transmission strategy. In this paper, we propose an attention-based multi-user Fo V joint prediction approach to improve the accuracy, establishing a probability model of watching video tiles for users and applying Long Short-Term Memory (LSTM) network and DBSCAN clustering method. Furthermore, we present an adaptive 360-degree video streaming approach based on deep reinforcement learning (DRL), using A3C algorithm to optimize the QoE. The real-world trace-driven experiments demonstrate that our approach achieves about 8 % gains on user Fo V prediction precision and an increase at least 20 % on user QoE compared with the benchmarks.
Yuanhong Zhang, Junquan Liu, Haipeng Du, Weizhan Zhang
ISCC4
2022 PRIOR: deep reinforced adaptive video streaming with attention-based throughput prediction
abstract
Video service providers have deployed dynamic video bitrate adaptation services to fulfill user demands for higher video quality. However, fluctuations and instability of network conditions inhibit the performance promotion of adaptive bitrate (ABR) algorithms. Existing rule-based approaches fail to guarantee accurate throughput estimates, and learning-based algorithms are considerably sensitive to the variability of network. Therefore, how to gain effective and stable throughput estimates has become one of the critical challenges to further enhancing ABR methods. To eliminate this concern, we propose PRIOR, an ABR algorithm that fuses an effective throughput prediction module and a state-of-the-art multi-agent reinforcement learning method to provide a high quality of experience (QoE). PRIOR aims to maximize the QoE metric by straightforwardly utilizing accurate throughput estimates rather than past throughput measurements. Specifically, PRIOR employs a light-weighted prediction module with attention mechanism to obtain effective future throughput. Considering the excellent features introduced by the HTTP/3 protocol, we apply PRIOR to trace-driven simulations and real-world scenarios over HTTP/1.1 and HTTP/3. Trace-driven emulation illustrates that PRIOR outperforms existing ABR schemes over HTTP/1.1 and HTTP/3, and our prediction module can also reinforce the performance of other ABR algorithms. Extensive results on real-world evaluation demonstrate the superiority of PRIOR over existing state-of-the-art ABR schemes.
Danfu Yuan, Yuanhong Zhang, Weizhan Zhang, Xuncheng Liu, Haipeng Du
NOSSDAV5
2021 Adaptive Video Streaming Using Dynamic Server Push over HTTP/2
abstract
With the increasing popularity of video services, HTTP adaptive streaming (HAS) has become the mainstream technology for media streaming distribution. In traditional HAS over HTTP/1.1, the HAS server responds to each request from the client individually. This process adds additional round-trip time, resulting in underestimation of available bandwidth. As a result, the HAS client chooses a lower bitrate, which reduces network utilization and the user's quality of experience. In recent years, the HTTP/2 protocol has emerged, which allows server to actively push multiple data segments to the client. Pushing multiple segments can reduce the negative impact of network latency on estimating available bandwidth, thereby increasing the user's request bitrate and video quality. However, when the network is unstable, the more video segments that are pushed by the server, the more challenges the client encounters in responding to network fluctuations i n time, causing playback stalling and poor user experience. Therefore, this paper proposes a dynamic server push algorithm over HTTP/2, which chooses a different number of segments for server push according to network fluctuations. For the evaluation results, relative to its benchmarks, the proposed approach improves the average video request bitrate while minimizing the probability of playback stalling.
Shouqin Huang, Weizhan Zhang, Haipeng Du
CSCWD4
2021 Dynamic Push for HTTP Adaptive Streaming with Deep Reinforcement Learning
Haipeng Du, Danfu Yuan, Weizhan Zhang
ICPADS1
2021 QoE-driven HAS Live Video Channel Placement in the Media Cloud
abstract
HTTP adaptive streaming (HAS) technology has been increasingly employed by video service providers (VSPs) due to its prominent benefits such as reducing interruptions of video playback and achieving higher bandwidth utilization and outstanding quality of experience (QoE). And many VSPs have deployed HAS applications in the media cloud to provide large-scale video streaming services. At present, research into the media cloud typically focuses on the management and optimization of cloud resources, such as the placement and migration of virtual machines in media cloud data centers. However, considering the HASlive video streamingservice, existing related works have not adequately discussed the specific impact of the consumption of computing and bandwidth resources of media cloud servers on the user experience (QoE), particularly under the resource constraints in the media cloud. In this paper, we first investigate and formulate the computing and bandwidth resource consumption characteristics of HAS live video streaming with different frame rates and resolutions, and we further establish a resources-aware QoE model to quantify the user experience oflive video channels(i.e., programs). Then, based on the model, we present a QoE-driven HAS live video channel placement approach (including a placement algorithmHCPand a rescheduling algorithmHCR) to optimize the channel allocation in media cloud servers, aiming to maximize the average user QoE. We abstract the maximization problem into an MMKP problem, and employ a heuristic solution to address this problem. The experimental results demonstrate the effectiveness of our proposed approach compared with benchmark solutions.
Junquan Liu, Weizhan Zhang, Shouqin Huang, Haipeng Du
IEEE Trans. Multim.4
2020 Deploying Fused Sharable Video Interaction Channels in Mobile Cloud
abstract
A plethora of mobile video applications involving all aspects of social life are becoming increasingly prevalent. However, for resource-hungry and delay-sensitive multi-view video applications such as multi-channel video conferencing and 3D videos, the hardware resources of mobile terminals becomes a bottleneck in concurrently decoding multiple videos. Fusing multiple views into a single-view video stream in the cloud before transmission can unload the computation of mobile terminals. But the delay increment caused by video fusing in this cloudbased multi-view video streaming makes the strategy of deploying such applications a new area to examine. In this paper, in the interest of deployment with higher resource utilization and less latency, first, together with a load model of the fused sharable video interaction process, a channel admission control algorithm that targets maximizing the supportive capacity of the cloud is introduced. Subsequently, a channel deployment algorithm based on load coordination between the cloud and mobile clients is proposed to ensure that the delay caused by cloud processing is acceptable while minimizing the terminal computing load.
Xuanyu Wang 0001, Haipeng Du, Weizhan Zhang
GLOBECOM2
2018 Integrated Bandwidth Variation Pattern Differentiation for HTTP Adaptive Streaming over 4G Cellular Networks
abstract
HTTP adaptive streaming (HAS) is the state-of-the-art technology for improving the quality of user experience under conditions of time-varying available bandwidth. Developing the bitrate adaptation algorithm becomes more challenging with the transition to 4G cellular networks. The features of bandwidth variation caused by changes in radio channel quality are significantly different from the pattern caused by changes in radio channel resources. In this paper, we propose a bitrate adaptation algorithm for 4G cellular networks with bandwidth variation pattern differentiation. By investigating the bandwidth data profiles collected from an actual 4G cellular network with the field test approach, we distinguish the pattern of bandwidth capacity variations as sustained fluctuations and instantaneous hopping. With bandwidth variation pattern differentiation, the proposed algorithm performs a smoothed bitrate adaptation to sustained fluctuations and an instant bitrate adaptation to instantaneous hopping. Performance evaluations obtained on a 4G cellular network testbed demonstrate that the algorithm achieves reductions in bitrate switching frequency and playback stalls while increasing the average bitrate.
Haipeng Du, Weizhan Zhang, Xuanyu Wang 0001
IPCCC1
2017 LTE-EMU: A High Fidelity LTE Cellar Network Testbed for Mobile Video Streaming
Haipeng Du, Weizhan Zhang, Yunhui Huang
Mob. Networks Appl.1
2009 Transmission Cost of P2P Multicasting
abstract
In peer-to-peer streaming system, multicasting tree is essential to application performance and network transmission resource usage. Currently, available multicasting tree construction strategy optimize for source to client delay, load balance, etc. But network resource occupation of these strategy is not carefully studied. In this paper, network resource usage of application is analyzed. Then, metrics of network resource usage is proposed. Accordingly, Transmission Resource Saving Tree(TRST) is proposed. In order to make our experimental reliable, transmission delay and jitter are sampled from PlanetLab. Contrast experiments is carried out. According to the results, our proposed strategy consumes less network resources to serve the same set of peers.
Weimei Lv, Haipeng Du, Chen Chen 0022
CCNC5