Haodan Zhang

dblp:272/6681 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0003-2499-986XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Range-Level Preloading With Scalable Watch-Time Estimation for Billion-User Streaming Systems
abstract
Short-video platforms have grown rapidly by allowing users to browse rich media content through seamless swiping. However, the inherently random nature of swipe behavior creates significant challenges for bandwidth efficiency and playback continuity, often resulting in stalls and unnecessary data transfers. We present OffLoad, a new preloading framework that enhances bandwidth efficiency and playback quality using range-based downloading, which generalizes traditional chunk-based preloading to arbitrary-length segments for finer-grained control. At the core of OffLoad is a two-dimensional watch-time estimation model that jointly captures user preferences and video characteristics. Guided by this estimator, OffLoad introduces a hybrid preloading algorithm that integrates heuristic rules with a learning-based module trained directly on large-scale production data, enabling strong generalization in deployment. Following extensive system-level optimization, OffLoad has been deployed on a commercial short-video platform for more than six months. Our A/B testing results show that OffLoad increases overall user watch-time by 1.1‰, while simultaneously reducing 0.13% rebuffering events and 4.92% of bandwidth consumption.
Guanyan Peng, Haodan Zhang, Zhen Wang 0071, Pengjin Xie, Liang Liu 0001, Huadong Ma
IEEE Trans. Netw.5
2025 3DGCoding: Novel Framework for 3D Gaussian Video Incremental Training and Coding
abstract
Free-viewpoint videos (FVVs) streaming plays a pivotal role in driving the progress of immersive AR/VR applications. To be streamable, FVVs should simultaneously satisfy perframe decoding, compactness, and real-time rendering. Neural radiance based works suffer from low rendering speed, while previous 3D Gaussians (3DGs) based works necessitate large data volumes and pose challenges for compact representation. To support FVVs streaming, we present 3DGCoding, a novel framework for efficient reconstruction and coding of 3DGs for dynamic scenes. We design a learnable 3D mask to partition 3DGs into static and dynamic components, capturing all types of changes and making 3DGs quantitatively and attributively consistent for coding. Then, we propose a 2D-based coding process to further address inter-frame redundancy, achieving independent decoding and smaller data volumes. Experiments confirm that 3DGCoding could stream with 0.49MB per frame, and achieves competitive performance compared with state-of-the-art FVVs streaming methods.
Peiheng Wang, Haodan Zhang, Quanlu Jia, Jiangkai Wu, Haoyang Wang 0015, Xinggong Zhang
ICME2
2025 DeLoad: Demand-Driven Short-Video Preloading with Scalable Watch-Time Estimation
abstract
Short video streaming has become a dominant paradigm in digital media, characterized by rapid swiping interactions and diverse media content. A key technical challenge is designing an effective preloading strategy that dynamically selects and prioritizes download tasks from an evolving playlist, balancing Quality of Experience (QoE) and bandwidth efficiency under practical commercial constraints. However, real-world analysis reveals critical limitations of existing approaches: (1) insufficient adaptation of download task sizes to dynamic conditions, and (2) watch-time prediction models that are difficult to deploy reliably at scale. In this paper, we propose DeLoad, a novel preloading framework that addresses these issues by introducing dynamic task sizing and a practical, multi-dimensional watch-time estimation method. Additionally, a Deep Reinforcement Learning (DRL)-enhanced agent is trained to optimize the download range decisions adaptively. Extensive evaluations conducted on an offline testing platform, leveraging massive real-world network data, demonstrate that DeLoad achieves significant improvements in QoE metrics (34.4%-87.4% gain). Furthermore, after deployment on a large-scale commercial short-video platform, DeLoad has increased overall user watch-time by 0.9‰ while simultaneously reducing rebuffering events and 3.76% bandwidth consumption.
Guanyan Peng, Haodan Zhang, Zhen Wang 0071, Pengjin Xie, Liang Liu 0001
ACM Multimedia4
2024 SRFC: Scalable Radiance Fields Streaming with Planar Codec
abstract
Volumetric videos afford comprehensive and immer-sive viewing experiences with six degrees of freedom (6DoF) for navigation, allowing users to move freely within a three-dimensional space. Radiance fields (RF) is emerging to reproduce photorealistic 3D scenes, which achieves lighting consistency and realistic transfer between the real and virtual worlds. In this paper, we investigate how to deliver photorealistic volumetric video under the network bandwidth constraints and low-power device. We design a novel Scalable Radiance Fields Video Streaming, SRFC, to enable streaming RF video with planar codec. An Orthogonal Layered Depth Image (OLDI) mapping is introduced to map 3D to view-dependent 2D plane images, in order to reduce bitrates and decoding complexity by the commercial 2D codec. Moreover, to address the issue of deviation of the user viewpoint, we propose a dual-tiered structure in radiance fields and optimizes user's perceptual quality by adaptive bitrate and viewpoint decision. The evaluations demonstrate that SRFC is capable of reducing bandwidth requirements by 4 × and increasing decoding frame rate by nearly 3 × while maintaining satisfactory user perception of video quality.
Quanlu Jia, Haodan Zhang, Haoyang Wang 0015, Jiangkai Wu, Xinggong Zhang, Zongming Guo
ICC2
2023 QUTY: Towards Better Understanding and Optimization of Short Video Quality
abstract
Short video applications such as TikTok and Instagram have attracted tremendous attention recently. However, it is very limited for industry and academia to understand the user's Quality of Experience (QoE) on short video, let alone how to improve the QoE in short video streaming.
Haodan Zhang, Yixuan Ban, Zongming Guo, Zhimin Xu 0001, Yue Wang 0032, Xinggong Zhang
MMSys1
2023 ActRay: Online Active Ray Sampling for Radiance Fields
abstract
Thanks to the high-quality reconstruction and photorealistic rendering, the Neural Radiance Field (NeRF) has garnered extensive attention and has been continuously improved. Despite its high visual quality, the prohibitive training time limits its practical application. Although significant acceleration has been achieved, it is still far from real-time training, due to the need for tens of thousands of iterations. In this paper, a feasible solution is to reduce the number of required iterations by always training the rays with the highest loss values, instead of the traditional method of training each ray with a uniform probability. To this end, we propose an online active ray sampling strategy, ActRay. Specifically, to avoid the substantial overhead of calculating the actual loss values for all rays in each iteration, a rendering-gradient-based loss propagation algorithm is presented to efficiently estimate the loss values. To further narrow the gap between the estimated loss and the actual loss, an online learning algorithm based on the Upper Confidence Bound (UCB) is proposed to control the sampling probability of the rays, thereby compensating for the bias in loss estimation. We evaluate ActRay on both real-world and synthetic scenes, and the promising results show that it accelerates radiance field training by 6.5x. Besides, we test ActRay under all kinds of radiance field representations (implicit, explicit, and hybrid), proving that it is general and effective to different representations. We believe this work will contribute to the practical application of radiance fields, because it has taken a step closer to real-time radiance field training. ActRay is open-source at: https://pku-netvideo.github.io/actray/.
Jiangkai Wu, Yunpeng Tan, Quanlu Jia, Haodan Zhang, Xinggong Zhang
SIGGRAPH Asia5
2023 RAM360: Robust Adaptive Multi-Layer 360$^\circ$ Video Streaming With Lyapunov Optimization
abstract
Viewport-adaptive streaming approaches are emerging as the most promising way to deliver high-quality 360 videos over mobile networks. However, the viewport prediction is only reliable within a short prediction window, i.e., a short playback buffer, which conicts with maintaining a long buffer to avoid playback rebuffering. To deal with this problem, we present RAM360, a Robust Adaptive Multi-layer 360 video streaming system, to ensure high viewport quality and low stall ratio concurrently. We make three technical contributions. First, we design a QoE-driven robust multi-layer streaming framework, where each chunk is encoded by multiple independent layers with different quality levels. The client dynamically decides which chunk and layer to be downloaded by their QoE contributions. Thus, the base-layer could be prefetched to avoid the risk of stalling while the viewport quality is improved by downloading enhancement layer. Second, we establish a novel QoE model to represent the quality of whole playback session, not that of chunk. It aims to maximize the overall QoE of playback session. Third, we introduce the Lyapunov optimization theory to solve the QoE optimization problem, which is an online algorithm with near-optimality solution. We demonstrate that RAM360 can significantly outperform the existing schemes regarding viewport quality, stall ratio, and QoE through extensive experiments with public datasets.
Haodan Zhang, Yixuan Ban, Zongming Guo, Xinggong Zhang
IEEE Trans. Multim.1
2020 MA360: Multi-Agent Deep Reinforcement Learning Based Live 360-Degree Video Streaming on Edge
abstract
The mobile edge caching has made video service providers deliver live 360-degree videos worldwide. However, these services still suffer from the huge network traffic on the core network due to the spherical nature and the diverse requests generated from large user populations. It is challenging to optimize the Quality of Experience (QoE) and the bandwidth consumption simultaneously under the significant number of users as well as dynamic network and playback status. In this paper, we propose a Multi-Agent deep reinforcement learning based 360-degree video streaming system, named MA360, to tackle this multi-user live 360-degree video streaming problem in the context of the edge cache network. Specifically, MA360 employs the Mean Field Actor-Critic (MFAC) algorithm to make clients collaboratively and distributively request tiles aiming at maximizing the overall QoE while minimizing the total bandwidth consumption. Experiments over real-world datasets show that MA360 can improve the QoE while significantly reducing the bandwidth consumption compared with several state-of-the-art edge-assisted 360-degree video streaming strategies.
Yixuan Ban, Yuanxing Zhang, Haodan Zhang, Xinggong Zhang, Zongming Guo
ICME3
2020 APL: Adaptive Preloading of Short Video with Lyapunov Optimization
abstract
Short video applications, like TikTok, have attracted many users across the world. It can feed short videos based on users' preferences and allow users to slide the boring content anywhere and anytime. To reduce the loading time and keep playback smoothness, most of the short video apps will preload the recommended short videos in advance. However, these apps preload short videos in fixed size and fixed order, which can lead to huge playback stall and huge bandwidth waste. To deal with these problems, we present an Adaptive Preloading mechanism for short videos based on Lyapunov Optimization, also called APL, to achieve near-optimal playback experience, i.e., maximizing playback smoothness and minimizing bandwidth waste considering users' sliding behaviors. Specifically, we make three technical contributions: (1) We design a novel short video streaming framework which can dynamically preload the recommended short videos before the current video is downloaded completely. (2) We formulate the preloading problem into a playback experience optimization problem to maximize the playback smoothness and minimize the bandwidth waste. (3) We transform the playback experience optimization problem during the whole viewing process into a single-step greedy algorithm based on the Lyapunov optimization theory to make the online decisions during playback. Through extensive experiments based on the real datasets that generously provided by TikTok, we demonstrate that APL can reduce the stall ratio by 81%/12% and bandwidth waste by 11%/31% compared with no-preloading/fixed-preloading mechanism.
Haodan Zhang, Yixuan Ban, Xinggong Zhang, Zongming Guo, Zhimin Xu 0001, Shengbin Meng, Yue Wang 0032
VCIP1