Chao Zhou 0003

dblp:72/4184-3 · DBLP profile ↗
← Back
57ranked-venue papers
13as first author
30since 2021 · last 2026
0000-0003-2969-3042ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 39 · 11 first-author · 19 since 2021Computer networks · 13 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 11 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Systems, architecture and hardware · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Retrieval-augmented image harmonization
Haolin Wang 0004, Ming Liu 0018, Zifei Yan, Chao Zhou 0003, Longan Xiao, Wangmeng Zuo
Pattern Recognit.4
2026 Deep Network-Based Adaptive Quantization for Practical Video Coding
abstract
The optimization of block-level quantization parameters (QP) is critical to improving the performance of practical block-based video compression encoders, but the extremely large optimization space makes it challenging to solve. Existing solutions, e.g. HEVC encoder x265, usually add some optimization constraints of the block-independent assumption and linear distortion propagation model, which limits compression efficiency improvement to a certain extent. To address this problem, a deep learning-based encoder-only adaptive quantization method (DAQ) is proposed in this paper, where a deep network is designed to adaptively model the joint temporal propagation relationship of quantization among blocks. Specifically, DAQ consists of two phases: in the training phase, considering the heavy searching cost of the traditional codec, we introduce a well-designed end-to-end learned block-based video compression network as an effective training proxy tool for the deep encoder-side network. While in the deployment phase, the trained deep network is applied to jointly predict all block QPs in a frame for the traditional encoder. Besides, our network deploys only on the encoder side without changing the standard decoder and has very low inference complexity, making it able to apply in practice. At last, we deploy DAQ in HEVC and VVC encoder for performance comparison, and the experimental results demonstrate that DAQ significantly outperforms practically used x265 with on average 15.0%, 10.9% BD-rate reduction under the SSIM and PSNR, and also achieves 12.5%, 5.0% coding gain than VTM. Moreover, for deploying deep video codec in practice, this work provides a new insight for optimizing the encoder parameters with a large space.
Hewei Liu, Jiawen Gu, Dengchao Jin, Meng Lei, Chao Zhou 0003
IEEE Trans. Circuits Syst. Video Technol.7
2026 LVT: A Learned Video Transcoding Framework
abstract
With the exponential growth of video traffic and the continuous evolution of video coding standards, video transcoding has become essential for existing bitstreams to benefit from the advanced features of new video compression technologies. Typically, video transcoding involves decoding an existing bitstream and re-encoding the decoded sequence into a target format. A key challenge in transcoding is the inevitable presence of compression artifacts in the decoded sequences, which, if not properly addressed, can degrade transcoding efficiency by causing suboptimal bit allocation and disrupting core coding processes. In this article, a learned video transcoding framework (LVT) is proposed to optimize video transcoding, leveraging coding priors from the input bitstream to guide the transcoding process. In the framework, to mitigate the adverse effects of compression artifacts, a Coding Priors-Guided Spatial Feature Transform module is designed, which utilizes coding prior features to adaptively modulate intermediate features through spatial affine transformations, enhancing bit allocation and suppressing artifacts. Additionally, a Coding Priors-Guided Quality Adapter module is proposed to generate a compression degradation representation using coding priors, which dynamically interacts with intermediate features to enable the network to perceive and adapt to different levels of degradation in the input video. Furthermore, a Motion Vectors-Guided Flow Refinement module is proposed to reduce prediction errors caused by artifacts. It refines optical flow predictions by using motion vectors from the bitstream as auxiliary information. Extensive experiments demonstrate that our framework outperforms both existing traditional and learned video codecs in transcoding performance, achieving an average bitrate saving of 20.3% compared to the H.266/VVC reference software VTM under the practical YUV420 setting measured with PSNR.
Nianxiang Fu, Daiqin Yang, Zhenan Lin, Chao Zhou 0003
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Plug-and-Play Tri-Branch Invertible Block for Image Rescaling
abstract
High-resolution (HR) images are commonly downscaled to low-resolution (LR) to reduce bandwidth, followed by upscaling to restore their original details. Recent advancements in image rescaling algorithms have employed invertible neural networks (INNs) to create a unified framework for downscaling and upscaling, ensuring a one-to-one mapping between LR and HR images. Traditional methods, utilizing dual-branch based vanilla invertible blocks, process high-frequency and low-frequency information separately, often relying on specific distributions to model high-frequency components. However, processing the low-frequency component directly in the RGB domain introduces channel redundancy, limiting the efficiency of image reconstruction. To address these challenges, we propose a plug-and-play tri-branch invertible block (T-InvBlocks) that decomposes the low- frequency branch into luminance (Y) and chrominance (CbCr) components, reducing redundancy and enhancing feature processing. Additionally, we adopt an all-zero mapping strategy for high-frequency components during upscaling, focusing essential rescaling information within the LR image. Our T-InvBlocks can be seamlessly integrated into existing rescaling models, improving performance in both general rescaling tasks and scenarios involving lossy compression. Extensive experiments confirm that our method advances the state of the art in HR image reconstruction.
Jingwei Bao, Jinhua Hao, Ming Sun 0008, Chao Zhou 0003, Shuyuan Zhu
AAAI5
2025 KVQ: Boosting Video Quality Assessment via Saliency-guided Local Perception
abstract
Video Quality Assessment (VQA), which intends to predict the perceptual quality of videos, has attracted increasing attention. Due to factors like motion blur or specific distortions, the quality of different regions in a video varies. Recognizing the region-wise local quality within a video is beneficial for assessing global quality and can guide us in adopting fine-grained enhancement or transcoding strategies. Due to the heavy cost of annotating region-wise quality, the lack of ground truth constraints from relevant datasets further complicates the utilization of local perception. Inspired by the Human Visual System (HVS) that links global quality to the local texture of different regions and their visual saliency, we propose a Kaleidoscope Video Quality Assessment (KVQ) framework, which aims to effectively assess both saliency and local texture, thereby facilitating the assessment of global quality. Our framework extracts visual saliency and allocates attention using Fusion-Window Attention (FWA) while incorporating a Local Perception Constraint (LPC) to mitigate the reliance of regional texture perception on neighboring areas. KVQ obtains significant improvements across multiple scenarios on five VQA benchmarks compared to SOTA methods. Furthermore, to assess local perception, we establish a new Local Perception Visual Quality (LPVQ) dataset with region-wise annotations. Experimental results demonstrate the capability of KVQ in perceiving local distortions. KVQ models and the LPVQ dataset will be available at https://github.com/qyp2000/KVQ.
Yunpeng Qu, Kun Yuan 0003, Qizhi Xie, Ming Sun 0008, Chao Zhou 0003
CVPR5
2025 Deep Adaptive Quantization for Practical Video Compression
abstract
In this work, we propose a deep learning-based adaptive quantization method to promote video coding performance. Due to inter-prediction and reference mechanism, the block-level quantization parameter (QP) not only influences current block distortion but also has complex temporal propagation effects on subsequent coding frames. Our idea is to utilize a deep network to model the complex temporal propagation relationship of quantization. As shown in Fig. 1, the deep network directly predicts all block-level QPs of the frame for the traditional encoder without changing the standard decoder. Since our network deploys only on the encoder side and has low inference complexity, it can be easily applied in practice. In addition, we use a learned coding network as a proxy of the traditional codec to train our network.
Hewei Liu, Jiawen Gu, Dengchao Jin, Meng Lei, Chao Zhou 0003
DCC7
2025 Content-Aware Motion Compensated Temporal Filter for Video Coding
abstract
Video coding achieves efficient compression by exploiting the spatial and temporal correlations within the video signal. However, the noise in the source signal corrupts such correlation and impairs the coding performance. Motion compensated temporal filter (MCTF) [1] is a pre-processing tool that removes certain noise from the source signal, thereby enhancing temporal correlations among adjacent frames. Although numerous efforts have been made to optimize MCTF, MCTF still lacks flexibility in filtering for diverse video content, and its filtering efficiency is still limited. In this paper, we propose the Content-Aware MCTF method (CAMCTF) to enhance the filtering adaptability of MCTF for diverse video contents. The CAMCTF adaptively adjusts the filtering block sizes based on the Sum of Square Error (SSE) and Motion Vector (MV) information calculated during the motion estimation (ME) process in MCTF, and offset weighting method is applied to improve the prediction quality of filtering blocks. Performance was evaluated on top of Versatile Video Coding (VVC) reference software VTM-23.4. The VVC Common Test Conditions (CTC) [2] with QPs 22, 27, 32, 37 are used. As shown in Table 1, CAMCTF achieves an overall of 1.07% and 0.79% luma BD-rate gains in Random Access (RA) and Low Delay (LD) configurations, respectively. The encoder complexity is 112% and 110% for RA and LD configurations, respectively, without decoder complexity increasing.
Yunrui Jian, Meng Lei, Weilun Feng, Zhenan Lin, Chao Zhou 0003
DCC7
2025 Content-Adaptive Motion Compensated Temporal Filter for Versatile Video Coding
abstract
The exploitation of spatial and temporal correlation within video signals is a cornerstone of video coding, which is extensively employed to achieve high coding efficiency. The presence of noise in video signal deteriorates the correlations, leading to a degradation in coding efficiency. Motion Compensated Temporal Filter (MCTF) is a pre-processing tool designed to remove noise from the source signal, thereby enhancing temporal correlations among adjacent frames. In this paper, a Content-Adaptive MCTF method (CAMCTF) is proposed to enhance the filtering adaptability of MCTF for diverse video contents. The proposed CAMCTF method consists of Quadtree Block Partitioning Scheme (QBPS), Offset Block Weighted Compensation (OBWC) and Structural Similarity (SSIM) based Filtering Weight Adjustment (SSFWA). Specifically, QBPS generates size-adaptive Motion Compensation Block (MCB) to better cater to the video content characteristics. Subsequently, OBWC is employed to address the uneven prediction quality of MCBs. Furthermore, the filtering weights of each MCB are fine-tuned by SSFWA based on the SSIM index. The simulation results show that the proposed CAMCTF method can achieve 1.23% and 0.94% BD-rate gains for RA and LD configurations, respectively, on top of Versatile Video Coding (VVC) reference software VTM-23.4.
Yunrui Jian, Xueli Cheng, Weilun Feng, Zhenan Lin, Chao Zhou 0003
ICME7
2025 Visual Autoregressive Modeling for Image Super-Resolution
abstract
Image Super-Resolution (ISR) has seen significant progress with the introduction of remarkable generative models. However, challenges such as the trade-off issues between fidelity and realism, as well as computational complexity, have also posed limitations on their application. Building upon the tremendous success of autoregressive models in the language domain, we propose VARSR, a novel visual autoregressive modeling for ISR framework with the form of next-scale prediction. To effectively integrate and preserve semantic information in low-resolution images, we propose using prefix tokens to incorporate the condition. Scale-aligned Rotary Positional Encodings are introduced to capture spatial structures and the diffusion refiner is utilized for modeling quantization residual loss to achieve pixel-level fidelity. Image-based Classifier-free Guidance is proposed to guide the generation of more realistic images. Furthermore, we collect large-scale data and design a training process to obtain robust generative priors. Quantitative and qualitative results show that VARSR is capable of generating high-fidelity and high-realism images with more efficiency than diffusion-based methods. Our codes are released at https://github.com/quyp2000/VARSR.
Yunpeng Qu, Kun Yuan 0003, Jinhua Hao, Kai Zhao 0011, Qizhi Xie, Ming Sun 0008, Chao Zhou 0003
ICML7
2025 Accelerating Diffusion-based Super-Resolution with Dynamic Time-Spatial Sampling
abstract
Diffusion models have gained attention for their success in modeling complex distributions, achieving impressive perceptual quality in SR tasks. However, existing diffusion-based SR methods often suffer from high computational costs, requiring numerous iterative steps for training and inference. Existing acceleration techniques, such as distillation and solver optimization, are generally task-agnostic and do not fully leverage the specific characteristics of low-level tasks like super-resolution (SR). In this study, we analyze the frequency- and spatial-domain properties of diffusion-based SR methods, revealing key insights into the temporal and spatial dependencies of high-frequency signal recovery. Specifically, high-frequency details benefit from concentrated optimization during early and late diffusion iterations, while spatially textured regions demand adaptive denoising strategies. Building on these observations, we propose the Time-Spatial-aware Sampling strategy (TSS) for the acceleration of Diffusion SR without any extra training cost. TSS combines Time Dynamic Sampling (TDS), which allocates more iterations to refining textures, and Spatial Dynamic Sampling (SDS), which dynamically adjusts strategies based on image content. Extensive evaluations across multiple benchmarks demonstrate that TSS achieves state-of-the-art (SOTA) performance with significantly fewer iterations, improving MUSIQ scores by 0.2~3.0 and outperforming the current acceleration methods with only half the number of steps.
Qijie Wang, Ming Sun 0008, Haowei Zhu, Chao Zhou 0003, Bin Wang 0021
IJCAI5
2025 Towards User-level QoE: Large-scale Practice in Personalized Optimization of Adaptive Video Streaming
abstract
Traditional optimization methods based on system-wide Quality of Service (QoS) metrics have approached their performance limitations in modern large-scale streaming systems. However, aligning user-level Quality of Experience (QoE) with algorithmic optimization objectives remains an unresolved challenge. Therefore, we propose LingXi, the first large-scale deployed system for personalized adaptive video streaming based on user-level experience. LingXi dynamically optimizes the objectives of adaptive video streaming algorithms by analyzing user engagement. Utilizing exit rate as a key metric, we investigate the correlation between QoS indicators and exit rates based on production environment logs, subsequently developing a personalized exit rate predictor. Through Monte Carlo sampling and online Bayesian optimization, we iteratively determine optimal parameters. Large-scale A/B testing utilizing 8% of traffic on Kuaishou, one of the largest short video platforms, demonstrates LingXi's superior performance. LingXi achieves a 0.15% increase in total viewing time, a 0.1% improvement in bitrate, and a 1.3% reduction in stall time across all users, with particularly significant improvements for low-bandwidth users who experience a 15% reduction in stall time.
Lianchen Jia, Chao Zhou 0003, Chaoyang Li 0002, Jiangchuan Liu, Lifeng Sun
SIGCOMM2
2024 KVQ: Kwai Video Quality Assessment for Short-form Videos
abstract
Short-form UGC video platforms, like Kwai and TikTok, have been an emerging and irreplaceable mainstream media form, thriving on user-friendly engagement, and kaleidoscope creation, etc. However, the advancing content-generation modes, e.g., special effects, and sophisticated processing workflows, e.g., de-artifacts, have introduced significant challenges to recent UGC video quality assessment: (i) the ambiguous contents hinder the identification of quality-determined regions. (ii) the diverse and complicated hybrid distortions are hard to distinguish. To tackle the above challenges and assist in the development of short-form videos, we establish the first large-scale Kwai short Video database for Quality assessment, termed KVQ, which comprises 600 user-uploaded short videos and 3600 processed videos through the diverse practical processing workflows, including preprocessing, transcoding, and enhancement. Among them, the absolute quality score of each video and partial ranking score among indistinguish samples are provided by a team of professional researchers specializing in image processing. Based on this database, we propose the first short-form video quality evaluator, i.e., KSVQE, which enables the quality evaluator to identify the quality-determined semantics with the content understanding of large vision language models (i.e., CLIP) and distinguish the distortions with the distortion understanding module. Experimental results have shown the effectiveness of KSVQE on our KVQ database and popular VQA databases. The project can be found at https://lixinustc.github.io/projects/KVQ/.
Yiting Lu, Xin Li 0082, Yajing Pei, Kun Yuan 0003, Qizhi Xie, Yunpeng Qu, Ming Sun 0008, Chao Zhou 0003, Zhibo Chen 0001
CVPR8
2024 PTM-VQA: Efficient Video Quality Assessment Leveraging Diverse PreTrained Models from the Wild
abstract
Video quality assessment (VQA) is a challenging problem due to the numerous factors that can affect the perceptual quality of a video, e.g., content attractiveness, distortion type, motion pattern, and level. However, annotating the Mean opinion score (MOS) for videos is expensive and time-consuming, which limits the scale of VQA datasets, and poses a significant obstacle for deep learning-based methods. In this paper, we propose a VQA method named PTM-VQA, which leverages PreTrained Models to transfer knowledge from models pretrained on various pre-tasks, enabling benefits for VQA from different aspects. Specifically, we extract features of videos from different pretrained models with frozen weights and integrate them to generate representation. Since these models possess var-ious fields of knowledge and are often trained with labels irrelevant to quality, we propose an Intra-Consistency and Inter-Divisibility (ICID) loss to impose constraints on features extracted by multiple pretrained models. The intra-consistency constraint ensures that features extracted by different pretrained models are in the same unified quality-aware latent space, while the inter-divisibility introduces pseudo clusters based on the annotation of samples and tries to separate features of samples from different clusters. Furthermore, with a constantly growing number of pretrained models, it is crucial to determine which models to use and how to use them. To address this problem, we propose an efficient scheme to select suitable candidates. Models with better clustering performance on VQA datasets are chosen to be our candidates. Extensive experiments demonstrate the effectiveness of the proposed method.
Kun Yuan 0003, Mading Li, Muyi Sun, Ming Sun 0008, Jiachao Gong, Jinhua Hao, Chao Zhou 0003, Yansong Tang
CVPR8
2024 CPGA: Coding Priors-Guided Aggregation Network for Compressed Video Quality Enhancement
abstract
Recently, numerous approaches have achieved notable success in compressed video quality enhancement (VQE). However, these methods usually ignore the utilization of valuable coding priors inherently embedded in compressed videos, such as motion vectors and residual frames, which carry abundant temporal and spatial information. To remedy this problem, we propose the Coding Priors-Guided Aggregation (CPGA) network to utilize temporal and spatial information from coding priors. The CPGA mainly consists of an inter-frame temporal aggregation (ITA) module and a multi-scale non-local aggregation (MNA) module. Specifically, the ITA module aggregates temporal information from consecutive frames and coding priors, while the MNA module globally captures spatial information guided by residual frames. In addition, to facilitate research in VQE task, we newly construct the Video Coding Priors (VCP) dataset, comprising 300 videos with various coding priors extracted from corresponding bitstreams. It remedies the shortage of previous datasets on the lack of coding information. Experimental results demonstrate the superiority of our method compared to existing state-of-the-art methods. The code and dataset will be released at https://github.com/VQE-CPGA/CPGA.
Jinhua Hao, Yukang Ding, Yu Liu 0091, Qiao Mo, Ming Sun 0008, Chao Zhou 0003, Shuyuan Zhu
CVPR7
2024 Asymmetric Motion Vector Refinement for Future Video Coding
abstract
Efficiently improving the accuracy of motion vector (MV) in merge candidate list is a critical issue in terms of the advanced inter coding technologies. The decoder side motion vector refinement (DMVR) and merge mode with motion vector differences (MMVD) are utilized to refine the MV obtained from merge mode in Versatile Video Coding (VVC). Nevertheless, both DMVR and MMVD operate under the assumption of symmetric motion when adjusting bi-prediction MV. This assumption may result in inaccuracies adjusting where the motion is asymmetric, leading to imprecise MV. To address this issue, we propose the asymmetric motion vector refinement (ASMVR) approach to refine asymmetric motion more accurately for future video coding in this paper. Specifically, ASMVR is formulated by the asymmetric MMVD (asy-MMVD) and asymmetric DMVR (asy-DMVR) schemes, which are compatible with MMVD and DMVR in VVC respectively. Four asymmetric MV refinement templates are devised to capture varying motion scenes, and the optimal one is derived through bilateral matching and rate-distortion optimization. Moreover, meticulously designed fast algorithms are implemented to bypass unnecessary candidate evaluations, thereby effectively reducing both encoding and decoding complexities. The simulation result shows that on top of the VVC Test Model (VTM-22.1), ASMVR achieves 1.52% BD-rate saving for random access (RA).
Yunrui Jian, Zhenan Lin, Meng Lei, Chao Zhou 0003
DCC7
2024 OAPT: Offset-Aware Partition Transformer for Double JPEG Artifacts Removal
Qiao Mo, Yukang Ding, Jinhua Hao, Ming Sun 0008, Chao Zhou 0003, Feiyu Chen 0001, Shuyuan Zhu
ECCV (22)6
2024 A New Dataset and Framework for Real-World Blurred Images Super-Resolution
Ming Sun 0008, Chao Zhou 0003, Bin Wang 0021
ECCV (28)3
2024 XPSR: Cross-Modal Priors for Diffusion-Based Image Super-Resolution
Yunpeng Qu, Kun Yuan 0003, Kai Zhao 0011, Qizhi Xie, Jinhua Hao, Ming Sun 0008, Chao Zhou 0003
ECCV (11)7
2024 Dancing with Shackles, Meet the Challenge of Industrial Adaptive Streaming via Offline Reinforcement Learning
abstract
Adaptive video streaming has been studied for over 10 years and has demonstrated remarkable performance. However, adaptive video streaming is not an independent algorithm but relies on other components of the video system. Consequently, as other components undergo optimization, the gap between the traditional simulator and the real-world system continues to grow which makes the adaptive video streaming algorithm must adapt to these variations.In order to address the challenges facing industrial adaptive video streaming, we introduce a novel offline reinforcement learning framework called Backwave. This framework leverages history logs to reduce the sim-real gap. We propose new metrics based on counterfactual reasoning to evaluate its performance and we integrate expert knowledge to generate valuable data to mitigate the issue of data override. Furthermore, we employ curriculum learning to minimize additional errors.We deployed Backwave on a mainstream commercial short video platform, Kuaishou. In a series of A/B tests conducted nearly one month with over 400M daily watch times, Backwave consistently outperforms prior algorithms. Specifically, Backwave reduces stall time by 0.45% to 8.52% while maintaining comparable video quality and Backwave demonstrates improvements in average play duration by 0.12% to 0.16%, and overall play duration by 0.12% to 0.26%.
Lianchen Jia, Chao Zhou 0003, Tianchi Huang, Chaoyang Li 0002, Lifeng Sun
INFOCOM2
2024 QPT-V2: Masked Image Modeling Advances Visual Scoring
abstract
Quality assessment and aesthetics assessment aim to evaluate the perceived quality and aesthetics of visual content. Current learning-based methods suffer greatly from the scarcity of labeled data and usually perform sub-optimally in terms of generalization. Although masked image modeling (MIM) has achieved noteworthy advancements across various high-level tasks (e.g., classification, detection etc.). In this work, we take on a novel perspective to investigate its capabilities in terms of quality- and aesthetics-awareness. To this end, we propose Quality- and aesthetics-aware pretraining (QPT V2), the first pretraining framework based on MIM that offers a unified solution to quality and aesthetics assessment. To perceive the high-level semantics and fine-grained details, pretraining data is curated. To comprehensively encompass quality- and aesthetics-related factors, degradation is introduced. To capture multi-scale quality and aesthetic information, model structure is modified. Extensive experimental results on 11 downstream benchmarks clearly show the superior performance of QPT V2 in comparison with current state-of-the-art approaches and other pretraining paradigms. Code and models will be released at https://github.com/KeiChiTse/QPT-V2.
Qizhi Xie, Kun Yuan 0003, Yunpeng Qu, Mingda Wu, Ming Sun 0008, Chao Zhou 0003, Jihong Zhu 0001
ACM Multimedia6
2024 Meet Challenges of RTT Jitter, A Hybrid Internet Congestion Control Algorithm
abstract
Congestion control has been a fundamental research focus in web transmission for over 30 years. However, with diverse network scenarios like cellular networks and WiFi, traditional models might no longer accurately describe current network conditions -- we empirically observe that the minimum round-trip time (RTTmin) still varies under different network conditions, challenging the assumption of its constancy in traditional models. In this paper, we model it as a normal distribution based on our measurements and propose a novel congestion control algorithm LingBo. LingBo consists of two phases: an offline trained decision model to achieve goals under different RTTmin distributions, and an online perception scheme to detect the current RTTmin distribution. We evaluate LingBo in various network environments and find it consistently performs well in terms of power metric and throughput compared to recent state-of-the-art baselines. Our code is available at https://github.com/thumedia/LingBo.
Lianchen Jia, Chao Zhou 0003, Tianchi Huang, Chaoyang Li 0002, Lifeng Sun
WWW2
2023 Buffer Awareness Neural Adaptive Video Streaming for Avoiding Extra Buffer Consumption
Tianchi Huang, Chao Zhou 0003, Rui-Xiao Zhang, Chenglei Wu, Lifeng Sun
INFOCOM2
2023 RDladder: Resolution-Duration Ladder for VBR-encoded Videos via Imitation Learning
Lianchen Jia, Chao Zhou 0003, Tianchi Huang, Chaoyang Li 0002, Lifeng Sun
INFOCOM2
2022 Learned Internet Congestion Control for Short Video Uploading
abstract
Short video uploading service has become increasingly important, as at least 30 million videos are uploaded per day. However, we find that existing congestion control (CC) algorithms, either heuristics or learning-based, are not applicable for video uploading -- i.e., lacking in the design of the fundamental mechanism and being short of leveraging network modeling. We present DuGu, a novel learning-based CC algorithm designed by considering the unique proprieties of video uploading via the probing phase and internet networking via the control phase. During the probing phase, DuGu leverages the transmission gap of uploading short videos to actively detect the network metrics to better understand network dynamics. DuGu uses a neural network~(NN) to avoid congestion during the control phase. Here, instead of using handcrafted reward functions, the NN is learned by imitating the expert policy given by the optimal solver, improving both performance and learning efficiency. To build this system, we construct an omniscient-like network emulator, implement an optimal solver and collect a large corpus of real-world network traces to learn expert strategies. Trace-driven and real-world A/B tests reveal that DuGu supports multi-objective and rivals or outperforms existing CC algorithms across all considered scenarios.
Tianchi Huang, Chao Zhou 0003, Lianchen Jia, Rui-Xiao Zhang, Lifeng Sun
ACM Multimedia2
2022 PDAS: Probability-Driven Adaptive Streaming for Short Video
abstract
To improve Quality of Experience (QoE) for short video applications, most commercial companies adopt preloading and adaptive streaming technologies concurrently. Though preloading can reduce rebuffering, it may greatly waste bandwidth if the downloaded video chunks are not played. Also, each short video's downloading competes against others, which makes the existing adaptive streaming technologies fail to optimize the QoE for all videos. In this paper, we propose PDAS, a Probability-Driven Adaptive Streaming framework, to minimize the bandwidth waste while guaranteeing QoE simultaneously. We formulate PDAS into an optimization problem, where a probabilistic model is designed to describe the swiping events. Then, the maximum preload size is controlled by the proposed probability-driven max-buffer model, which reduces the bandwidth waste by proactively sleeping. At last, the optimization problem is solved by jointly deciding the preload order and preload bitrate. Extensive experimental results demonstrate that PDAS achieves almost 22.34% gains on QoE and 22.80% reductions on bandwidth usage against the existing methods. As for online evaluation, PDAS ranks first in the ACM MM 2022 Grand Challenge: Short Video Streaming.
Chao Zhou 0003, Yixuan Ban, Yangchao Zhao
ACM Multimedia1
2022 Learning Tailored Adaptive Bitrate Algorithms to Heterogeneous Network Conditions: A Domain-Specific Priors and Meta-Reinforcement Learning Approach
abstract
Internet adaptive video streaming is a typical form of video delivery that leverages adaptive bitrate (ABR) algorithms to provide video services with high quality of experience (QoE) for various users in diverse and unique network conditions. Such heterogeneous network environments, which can be viewed as exogenous input processes, often lead to the unstable performance of ABR algorithms. Unfortunately, learning-based ABR algorithm which generated by state-of-the-art reinforcement learning (RL) technologies achievesgood average performancebut fails to perform well in all kinds of network conditions. In this work, considering the video playback process as the Input-driven Markov Decision Process (IMDP), we propose$\text{A}^{2}$BR (Adaptation of ABR), a novel meta-RL ABR approach.$\text{A}^{2}$BR is mainly composed of an online stage and an offline stage. It leverages meta-RL to learn an initial meta-policy with various network conditions at the offline stage and makes decisions in personalized network conditions at the online stage. At the same time, we continually optimize the meta-policy to the tailor-made ABR policy for varying the current network environment within few shots. Moreover, in order to improve the learning efficiency, we fully utilize domain knowledge for implementing a virtual player to replay the previously experienced network. Using trace-driven experiments on various scenarios including different vehicles, users, network types, and heterogeneous user-preferences, we show that$\text{A}^{2}$BR outperforming recent ABR approaches with rapidly adapting to the personalized QoE metrics and specific network conditions. Testbed experimental results also illustrate the superiority of$\text{A}^{2}$BR in adapting to the unseen environments.
Tianchi Huang, Chao Zhou 0003, Rui-Xiao Zhang, Chenglei Wu, Lifeng Sun
IEEE J. Sel. Areas Commun.2
2022 Cratus: A Lightweight and Robust Approach for Mobile Live Streaming
abstract
Live video applications are getting popular, and content providers widely use adaptive bitrate (ABR) streaming to improve QoE while maintaining low latency. However, users’ increasing preference to watch videos on mobile devices poses great challenges for ABR algorithm due to the dramatically varying cellular network. Existing learn-based ABR algorithms face difficulties to generalize to various network conditions because of their reliance on training traces, and model/rule-based ABR schemes suffer from rebuffering under low latency constraint since they cannot robustly control the buffer occupancy within a small range. To address it, this work proposes Cratus, a lightweight and robust ABR algorithm for mobile live streaming, which achieves high QoE and low latency by accurately regulating the buffer at a small level. To enhance the control ability, Cratus controls the buffer dynamic behavior rather than the buffer occupancy. By using sliding mode control approach, Cratus robustly controls the buffer dynamic and ensures that the buffer occupancy is bounded around the target level regardless of network uncertainties. Trace-driven experiments show that Cratus outperforms existing ABRs: average QoE is increased by 12.3 to 28.6 percent, and rebuffering time is limited within 0.8$s$on average, which is reduced by 53.5 to 92.3 percent.
Bo Wang 0066, Mingwei Xu 0001, Fengyuan Ren, Chao Zhou 0003
IEEE Trans. Mob. Comput.4
2021 Deadline and Priority-aware Congestion Control for Delay-sensitive Multimedia Streaming
abstract
Most applications of interactive multimedia require the data to arrive within the specific acceptable end-to-end latency (i.e., meeting deadline). To avoid efforts being wasted, the content must reach the destination before the deadline. In our work, we propose DAP (Deadline And Priority-aware congestion control) to achieve high throughput within acceptable end-to-end latency, especially to send high-priority packets while meeting deadline requirements. DAP is mainly composed of two modules: i) the scheduler decides which packet should be sent at first w.r.t the reward function with fully considering the packets' priority, deadline, and current network conditions. ii) the deadline-sensitive congestion control module transmits packets with high efficiency while guaranteeing the end-to-end latency. Specifically, we propose an improved packet-pair scheme to adjust the best congestion window corresponding to the Bandwidth-Delay Product and to update the instant sending rate by current queue length. Experimental results demonstrate the significant performance of our scheme and DAP ranks first in both the training phase and final phase of the ACM MM 2021 Grand Challenge: Meet Deadline Requirements.
Chao Zhou 0003, Tianchi Huang
ACM Multimedia1
2021 Improving the Performance of Online Bitrate Adaptation with Multi-Step Prediction Over Cellular Networks
abstract
Video streaming over mobile is flourishing, and most commercial players use adaptive bitrate (ABR) streaming to deliver video in varying network conditions. Using network capacity and buffer occupancy as system states, ABR algorithms adjust bitrate based on the instantaneous system states, which is able to adapt to network changes in real-time and ensure high quality of experience (QoE). However, they are incapable of providing good QoE over mobile. Due to the high dynamic characteristics of cellular network, the system states change rapidly over time. The instantaneous state-based adaptation can induce significant video quality fluctuation which greatly degrades QoE. In this paper, we propose an online ABR algorithm called MSPC to provide good QoE in cellular network. To balance the conflict between rapid adaptation and smooth bitrate, MSPC utilizes the multi-step prediction of future system states to select bitrates instead of the instantaneous current states. At the same time, it controls the buffer occupancy to eliminate the impact of prediction error on performance. We implement MSPC on a reference video player with performance evaluated based on realistic cellular traces. Experimental results show that MSPC reduces the bitrate change of existing online algorithms by 62.4 percent on average while maintaining high bitrates and achieving zero rebuffering over 97.83 percent of all tested sessions.
Bo Wang 0066, Fengyuan Ren, Jiahai Yang 0001, Chao Zhou 0003
IEEE Trans. Mob. Comput.4
2021 Practically Deploying Heavyweight Adaptive Bitrate Algorithms With Teacher-Student Learning
abstract
Major commercial client-side video players employ adaptive bitrate (ABR) algorithms to improve the user quality of experience (QoE). With the evolvement of ABR algorithms, increasingly complex methods such as neural networks have been adopted to pursue better performance. However, these complex methods are too heavyweight to be directly deployed in client devices with limited resources, such as mobile phones. Existing solutions suffer from a trade-off between algorithm performance and deployment overhead. To make the deployment of sophisticated ABR algorithms practical, we propose PiTree, a general, high-performance, and scalable framework that can faithfully convert sophisticated ABR algorithms into decision trees with teacher-student learning. In this way, network operators can train complex models offline and deploy converted lightweight decision trees online. We also present theoretical analysis on the conversion and provide two upper bounds of the prediction error during the conversion and the generalization loss after conversion. Evaluation on three representative ABR algorithms with both trace-driven emulation and real-world experiments demonstrates that PiTree could convert ABR algorithms into decision trees with <; 3% average performance degradation. Moreover, compared to original deployment solutions, PiTree could save considerable operating expenses for content providers.
Zili Meng, Yaning Guo, Yixin Shen 0002, Chao Zhou 0003, Minhu Wang, Jia Zhang 0010, Mingwei Xu 0001, Chen Sun 0005, Hongxin Hu
IEEE/ACM Trans. Netw.5
2020 High Efficiency Live Video Streaming With Frame Dropping
abstract
HTTP based video streaming is widely adopted in video service and its adaptive bitrate algorithms have attracted a lot of studies in recent years. However, most of the algorithms are designed for on-demand video service and are not suitable for live video streaming service, which is sensitive to end-to-end latency. In this paper, we propose a live video streaming framework based on HTTP/2 which enables frame dropping for low latency environment. Firstly, we formulate live video streaming model and QoE model considering frame dropping. Then, an optimization problem is formulated aiming at high quality of video streaming. Furthermore, to solve this problem, we propose a live video streaming adaptation algorithm with frame dropping based on Model Predictive Control (MPC). Finally, extensive experiments are conducted to evaluate the proposed method over realistic traces with general Adaptive Bitrate Algorithms (ABR). Compared with the optimal solution, the proposed method achieves comparable performance with only 8.06% quality loss.
Shanshe Wang, Xinfeng Zhang 0001, Chao Zhou 0003, Siwei Ma 0001
ICIP4
2020 Stick: A Harmonious Fusion of Buffer-based and Learning-based Approach for Adaptive Streaming
abstract
Off-the-shelf buffer-based approaches leverage a simple yet effective buffer-bound to control the adaptive bitrate (ABR) streaming system. Nevertheless, such approaches in standard parameters fail to always provide high quality of experience (QoE) video streaming services under all considered network conditions. Meanwhile, state-of-the-art learning-based ABR approach Pensieve outperforms existing schemes but is impractical to deploy. Therefore, how to harmoniously fuse the buffer-based and learning-based approach has become a key challenge for further enhancing ABR methods. In this paper, we propose Stick, an ABR algorithm that fuses the deep learning method and traditional buffer-based method. Stick utilizes the deep reinforcement learning (DRL) method to train the neural network, which outputs the buffer-bound to control the buffer-based approach for maximizing the QoE metric with different parameters. Trace-driven emulation illustrates that Stick betters Pensieve by 3.5% - 9.41% with an overhead reduction of 88%. Moreover, aiming to further reduce the computational costs while preserving the performances, we propose Trigger, a light-weighted neural network that determines whether the buffer-bound should be adjusted. Experimental results show that Stick+Trigger rivals or outperforms existing schemes in average QoE by 1.7%-28%, and significantly reduces the Stick's computational overhead by 24%-61%. Meanwhile, we show that Trigger also helps other ABR schemes mitigate the overhead. Extensive results on real-world evaluation demonstrate the superiority of Stick over existing state-of-the-art approaches.
Tianchi Huang, Chao Zhou 0003, Rui-Xiao Zhang, Chenglei Wu, Xin Yao 0003, Lifeng Sun
INFOCOM2
2020 Quality-Aware Neural Adaptive Video Streaming With Lifelong Imitation Learning
abstract
Existing Adaptive Bitrate (ABR) algorithms pick future video chunks' bitrates via fixed rules or offline trained models to ensure good quality of experience (QoE) for Internet video. Nevertheless, data analysis demonstrates that a good ABR algorithm is required to continually and fast update for adapting itself to time-varying network conditions. Therefore, we propose Comyco, a video quality-aware learning-based ABR approach that enormously improves recent schemes by i) picking the chunk with higher perceptual video qualities rather than video bitrates; ii) training the policy via imitating expert trajectories given by the expert strategy; iii) employing the lifelong learning method to continually train the model w.r.t the fresh trace collected by the users. To achieve this, we develop a complete quality-aware lifelong imitation learning-based ABR system, construct quality-based neural network architecture, collect a quality-driven video dataset, and estimate QoE metrics with video quality features. Using trace-driven and real-world experiments, we demonstrate Comyco reaches 1700-fold improvements in the number of samples required and 16-fold speedup in the training time compared with the prior work. Meanwhile, Comyco outperforms existing methods, with the improvements on average QoE of 7.5%-16.79%. Moreover, experimental results on continual training also illustrate that lifelong learning helps Comyco further improve the average QoE of 1.07%-9.81% in comparison to the offline trained model.
Tianchi Huang, Chao Zhou 0003, Xin Yao 0003, Rui-Xiao Zhang, Chenglei Wu, Lifeng Sun
IEEE J. Sel. Areas Commun.2
2019 Hybrid Control-Based ABR: Towards Low-Delay Live Streaming
abstract
Video content providers are increasingly interested in interactive live streaming since user engagement increases their revenues. To provide high quality of experience (QoE), it is critical to design a low delay adaptive bitrate (ABR) algorithm, but which is lacked in existing studies. The low delay constraint poses much more challenges to achieve high bitrate and low rebuffering. For example, low delay requires the player to maintain a small playback buffer, which, however, increases the risk of rebuffering. This work designs a low delay ABR algorithm called HCA which provides good QoE by controlling the buffer occupancy at a low but non-empty level. To achieve accurate control, HCA uses the hybrid of feedback and feedforward control to regulate the buffer dynamic (buffer occupancy and its variation) based on predictions of future network condition (throughput and its variation). Trace-driven experiments show that HCA achieves zero rebuffering for 98% of all traces while ensuring high bitrate.
Bo Wang 0066, Fengyuan Ren, Chao Zhou 0003
ICME3
2019 Comyco: Quality-Aware Adaptive Video Streaming via Imitation Learning
abstract
Learning-based Adaptive Bit Rate~(ABR) method, aiming to learn outstanding strategies without any presumptions, has become one of the research hotspots for adaptive streaming. However, it is still suffering from several issues, i.e., low sample efficiency and lack of awareness of the video quality information. In this paper, we propose Comyco, a video quality-aware ABR approach that enormously improves the learning-based methods by tackling the above issues. Comyco trains the policy via imitating expert trajectories given by the instant solver, which can not only avoid redundant exploration but also make better use of the collected samples. Meanwhile, Comyco attempts to pick the chunk with higher perceptual video qualities rather than video bitrates. To achieve this, we construct Comyco's neural network architecture, video datasets and QoE metrics with video quality features. Using trace-driven and real world experiments, we demonstrate significant improvements of Comyco's sample efficiency in comparison to prior work, with 1700x improvements in terms of the number of samples required and 16x improvements on training time required. Moreover, results illustrate that Comyco outperforms previously proposed methods, with the improvements on average QoE of 7.5% - 16.79%. Especially, Comyco also surpasses state-of-the-art approach Pensieve by 7.37% on average video quality under the same rebuffering time.
Tianchi Huang, Chao Zhou 0003, Rui-Xiao Zhang, Chenglei Wu, Xin Yao 0003, Lifeng Sun
ACM Multimedia2
2019 Generalizing Rate Control Strategies for Realtime Video Streaming via Learning from Deep Learning
abstract
The leading learning-based rate control method, i.e., QARC, achieves state-of-the-art performances but fails to interpret the fundamental principles, and thus lacks the abilities to further improve itself efficiently. In this paper, we propose EQARC (Explainable QARC) via reconstructing QARC's modules, aiming to demystify how QARC works. In details, we first utilize a novel hybrid attention-based CNN+GRU model to re-characterize the original quality prediction network and reasonably replace the QARC's 1D-CNN layers with 2D-CNN layers. Using trace-driven experiment, we demonstrate the superiority of EQARC over existing state-of-the-art approaches. Next, we collect several useful information from each interpretable modules and learn the insight of EQARC. Following this step, we further propose AQARC (Advanced QARC), which is the light-weighted version of QARC. Experimental results show that AQARC achieves the same performances as the QARC with an overhead reduction of 90%. In short, through learning from deep learning, we generalize a rate control method which can both reach high performance and reduce computation cost.
Tianchi Huang, Rui-Xiao Zhang, Chenglei Wu, Xin Yao 0003, Chao Zhou 0003, Lifeng Sun
MMAsia5
2019 TFDASH: A Fairness, Stability, and Efficiency Aware Rate Control Approach for Multiple Clients Over DASH
abstract
Dynamic adaptive streaming over HTTP (DASH) has recently been widely deployed in the Internet and adopted in the industry. It, however, does not impose any adaptation logic for selecting the quality of video segments requested by clients and suffers from lackluster performance with respect to a number of desirable properties: efficiency, stability, and fairness when multiple players compete for a bottleneck link. In this paper, we propose a throughput-friendly DASH rate control scheme for video streaming with multiple clients over DASH to well balance the tradeoffs among efficiency, stability, and fairness. The core idea behind guaranteeing fairness and high efficiency (bandwidth utilization) is to avoid OFF periods during the downloading process for all clients, i.e., the bandwidth is in perfect-subscription or over-subscription with bandwidth utilization approach to 100%. We also propose a dual-threshold buffer model to solve the instability problem caused by the above idea. As a result, by integrating these novel components, we also propose a probability-driven rate adaption logic taking into account several key factors that most influence visual quality, including buffer occupancy, video playback quality, video bit-rate switching frequency and amplitude, to guarantee high-quality video streaming. Our experiments evidently demonstrate the superior performance of the proposed method.
Chao Zhou 0003, Chia-Wen Lin, Xinggong Zhang, Zongming Guo
IEEE Trans. Circuits Syst. Video Technol.1
2018 QARC: Video Quality Aware Rate Control for Real-Time Video Streaming based on Deep Reinforcement Learning
abstract
Real-time video streaming is now one of the main applications in all network environments. Due to the fluctuation of throughput under various network conditions, how to choose a proper bitrate adaptively has become an upcoming and interesting issue. To tackle this problem, most proposed rate control methods work for providing high video bitrates instead of video qualities. Nevertheless, we notice that there exists a trade-off between sending bitrate and video quality, which motivates us to focus on how to reach a balance between them.
Tianchi Huang, Rui-Xiao Zhang, Chao Zhou 0003, Lifeng Sun
ACM Multimedia3
2018 Delay-Constrained Rate Control for Real-Time Video Streaming with Bounded Neural Network
abstract
Rate control is widely adopted during video streaming to provide both high video qualities and low latency under various network conditions. However, despite that many work have been proposed, they fail to tackle one major problem: previous methods determine a future transmission rate as a single for value which will be used in an entire time-slot, while real-world network conditions, unlike lab setup, often suffer from rapid and stochastic changes, resulting in the failures of predictions.
Tianchi Huang, Rui-Xiao Zhang, Chao Zhou 0003, Lifeng Sun
NOSSDAV3
2017 Dynamic threshold based rate adaptation for HTTP live streaming
abstract
The Dynamic Adaptive Streaming over HTTP (DASH) is specified to cope with the changing network conditions and provide an adaptive bit-rate HTTP-based streaming solution. While there have been many researches of rate adaptation algorithms on adaptive HTTP streaming, much of the work is focused on Video on Demand (VoD) service - which is not same as live streaming. It is generally preferred to minimize the end-to-end delay and make full use of the bandwidth for live services. In this paper, we propose a buffer-based rate adaptation algorithm with dynamic threshold which can decrease the rate transitions and provide a seamless playback under a low latency requirement. The rate adaptation metrics not only take into account the momentary value of bandwidth but also consider its fluctuation as the recognition of bandwidth is crucial over small buffer. Experiments demonstrate that our proposed rate adaptation scheme outperforms the methods using fixed threshold or instant throughput.
Lan Xie, Chao Zhou 0003, Xinggong Zhang, Zongming Guo
ISCAS2
2016 An unequal error protection scheme for reliable peer-to-peer scalable video streaming
Chi-Wen Lo, Chao Zhou 0003, Chia-Wen Lin, Yung-Chang Chen
J. Vis. Commun. Image Represent.2
2016 mDASH: A Markov Decision-Based Rate Adaptation Approach for Dynamic HTTP Streaming
abstract
Dynamic adaptive streaming over HTTP (DASH) has recently been widely deployed in the Internet. It, however, does not impose any adaptation logic for selecting the quality of video fragments requested by clients. In this paper, we propose a novel Markov decision-based rate adaptation scheme for DASH aiming to maximize the quality of user experience under time-varying channel conditions. To this end, our proposed method takes into account those key factors that make a critical impact on visual quality, including video playback quality, video rate switching frequency and amplitude, buffer overflow/underflow, and buffer occupancy. Besides, to reduce computational complexity, we propose a low-complexity sub-optimal greedy algorithm which is suitable for real-time video streaming. Our experiments in network test-bed and real-world Internet all demonstrate the good performance of the proposed method in both objective and subjective visual quality.
Chao Zhou 0003, Chia-Wen Lin, Zongming Guo
IEEE Trans. Multim.1
2015 Unequal error protection for real-time video streaming using expanding window reed-solomon code
abstract
Expanding Window FEC is an emerging scheme for robust real-time video streaming over wireless networks, with low latency and reduction of error propagation. In this work, we focus on the problem of Expanding Window FEC redundancy allocation which has not been adequately addressed in current works. We first analyse the error probability of the adopted Expanding Window Reed-Solomon code (EW-RS), and introduce an equivalent error probability to simplify them. Then we are able to formulate the optimal redundancy allocation into a constrained nonlinear optimization problem, where by allocating the redundancy unequally considering the unequal importance of different frames and their dependency based on the expanding window, unequal error protection (UEP) is achieved and the overall distortion is minimized. Moreover, to reduce the computation complexity, a high-efficiency hill-climbing algorithm is developed to obtain the suboptimal allocation. At last, the experimental results demonstrate the effectiveness of both the proposed allocation scheme and solution algorithm.
Yufeng Geng, Xinggong Zhang, Chao Zhou 0003, Zongming Guo
ICIP3
2015 A fairness-aware smooth rate adaptation approach for dynamic HTTP streaming
abstract
Recently, Dynamic Adaptive Streaming over HTTP (DASH) has been widely deployed over the Internet. Under time-varying network conditions, it is, however, still a big challenge to provide smooth video bit-rate with high video quality, especially when multiple clients compete for the network resources where the fairness must be considered. In this paper, a fairness-aware smooth rate adaptation approach is designed for DASH under the scenario that multiple clients are competing for the network resources. To avoid the unfair bandwidth estimated by the client induced by the off-intervals during the downloading process, a probe-based bandwidth estimation method is designed which includes a logarithmic law based increase probing scheme and a conservative back-off based decrease probing scheme. Then, with the probed bandwidth, a dual-threshold based video bit-rate switching scheme is designed that buffer overflow/underflow is avoided, and smooth video bit-rate is also provided. The extensive experiments on our network testbed demonstrate that the proposed approach outperforms the existing schemes significantly.
Chao Zhou 0003, Xinggong Zhang, Zongming Guo
ICIP2
2015 Delay-constrained rate control for real-time video streaming over wireless networks
abstract
Rate control is a big challenge for real-time video streaming on the internet with the needs of low latency, bandwidth-consuming and stable video rate. However, most of the existing Internet congestion control protocols ignore these needs, and some of them use the packet loss event as congestion signal which is deviation especially over error-prone wireless networks. In this paper, we propose a delay-constrained rate control algorithm by locking queueing delay onto a desired objective. The shadow price of video rate is controlled by queueing delay. All flows adapt video rate according to distortion weight and shadow price so as to achieve a distributed bandwidth sharing with low latency, efficient utilization, and distortion fairness. A closed-loop rate control system is designed for the purpose of stable and agile control. The control parameters are analyzed using control-theoretic approach. Additionally, we construct a real-time wireless video streaming test-bed and conduct extensive experiments over it. Compared with the current widely used methods, the experimental results show that the proposed algorithm can achieve 3dB or more gains in PSNR, and better performance on bandwidth utilization, flow stability with well guaranteed multi-flow fairness.
Yufeng Geng, Xinggong Zhang, Chao Zhou 0003, Zongming Guo
VCIP4
2015 A Markov decision based rate adaption approach for dynamic HTTP streaming
abstract
In this paper, we propose a novel Markov decision-based rate adaption scheme for DASH aiming to maximize the quality of user experience. To this end, our proposed method takes into account those key factors that have critical impact on visual quality, including video playback quality, video rate switching frequency and amplitude, buffer overflow/underflow, and buffer occupancy. And a dynamic reward function is carefully designed under three scenarios of buffer occupancy to measure the effectiveness of each transfer decision. Besides, to reduce computational complexity, we propose a low-complexity greedy algorithm to make it suitable for real-time video streaming. Our experiments in the real-world Internet demonstrate the good performance of the proposed method in terms of both objective and subjective visual quality.
Chao Zhou 0003, Chia-Wen Lin
VCIP1
2015 A Novel JSCC Scheme for UEP-Based Scalable Video Transmission Over MIMO Systems
abstract
In this paper, we propose a novel joint source-channel coding (JSCC) scheme for scalable video transmission over multiple-input multiple-output (MIMO) systems. By exploiting the diversity of MIMO antennas and forward error correction (FEC)-based protection, our method aims to provide unequal error protection (UEP) for the video layers, which are mapped to appropriate antennas. Moreover, JSCC is also considered that we extract a proper subset of video layers and allocate suitable FEC redundancy to them. Jointly considering video layer extraction, FEC rate allocation, and video layer scheduling, we are able to achieve UEP so as to minimize end-to-end distortion. We formulate the scheme as a nonlinear integer optimization problem, which is known to be NP-hard. To find a near-optimal solution efficiently, we propose a low-complexity branch-and-bound algorithm, which partitions the original problem into a series of subproblems by a video layer branching technique. In each branch, the upper and lower distortion bounds are derived. In particular, we transform the video layer scheduling subproblem into a 0/1 multiple knapsack problem, which is NP-complete, and employ an evolutionary Lagrangian method to find a solution efficiently. For the FEC allocation subproblem, a Lagrange duality algorithm with fuzzy surrogate subgradient is proposed. The experimental results demonstrate that the proposed method has good efficiency while achieving close performance to the optimal results.
Chao Zhou 0003, Chia-Wen Lin, Xinggong Zhang, Zongming Guo
IEEE Trans. Circuits Syst. Video Technol.1
2014 Joint multi-CDN and LT-coding for video transport over HTTP
abstract
Video transport over HTTP is becoming more and more popular. Many video service providers construct huge content distribution networks(CDN) to support HTTP streaming service, however, they seldom exploit the benefits of multiple servers to achieve higher bandwidth and reliability by parallel streaming. In this paper, we study the problem of jointing multi-CDN and LT-coding for video transport over HTTP. Using LT coding, a client could download the same video segment from multiple servers without considering data segmentation and server scheduling issue. Thus, we are able to treat all CDN servers as a virtual server with higher bandwidth and reliability. To reduce the ACK overhead, a stochastic model is designed to predict the amount of data to be sent from each server while guaranteeing the decoding probability. Compared with the existing schemes, the experimental results show that our proposed scheme obtains less overhead and fewer number of HTTP requests. Besides, we also achieve better video quality and better robustness to fluctuant bandwidth.
Chao Zhou 0003, Xinggong Zhang, Zongming Guo
ISCAS2
2014 Probabilistic chunk scheduling approach in parallel multiple-server DASH
abstract
Recently parallel Dynamic Adaptive Streaming over HTTP (DASH) has emerged as a promising way to supply higher bandwidth, connection diversity and reliability. However, it is still a big challenge to download chunks sequentially in parallel DASH due to heterogeneous and time-varying bandwidth of multiple servers. In this paper, we propose a novel probabilistic chunk scheduling approach considering time-varying bandwidth. Video chunks are scheduled to the servers which consume the least time while with the highest probability to complete downloading before the deadline. The proposed approach is formulated as a constrained optimization problem with the objective to minimize the total downloading time. Using the probabilistic model of time-varying bandwidth, we first estimate the probability of successful downloading chunks before the playback deadline. Then we estimate the download time of chunks. A near-optimal solution algorithm is designed which schedules chunks to the servers with minimal downloading time while the completion probability is under the constraint. Compared with the existing schemes, the experimental results demonstrate that our proposed scheme greatly increases the number of chunks that are received orderly.
Chao Zhou 0003, Xinggong Zhang, Zongming Guo
VCIP2
2014 A Control-Theoretic Approach to Rate Adaption for DASH Over Multiple Content Distribution Servers
abstract
Recently, dynamic adaptive streaming over HTTP (DASH) has been widely deployed on the Internet. However, the research about DASH over multiple content distribution servers (MCDS-DASH) is limited. Compared with traditional single-server DASH, MCDS-DASH is able to offer expanded bandwidth, link diversity, and reliability. It is, however, a challenging problem to smooth video bitrate switching over multiple servers due to their diverse bandwidths. In this paper, we propose a block-based rate adaptation method considering both the diverse bandwidths and feedback buffered video time. In our method, multiple fragments are grouped into a block and the fragments are downloaded in parallel from multiple servers. We propose to adapt video bitrate at the block level rather than at the fragment level. By dynamically adjusting the block length and scheduling fragment requests to multiple servers, the requested video bitrates from the multiple servers are synchronized, making the fragments download in an orderly way. Then, we propose a control-theoretic approach to select an appropriate bitrate for each block. By modeling and linearizing the rate adaption system, we propose a novel proportional-derivative controller to adapt video bitrate with high responsiveness and stability. Theoretical analysis and extensive experiments on our network testbed and the Internet demonstrate the good efficiency of the proposed method.
Chao Zhou 0003, Chia-Wen Lin, Xinggong Zhang, Zongming Guo
IEEE Trans. Circuits Syst. Video Technol.1
2013 Adaptive channel scheduling for Scalable Video broadcasting over MIMO wireless networks
abstract
Video broadcasting is an efficient way to deliver video content to multiple receivers. However, due to heterogeneous channel conditions of users, it is challenging to minimize all users' video transmission distortion in MIMO broadcasting. In this paper, we investigate the channel scheduling problem, which maps Scalable Video layers to MIMO heterogenous channels to protect video layers unequally, so as to minimize overall received video distortion for all users. We formulate this problem into an integer non-linear optimization problem, which is hard to be solved. An efficient near-optimal algorithm is proposed, which is based on simulated-annealing theory. Experimental results demonstrate the efficiency of our proposed algorithm, and its performance is very close to the optimal results. Compared with the existing MIMO scheduling methods, the proposed scheme significantly improves the overall quality of video broadcasting in MIMO networks.
Chao Zhou 0003, Xinggong Zhang, Zongming Guo
ISCAS1
2013 A control theory based rate adaption scheme for dash over multiple servers
abstract
Recently, Dynamic Adaptive Streaming over HTTP (DASH) has been widely deployed in the Internet. However, the research about DASH over Multiple Content Distribution Servers (MCDS) is few. Compared with traditional single-server-DASH, MCDS are able to offer expanded bandwidth, link diversity, and reliability. It is, however, a challenging problem to smooth video bitrate switchings over multiple servers due to their diverse bandwidths. In this paper, we propose a block-based rate adaptation method considering both the diverse bandwidths and feedback buffered video time. Multiple fragments are grouped into a block, and the fragments are downloaded in parallel from multiple servers. We propose to adapt video bitrate at the block level rather than at the fragment level. By dynamically adjusting the block length and scheduling fragment requests to multiple servers, the requested video bitrates from the multiple servers are synchronized, making the fragments downloaded orderly. Then, we propose a control-theoretic approach to select an appropriate bitrate for each block. By modeling and linearizing the rate adaption system, we propose a novel Proportional-Derivative (PD) controller to adapt video bitrate with high responsiveness and stability. Theoretical analysis and extensive experiments on the Internet demonstrate the good efficiency of our DASH designs.
Chao Zhou 0003, Xinggong Zhang, Zongming Guo
VCIP1
2013 Optimal adaptive channel scheduling for scalable video broadcasting over MIMO wireless networks
abstract
Video broadcasting is an efficient way to deliver video content to multiple receivers. However, due to heterogeneous channel conditions in MIMO wireless networks, it is challenging for video broadcasting to map scalable video layers to proper MIMO transmit antennas to minimize the average overall video transmission distortion. In this paper, we investigate the channel scheduling problem for broadcasting scalable video content over MIMO wireless networks. An adaptive channel scheduling based unequal error protection (UEP) video broadcasting scheme is proposed. In the scheme, video layers are protected unequally by being mapped to appropriate antennas, and the average overall distortion of all receivers is minimized. We formulate this scheme into a non-linear combinatorial optimization problem. It is not practical to solve the problem by an exhaustive search method with heavy computational complexity. Instead, an efficient branch-and-bound based channel scheduling algorithm, named TBCS, is developed. TBCS finds the global optimal solution with much lower complexity. The complexity is further reduced by relaxing the termination condition of TBCS, which produces a (1 − ε)-optimal solution. Experimental results demonstrate both the effectiveness and efficiency of our proposed scheme and algorithm. As compared with some existing channel scheduling methods, TBCS improves the quality of video broadcasting across all receivers significantly.
Chao Zhou 0003, Xinggong Zhang, Zongming Guo
Comput. Networks1
2012 A novel JSCC scheme for scalable video transmission over MIMO systems
abstract
MIMO recently emerges as one of promising techniques for wireless video streaming. It is still a challenge to provide un-equal error protections by joint source-channel coding (JSCC) over multiple diverse MIMO sub-channels. In this paper, a joint source-channel coding and antenna mapping scheme for scalable video transmission over MIMO systems is proposed. Bandwidth are elaborately allocated between video source and channel protections by layer extracting and FEC coding. For the extracted layers, we determine i) which antenna will they be transmitted over and ii) how much redundancy bits will be added for error protections. We formulate this scheme into a non-linear integer optimization problem, whose complexity is very high. Instead, a low-complexity branch-and-bound algorithm is presented. Source layers are partitioned into subsets of layers, and the selected layer are mapped to antennas using Min-max scheduling algorithm. By branching and pruning, the computation complexity are reduced significantly. We carry out extensive numerical experiments under various network conditions. The results demonstrate our algorithm's efficiency and the overall transmission quality is improved significantly.
Xinggong Zhang, Chao Zhou 0003, Zongming Guo
ICIP2
2012 Cross-entropy based antenna selection for scalable video streaming over MIMO wireless networks
abstract
In this paper, we investigate the antenna selection (AS) problem for scalable video streaming over MIMO wireless networks. By scheduling scalable video layers over MIMO antennas with different signal strength, the video layers are transmitted with un-equal error protections. Considering layer dependencies and various antenna conditions, it is a non-linear combinatorial problem for AS to minimize the overall end-to-end distortion. To find the optimal solution with low complexity, a cross-entropy based solution, named CEBAS, is proposed. All solutions are indexed by unique binary strings, and the primal problem is reformulated to a binary combination problem. Then, random strings are generated using the probability distribution of solutions, which is updated by the cross-entropy optimization method. The feasibility of solution is guaranteed by our proposed projection strategy. CEBAS is iterative in nature and converges to the global optimum in probability. Simulation results reveal both the effectiveness and efficiency of our proposed algorithm. When comparing CEBAS against other existing algorithms, consistent superior performance has been observed.
Chao Zhou 0003, Xinggong Zhang, Zongming Guo
ICIP1
2012 A control-theoretic approach to rate adaptation for dynamic HTTP streaming
abstract
Recently, dynamic adaptive HTTP streaming has been widely used for video content delivery over Internet. However, it is still a challenge how to switch video bitrate under time-varying bandwidth. In this paper, we propose a novel control-theoretic approach to adapt video segments in dynamic HTTP streaming. The rate control is based on a sink-buffer, which has an overflow-threshold and an underflow-threshold. The objective is to maximize the playback quality while keeping the receiver buffer from either overflow or underflow. Using control theory, we formulate this rate control scheme as a proportional (P) control system, which exists oscillations and steady-errors. Furthermore, we design a proportional derivative (PD) controller to improve its adaptation performance. The conditions for stability and settling time of the PD controller are also derived. Numerous experiment results demonstrate the effectiveness of our proposed PD control scheme for dynamic HTTP streaming.
Chao Zhou 0003, Xinggong Zhang, Longshe Huo, Zongming Guo
VCIP1
2010 Collision-detection based rate-adaptation for video multicasting over IEEE 802.11 wireless networks
abstract
Wireless video multicasting/broadcasting is an efficient method for simultaneous transmission of data to a group of users. But the multicasting rates are fixed in current IEEE 802.11 PHYs standard. In this paper, we propose a novel collision-detection based rate-adaptation scheme (CDRA), which fully exploits the potential of rate adaptation capability of wireless physical layer, to improve service qualities of video multicasting. The received signal strength indication (RSSI) and packet error ratio (PER) are comprehensively used to detect collision. The PER-guided rate adjustment algorithm is performed when no collision happens. Otherwise the collision-avoid mechanism works. By detecting the collision, our scheme could adaptively select the maximum data rates for video multicasting. We construct a practical multicasting test-bed in IEEE 802.11b network and carry out extensive experiments. The results show that CDRA achieves throughput gain up to 166% and PSNR gain to 139% compared with existing methods.
Chao Zhou 0003, Xinggong Zhang, Lichuan Lu, Zongming Guo
ICIP1