Jun Sun 0012

dblp:s/JunSun12 · DBLP profile ↗
← Back
57ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0001-6864-0803ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 49 · 8 first-author · 11 since 2021Databases, data management, data science and information retrieval · 6 · 4 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Systems, architecture and hardware · 3Computer networks · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Cross-View Interaction and Collaboration for Semi-supervised Time Series Classification
Maojing Shu, Xiaopeng Guo 0001, Jun Sun 0012
KSEM (1)3
2026 Enhancing Multivariate Time-Series Class Incremental Learning with Informative Spatiotemporal Consolidation
Maojing Shu, Xiaopeng Guo 0001, Jun Sun 0012
KSEM (1)3
2026 Efficient VVC Intra Partitioning With Shared Feature Extraction and Complexity-Aware Threshold Decision
abstract
Versatile Video Coding (VVC) achieves significant compression efficiency by adopting various techniques, especially the Quad-Tree Plus Multi-Type Tree (QTMT) partition structure, but this leads to a dramatic rise in encoding complexity. Some accelerating methods have relatively low accuracy in predicting partition modes, while others increase accuracy at the cost of larger network inference time. When setting thresholds for skipping partition modes, it is common to determine them empirically or based on prediction accuracy, often neglecting their impact on encoding complexity. To achieve a balance in accuracy and efficiency, we propose to only extract the feature map of each 32×32 CU once and share it among sub-CUs. Padding between layers is removed to keep consistency between training and inference. Considering that encoding complexity of each partition mode differs, an algorithm is designed to search for the optimal thresholds with a balance between RD performance and encoding complexity. Then, we integrate a gradient-based method and a novel partition mode reuse method into VTM, which achieve a better RD-complexity trade-off at a large time-saving rate. Lastly, we remove several existing accelerating methods in VTM and find that our method is a good substitute. Experiments implemented single-threaded on CPU demonstrate that our method can reduce 25.29%$\sim$73.98% complexity of VVC intra encoding with -0.04%$\sim$1.93% BDRate increase. All the code is available athttps://github.com/cppppp/FastIntra.
Jun Sun 0012
IEEE Trans. Multim.3
2025 STR-Saliency: Decomposition-based Perturbations to Generate Saliency Maps for Temporal Black-box Model Interpretation
abstract
The interpretability of deep black-box temporal models is crucial in modern machine learning. Identifying crucial time steps and temporal patterns is an important way in understanding how a black-box model makes a decision on a time series instance. Saliency methods are widely used for the interpretation of deep models. However, its application on temporal models faces two challenges: handling temporal relations and the selection of perturbation functions. In order to overcome these challenges, we propose Seasonal-Trend-Remainder Saliency (STR-Saliency), a new interpretation framework which decomposes a time series into three components then generates saliency maps on each component using a learning-based perturbation process. Our new method not only addresses the two issues above, but also produces more human-understandable saliency maps than previous methods. We evaluate our method on both synthetic and real-world datasets and it outperforms various baselines. Our code and detailed results are available at https://github.com/chenrunkai/STRSaliency.
Runkai Chen, Yuedi Chen, Haozhe Xu, Xiaopeng Guo 0001, Jun Sun 0012
ICASSP5
2025 Efficient Modeling and Low Complexity Implementation of Rate Estimation in Versatile Video Coding
abstract
In Versatile Video Coding (VVC), Rate-Distortion Optimized Quantization (RDOQ) is a widely adopted technique to strike a balance between bit rate and distortion. However, the computational complexity introduced by RDOQ poses significant challenges for real-time applications. To address this issue, we propose a low-complexity rate estimation model to accelerate the conventional RDOQ approach. Through comprehensive analysis, we have identified the generalized Gaussian distribution (GGD) as an efficient model for simulating the distribution of transform coefficients. Leveraging this insight, we developed a GGD-based rate estimator to expedite the selection of quantization levels. Experimental results demonstrate that the proposed model achieves an average encoding time reduction of 13.15%, with only a minimal impact on BD-Rate.
Jun Sun 0012
ICASSP4
2024 Programming Knowledge Tracing with Context and Structure Integration
Xiaopeng Guo 0001, Maojing Shu, Jun Sun 0012
KSEM (1)5
2024 STRANet: Soft-Target and Restriction-Aware Neural Network for Efficient VVC Intra Coding
abstract
The Versatile Video Coding (VVC) standard introduces a quad-tree with a nested multi-type tree (QTMT) partition structure to improve the rate-distortion (RD) performance, but this leads to a substantial increase in encoding complexity. Previous studies have labeled partition modes of CUs using hard targets (i.e., one-hot labels) generated by VVC reference software (VTM), which is challenging for neural networks to predict accurately. Furthermore, in earlier works, the VVC restrictions are not incorporated into convolutional neural network (CNN), not fully exploiting the predicting capacity of CNN. In this paper, we propose a novel soft-target and restriction-aware neural network (STRANet) to address these issues. Firstly, inspired by the observation that a CU may split differently under various circumstances, we collect these RD costs and precisely estimate the probability of each partition mode to generate a soft target. Secondly, our neural network incorporates QP and restriction type through attention modules so as to output predictions that are standard-compliant with simple post-processing. Thirdly, Window Attention Module, a combination of CNN and attention mechanism, is adopted to further enhance performance on GPU. Through the application of these methods, STRANet reduces encoding time by 51.84% and 61.00% with 0.44% and 0.84% Bjøntegaard delta bit-rate (BD-BR) increase, superior to state-of-the-art methods. The code has been released athttps://github.com/cppppp/STRANet.
Jun Sun 0012
IEEE Trans. Circuits Syst. Video Technol.4
2023 Generalized Compressed Video Restoration by Multi-Scale Temporal Fusion and Hierarchical Quality Score Estimation
abstract
Learning-based methods have achieved excellent performance for compressed video restoration (CVR) in recent years. However, existing networks aggregate multi-frame information inefficiently and are usually developed for specific quantization parameters (QPs), which are not convenient for practical usage. Moreover, current works only consider compressed video restoration in Constant QP (CQP) setting, but do not discuss the performance of the model in more realistic scenarios, e.g., Constant Rate Factor (CRF) and Constant Bitrate (CBR). In this paper, we propose a generalized quality-aware compressed video restoration network, namely QCRN. Specifically, to achieve multi-frame aggregation efficiently, we propose a multi-scale deformable temporal fusion. Meanwhile, QCRN decouples the global quality and local quality representations from input via the hierarchical quality score estimator, and then employs them to adjust the feature enhancement. Extensive experiments on compressed videos in various settings demonstrate that our proposed QCRN achieves favorable performance against state-of-the-art methods in terms of both quantitative metrics and visual quality.
Xiaopeng Guo 0001, Jun Sun 0012
ICME5
2023 Efficient Intra Coding through Hierarchical CU Partition Prediction for VVC
abstract
Versatile Video Coding (VVC) has adopted a quad-tree with a nested multi-type tree (QTMT) partition structure to improve the rate-distortion (RD) performance, but this greatly increases complexity due to the brute-force RDO process to search for the best partition type. Some methods cannot fully utilize partition information of deeper CUs when training, while others process large CUs as a whole and may neglect specific information inside smaller CUs. Therefore, in this paper, we propose a learning-based approach to predict the best partition modes for every CU while utilizing the partition modes of deeper CUs. Firstly, we model the searching for the optimal partition split in a hierarchical prediction way. Secondly, we predict a confidence score along with each prediction to select the accurate number of partition types. Lastly, we incorporate the prediction above into a framework with separate predictions for NS/QT/MTT split and MTT split, and modify it according to the mechanism of VTM. Experiments demonstrate that our method can reduce 47.21%∼53.97% complexity of VVC intra coding with negligible 0.81%∼1.26% BD-BR increase, superior to other state-of-the-art methods.
Jun Sun 0012
VCIP4
2023 FastCNN: Towards Fast and Accurate Spatiotemporal Network for HEVC Compressed Video Enhancement
abstract
Deep neural networks have achieved remarkable success in HEVC compressed video quality enhancement. However, most existing multiframe-based methods either deliver unsatisfactory results or consume a significant amount of resources to leverage temporal information of neighboring frames. For the sake of practicality, a thorough investigation of the architecture design of the video quality enhancement network regarding enhancement performance, model parameters, and running speed is essential. In this article, we first propose an efficient alignment module that can quickly and accurately aggregate the spatiotemporal information of neighboring frames. The proposed module estimates deformable offsets progressively in lower-resolution space motivated by the observation of offset correlations between adjacent pixels. Then, the quantization parameter (QP) that represents compression level prior knowledge is utilized to guide aligned feature enhancement. By combining alignment feature distillation with residual feature correction, we obtain an efficient QP attention block. To save the storage space of the network, we design a hash buffer to store QP embedding features. These efficient components allow our network to effectively exploit temporal redundancies and obtain favorable enhancement capability while maintaining a lightweight structure and fast running speed. Extensive experiments demonstrate that the proposed approach outperforms state-of-the-art methods over different QPs by up to 0.09 to 0.11 dB, whereas the inference time can be reduced by up to 69%.
Jun Sun 0012, Xiaopeng Guo 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2022 One-for-All: An Efficient Variable Convolution Neural Network for In-Loop Filter of VVC
abstract
Recently, many researches on convolution neural network (CNN) based in-loop filters have been proposed to improve coding efficiency. However, most existing CNN based filters tend to train and deploy multiple networks for various quantization parameters (QP) and frame types (FT), which drastically increases resources in training these models and the memory burdens for video codec. In this paper, we propose a novel variable CNN (VCNN) based in-loop filter for VVC, which can effectively handle the compressed videos with different QPs and FTs via a single model. Specifically, an efficient and flexible attention module is developed to recalibrate features according to QPs or FTs. Then we embed the module into the residual block so that these informative features can be continuously utilized in the residual learning process. To minimize the information loss in the learning process of the entire network, we utilize a residual feature aggregation module (RFA) for more efficient feature extraction. Based on it, an efficient network architecture VCNN is designed that can not only effectively reduce compression artifacts, but also can be adaptive to various QPs and FTs. To address training data imbalance on various QPs and FTs and improve the robustness of the model, a focal mean square error loss function is employed to train the proposed network. Then we integrate the VCNN into VVC as an additional tool of in-loop filters after the deblocking filter. Extensive experimental results show that our VCNN approach obtains on average 3.63%, 4.36%, 4.23%, 3.56% under all intra, low-delay P, low-delay, and random access configurations, respectively, which is even better than QP-Separate models.
Jun Sun 0012, Xiaopeng Guo 0001, Mingyu Shang
IEEE Trans. Circuits Syst. Video Technol.2
2021 An Efficient QP Variable Convolutional Neural Network Based In-loop Filter for Intra Coding
abstract
In this paper, a novel QP variable convolutional neural network based in-loop filter is proposed for VVC intra coding. To avoid training and deploying multiple networks, we develop an efficient QP attention module (QPAM) which can capture compression noise levels for different QPs and emphasize meaningful features along channel dimension. Then we embed QPAM into the residual block, and based on it, we design a network architecture that is equipped with controllability for different QPs. To make the proposed model focus more on examples that have more compression artifacts or is hard to restore, a focal mean square error (MSE) loss function is employed to fine tune the network. Experimental results show that our approach achieves 4.03% BD-Rate saving on average for all intra configuration, which is even better than QP-separate CNN models while having less model parameters.
Xiaopeng Guo 0001, Mingyu Shang, Jun Sun 0012
DCC5
2021 Enhancing Knowledge Tracing via Adversarial Training
abstract
We study the problem of knowledge tracing (KT) where the goal is to trace the students' knowledge mastery over time so as to make predictions on their future performance. Owing to the good representation capacity of deep neural networks (DNNs), recent advances on KT have increasingly concentrated on exploring DNNs to improve the performance of KT. However, we empirically reveal that the DNNs based KT models may run the risk of overfitting, especially on small datasets, leading to limited generalization. In this paper, by leveraging the current advances in adversarial training (AT), we propose an efficient AT based KT method (ATKT) to enhance KT model's generalization and thus push the limit of KT. Specifically, we first construct adversarial perturbations and add them on the original interaction embeddings as adversarial examples. The original and adversarial examples are further used to jointly train the KT model, forcing it is not only to be robust to the adversarial examples, but also to enhance the generalization over the original ones. To better implement AT, we then present an efficient attentive-LSTM model as KT backbone, where the key is a proposed knowledge hidden state attention module that adaptively aggregates information from previous knowledge hidden states while simultaneously highlighting the importance of current knowledge hidden state to make a more accurate prediction. Extensive experiments on four public benchmark datasets demonstrate that our ATKT achieves new state-of-the-art performance. Code is available at: https://github.com/xiaopengguo/ATKT.
Xiaopeng Guo 0001, Mingyu Shang, Maojing Shu, Jun Sun 0012
ACM Multimedia6
2021 Adaptive Deep Reinforcement Learning-Based In-Loop Filter for VVC
abstract
Deep learning-based in-loop filters have recently demonstrated great improvement for both coding efficiency and subjective quality in video coding. However, most existing deep learning-based in-loop filters tend to develop a sophisticated model in exchange for good performance, and they employ a single network structure to all reconstructed samples, which lack sufficient adaptiveness to the various video content, limiting their performances to some extent. In contrast, this paper proposes an adaptive deep reinforcement learning-based in-loop filter (ARLF) for versatile video coding (VVC). Specifically, we treat the filtering as a decision-making process and employ an agent to select an appropriate network by leveraging recent advances in deep reinforcement learning. To this end, we develop a lightweight backbone and utilize it to design a network set S containing networks with different complexities. Then a simple but efficient agent network is designed to predict the optimal network from S , which makes the model adaptive to various video contents. To improve the robustness of our model, a two-stage training scheme is further proposed to train the agent and tune the network set. The coding tree unit (CTU) is seen as the basic unit for the in-loop filtering processing. A CTU level control flag is applied in the sense of rate-distortion optimization (RDO). Extensive experimental results show that our ARLF approach obtains on average 2.17%, 2.65%, 2.58%, 2.51% under all-intra, low-delay P, low-delay, and random access configurations, respectively. Compared with other deep learning-based methods, the proposed approach can achieve better performance with low computation complexity.
Jun Sun 0012, Xiaopeng Guo 0001, Mingyu Shang
IEEE Trans. Image Process.2
2020 Multi-Gradient Convolutional Neural Network Based In-Loop Filter For Vvc
abstract
While recent researches on convolutional neural network (CNN) based in-loop filters for High Efficiency Video Coding (HEVC) have achieved great success, the performance of these models on the new standard Versatile Video Coding (VVC) may degrade due to many novel adopted techniques which make the compression process more fine and capture more image details. In this work, the performances on VVC of two CNN based in-loop filters proposed for HEVC are investigated and a multi-gradient convolutional neural network based in-loop filter (MGNLF) for VVC is proposed. The proposed model exploits the divergence and second derivative of frame, which contain much potential image structural information, like contour information, to restore more detail information to further improve the quality of frames. Experimental results demonstrate our approach can significantly improve the coding performance. On average, 3.29% BD-Rate reduction is achieved on luma component under all intra configuration compared with the original VVC with DBF and SAO enabled, which also outperforms other state-of-the-art approaches for VVC.
Yunchang Li, Jun Sun 0012
ICME3
2020 Enhanced Cu Partitioning Search Method for Intra Coding in HEVC
abstract
Most encoders in High Efficiency Video Coding (HEVC) standard determine the coding unit (CU) partitions in a depth-first full search manner. For intra coding, this method ignores the impact of partitioning on the subsequent blocks, which incurs less accurate prediction and worsens encoding efficiency performance. In this paper, we present an enhanced CU partitioning search method, where the various partitioning types are examined to effectively mitigate such impact. To reduce the additional complexity, we further refine the search candidates, early terminate the search process, and reuse encoding information. Experimental results show that compared with the unmodified HM software, the proposed method achieves a bit rate saving of 0.34% on average and 0.67% at most. Since our method doesn't modify the original definition of HEVC, the encoded bitstreams can be uncompressed by a standard HEVC decoder with no complexity increase.
Yunchang Li, Jun Sun 0012
ICME2
2020 Character Region Awareness Network For Scene Text Recognition
abstract
Recognizing text in natural scenes is still a very challenging task, due to arbitrary shapes, varying fonts, complex backgrounds and so on. Recently, some recognizers utilize Spatial Transform Network (STN) to rectify irregular text instances and achieve promising results. However, their robustness and accuracy are still limited, since rectification performance can be easily degraded by challenging samples. To tackle this issue, we propose a simple yet effective two-dimensional (2D) character attention module, which can enhance foreground text instances via character region awareness. By incorporating this with existing rectification pipeline, we build a novel scene text recognizer named Character Region Awareness Network (CRAN). Extensive experiments demonstrate that our CRAN outperforms previous methods nearly on all benchmarks of both regular and irregular text, particularly on SVT (+2.0%), SVTP (+1.5%) and CUTE80 (+2.1%).
Mingyu Shang, Jun Sun 0012
ICME3
2020 Page-Level Handwritten Word Spotting via Discriminative Feature Learning
Xiaopeng Guo 0001, Mingyu Shang, Jun Sun 0012
KSEM (1)4
2020 Efficient HEVC Downscale Transcoding Based on Coding Unit Information Mapping
Yunchang Li, Jun Sun 0012
MMM (1)3
2020 An Efficient Encoding Method for Video Compositing in HEVC
Yunchang Li, Jun Sun 0012
MMM (1)3
2019 An Efficient Logo Insertion Method for Video Coding in HEVC
abstract
Inserting a logo into HEVC video streams is highly demanded in video applications. In this paper, we present an efficient logo insertion method for video coding in HEVC. To reduce the impact of inserted logo, the proposed method mitigates the encoding dependence on logo by partitioning the video frame into separated regions. For lossless coding region, we reduce the bit rate overhead of lossless coding according to an error propagation model. For information reusing region, we partly re-encode the quality-loss area to maintain the encoding quality. The computational complexity is reduced by reusing information. Experimental results show that the proposed method achieves a significant speedup ratio of 52 times on average with only 1.43% bit rate increase compared with direct encoding using unmodified HM software.
Yunchang Li, Jun Sun 0012
MMSP3
2018 Gradient Based Interpolation for Intra Angular Prediction in HEVC
abstract
In the High Efficiency Video Coding (HEVC), the predicted pixels generated by the intra angular prediction are the same along the prediction direction. In this paper, an improved gradient based interpolation for intra angular prediction is proposed. The gradient is generated by both row and column reference samples according to the prediction direction and changed dynamically for each pixel. It is appended to the original intra prediction process to improve the performance of interpolation. This method describes the features of directional gradient changes, which improves the performance of intra interpolation. The method achieves up to 1.90% BD-Rate reduction and ignorable decoding time increase under intra main configuration based on HM 16.7.
Yushan Zheng, Jun Sun 0012, Zongming Guo
ISCAS3
2018 Efficient Rate Control Method for Logo Insertion Video Coding in HEVC
abstract
In order to improve the network transmission stability of logo inserted video, an efficient rate control method for HEVC logo insertion is proposed in this paper. Based on the acceleration algorithm, the proposed mehtod optimizes bit allocation process in region and CTU level. First, each logo inserted video frame is separated into different regions to stop pixel error propagation. Then bits for different regions are allocated according to their coding characteristics. At last, target bit of CTUs are adjusted according to the video content information. Experimental results show that the proposed rate control method can achieve 4.46% BD-Rate saving on average with only 2.26% speed loss compared with the acceleration algorithm.
Yunchang Li, Yingfan Zhang, Jun Sun 0012
VCIP4
2017 A Novel Two-Step Integer-pixel Motion Estimation Algorithm for HEVC Encoding on a GPU
Keji Chen, Jun Sun 0012, Zongming Guo, Dachuan Zhao 0002
MMM (2)2
2016 A general PID-based rate adaptation approach for TCP-based live streaming over mobile networks
abstract
TCP-based application-layer protocols are increasingly applied to commercial live video streaming systems. However, in unstable mobile networks, the throughput of TCP may fluctuate rapidly due to its transmission mechanism, causing undesired playback interruption. In this paper, we propose a general application-layer rate adaptation approach to cope with the variability in TCP throughput. We analyze the transmission process and evaluate the network condition using a multi-buffer model. With information obtained in this model, an algorithm based on the Proportional-Integral-Derivative (PID) controller is proposed to dynamically adjust the video bitrate in response to the throughput change. We have implemented an experimental mobile live streaming system employing this approach and achieved a significant improvement in playback continuity and bandwidth utilization.
Jiexi Wang, Shengbin Meng, Jun Sun 0012, Zongming Quo
ICME3
2016 Efficient arbitrary ratio downscale transcoding for HEVC
abstract
The arbitrary ratio transcoding usually introduces coding block grid misalignment, which results in difficulties to utilize the decoding information in coding blocks of source videos during the encoding phase. To reduce the large computation load of HEVC downscale transcoding in such situation, we propose an efficient transcoding method that we refer the decoded coding unit (CU) partitioning to accelerate partition decision, which takes up the most complexity in the encoding phase as well as the whole transcoding. First, we predict CU depth in pixel level according to decoded partitioning of source videos. Then, we propose adaptive rules to determine CU partitioning of target videos based on the prediction, so that we can make early CU splitting or pruning decision without complex recursive search. Experiments demonstrate that the proposed method achieves about 74% time reduction on average with acceptable BD-rate increase in the encoding phase compared to the encoder in reference software HM13.0.
Zhenan Lin, Keji Chen, Jun Sun 0012, Zongming Guo
VCIP4
2016 An adaptive intra-frame parallel method based on complexity estimation for HEVC
abstract
Parallelization is an efficient solution for addressing the increased computational complexity in High Efficiency Video Coding (HEVC). To improve the intra-frame parallelism, an adaptive parallel method is proposed based on an encoding complexity model for HEVC. First, by establishing the relationship between encoding complexity and Rate Distortion Optimization (RDO) process, the encoding complexity is measured by the merge skip modes and coding unit partition statistics. Then a greedy algorithm is proposed to partition each frame into several independent regions for parallelism, and the encoding complexity of each region is precisely controlled to achieve computational complexity balancing for better parallelism. Extensive experimental results show that the proposed method can achieve up to 3.34× speedup against wave-front parallel processing (WPP), and 1.19× speedup against tiles with acceptable encoding efficiency loss for low delay video encoding.
Keji Chen, Jun Sun 0012, Xiangyang Ji, Zongming Guo
VCIP3
2016 A Novel Wavefront-Based High Parallel Solution for HEVC Encoding
abstract
With a lot of enhanced coding tools introduced, High Efficiency Video Coding (HEVC) achieves significant improvement in coding efficiency at the cost of increased computational complexity. To efficiently reduce the encoding time of HEVC, a wavefront-based high parallel (WHP) solution integrating novel data-level and task-level methods is proposed in this paper. On data level, optimal single-instruction-multiple-data algorithms are designed for the enhanced coding tools, i.e., replacing the multiplication in motion compensation by add and shift operations with reduced instruction cycles, removing the transpose in transform via realignment of coefficients, and minimizing the memory access in sum of absolute difference/sum of squared differences calculation by fully reusing the registers. On task level, a novel inter-frame wavefront (IFW) method is developed by effectively decreasing the dependence of wavefront parallel processing (WPP). In addition, a coding tree block level parallelism analysis method is presented to prove the superior of IFW method compared with other HEVC representative parallel methods. Besides, a three-level thread management scheme is proposed to best exploit the parallelism of IFW method and achieve corresponding encoding speedup. Extensive experimental results show that, the overall WHP solution can bring up to $57.65\times $ , $65.55\times $ , and $88.17\times $ speedup for HEVC encoding of Wide Video Graphics Array, 720p and 1080p standard test sequences, while maintaining the same coding performance as with WPP. The proposed solution is also applied in several leading video companies in China, providing HEVC video service for more than 1.3 million users everyday.
Keji Chen, Jun Sun 0012, Yizhou Duan, Zongming Guo
IEEE Trans. Circuits Syst. Video Technol.2
2016 Adaptive Video Streaming With Optimized Bitstream Extraction and PID-Based Quality Control
abstract
To cope with the challenges brought about by bandwidth fluctuation and improve the experience of watching online videos, an adaptive video streaming system that can adjust video quality according to actual network conditions is proposed based on the scalable video coding (SVC) extension of H.264/AVC. First, a simple and effective linear error model is proposed and verified for quality scalability of SVC. The model exploits the linear feature of pixel value errors and can be used to accurately estimate the distortion caused by discarding any combination of enhancement data packets in an SVC bitstream. On that basis, a greedy-like algorithm is designed to assign each data packet a priority value according to its rate-distortion (R-D) impact, thus enabling R-D optimized bitstream extraction under certain bitrate constraints. Finally, the proportional-integral-derivative (PID) method is utilized to control the video quality adjustment and determine a suitable bitrate for transmission. By monitoring and predicting the past, current, and future bandwidth information, the PID-based quality control algorithm is able to reduce quality fluctuation, while still preserving a high quality level. Experimental results show that compared with the baseline software, the proposed system that integrates the above algorithms can achieve much lower video quality fluctuation, with PSNR variance reduced from 1.24 to 0.69, and at the same time deliver higher video quality, with the PSNR average increased by 0.83 dB.
Shengbin Meng, Jun Sun 0012, Yizhou Duan, Zongming Guo
IEEE Trans. Multim.2
2015 Software Solution for HEVC Encoding and Decoding
Shengbin Meng, Jun Sun 0012, Zongming Guo
MMM (2)2
2015 Towards Rate-Distortion analysis of general source distributions: Property and principles
abstract
This paper decouples the complex Rate-Distortion (R-D) analysis problem by inspecting the respective influence of source distribution and quantizer design on R-D performance. First, a universal R-D property is theoretically revealed that, for any source distribution can be expressed as the product of a Scaling Factor (SF) and its remaining part, SF does not affect its derivative R-D function. Second, efficient quantizer design principles are deduced for different source distributions, which can be used as convenient R-D performance classifier and indicator when the dead-zone plus uniform threshold scalar quantizer with nearly-uniform reconstruction quantizer (DZ+UTSQ/NURQ) is applied. These two contributions bring new insight and inspiration towards the R-D analysis of various different sources, being solid infrastructure to benefit various video/image applications.
Jun Sun 0012, Yizhou Duan, Jiasi Shen 0001, Zongming Guo
MMSP1
2015 A two-stage fast CU size decision method for HEVC intracoding
abstract
In HEVC coding standard, the encoder employs a quad-tree-based Coding Unit (CU) structure to adapt to various texture characteristics of images. Although it provides better compression performance, the computation load increases drastically. In this paper, a two-stage fast CU size decision method is proposed for HEVC intra coding. In the first stage, we utilize the combination of the weighted variance and the maximum value of the Hadamard costs of the four sub CUs to represent the CU complexity, and classify each CU into three categories: compound, homogeneous, and undetermined. The former two kinds of CUs will be made early split and pruned decisions respectively. In the second stage, an improved SATD-based estimated R-D cost is employed to decide whether the prev undetermined CUs should be early pruned. Experimental results demonstrate that compared with the reference software HM13.0, the proposed fast CU size decision method provides 49% time reduction with slight quality degradation using the HEVC all intra test condition. This proposal has been adopted to provide HEVC video and picture encoding services at the server-side for UC mobile browser which covers over 100 million people in China.
Jun Sun 0012, Yizhou Duan, Zongming Guo
MMSP2
2015 Corrections to "Novel Efficient HEVC Decoding Solution on General-Purpose Processors"
abstract
In the above paper [ibid., vol. 16, no. 7, p. 1915, Nov. 2014], the sentence "has provided HEVC service to over 1500 million people in China via the Xunlei Kankan video client" should have appeared as "has provided HEVC services to over 150 million people in China via the Xunlei Kankan video client." Also. J. Sun should have been noted as the corresponding author.
Yizhou Duan, Jun Sun 0012, Leju Yan, Keji Chen, Zongming Guo
IEEE Trans. Multim.2
2014 An efficient method for no-reference H.264/SVC bitstream extraction
abstract
This paper investigates the no-reference SVC bitstream extraction problem and presents an efficient solution to approximate the “optimal” extracted sub-stream. First, we introduce a linear error model to accurately estimate the distortion caused by discarding any combination of packets, even when the original sequence is not available. Then we propose a greedy algorithm to decide each packet's priority according to its R-D impact. The priority value of packets can be stored in the bitstream and used for R-D optimized extraction. Experimental results show that our bitstream extraction method can achieve a significant PSNR gain compared to the extractors of JSVM, without computational complexity increment. Comparison with other methods also demonstrates the advantage of the proposed method.
Shengbin Meng, Jun Sun 0012, Yizhou Duan, Zongming Guo
ICASSP2
2014 A PID-based quality control algorithm for SVC video streaming
abstract
Scalable Video Coding (SVC) makes it possible to change video quality dynamically according to real-time bandwidth. For quality control algorithms of SVC video streaming, the biggest challenge is to keep a video quality that is both smooth and as good as possible. In this paper, we first introduce a combined quality level scheme to describe SVC video quality in a unified way. Then an effective and efficient quality control algorithm for SVC video streaming is proposed based on the Proportional-Integral-Derivative (PID) control method. Extensive experiments show that the proposed algorithm improves 8.6% in video quality with 24.8% reduction in quality fluctuation compared with the existing packet delay feedback algorithm. The proposed algorithm has also been implemented in online video website www.7dlive.com and performs well in applications.
Shengbin Meng, Jun Sun 0012, Zongming Guo
ICC2
2014 Towards efficient wavefront parallel encoding of HEVC: Parallelism analysis and improvement
abstract
High Efficiency Video Coding (HEVC) is the new generation video coding standard which achieves significant improvement in coding efficiency. Although HEVC is promising in many applications, the increased computational complexity is a serious problem, which makes parallelization necessary in HEVC encoding. To better understand the bottleneck of parallelization and improve the encoding speed, in this paper, we propose a Coding Tree Blocks (CTB) level parallelism analysis method as well as a novel Inter-Frame Wavefront (IFW) parallel encoding method. First, by establishing the relationship between parallelism and dependence, parallelism is precisely described by CTB-level dependence as a criterion to evaluate different parallel methods of HEVC. On this basis, by effectively decreasing the dependence based on Wavefront Parallel Processing (WPP), IFW method is developed. Finally, with the proposed parallelism analysis method, IFW is theoretically proved to be of higher parallelism compared with other HEVC representative parallel methods. Extensive experimental results show that, the proposed method and implementation can bring up to 17.81x, 14.34x and 24.40x speedup for HEVC encoding of WVGA, 720p and 1080p standard test sequences with the same ignorable coding performance degradation as WPP, thus showing a promising technology for future large-scale HEVC video application.
Keji Chen, Yizhou Duan, Jun Sun 0012, Zongming Guo
MMSP3
2014 Highly optimized implementation of HEVC decoder for general processors
abstract
In this paper, we propose a novel design and optimized implementation of the HEVC decoder. First, a novel decoder prototype with refined decoding workflow and efficient memory management is designed. Then on this basis, a series of single-instruction-multiple-data (SIMD) based algorithms are used to speed up several time-consuming modules in HEVC decoding. Finally, a frame-based parallel framework is applied to exploit the multi-threading technology on multicore processors. With the highly optimized HEVC decoder, decoding speed of 246fps on Intel i7-2400 3.4GHz quad-core processor for 1080p videos and 52fps on ARM Cortex-A9 1.2GHz dual-core processor for 720p videos can be achieved in our experiments.
Shengbin Meng, Yizhou Duan, Jun Sun 0012, Zongming Guo
MMSP3
2014 A low-latency peer-to-peer live and VOD streaming system based on scalable video coding
abstract
This paper demonstrates a peer-to-peer (P2P) live and VOD streaming system called Immedia based on Scalable Video Coding (SVC) with low latency. With features of layered coding in SVC, Immedia system could automatically adapt output video rate to varying network. Especially, the live system combining P2P streaming media transmission techniques works better in this respect and greatly reduces server pressure in a large-scale environment.
Jun Sun 0012, Yanping Zhou, Yizhou Duan, Zongming Guo
VCIP1
2014 Towards simple and smooth rate adaption for VBR video in DASH
abstract
Rate adaption in Dynamic Adaptive Streaming over HTTP (DASH) is widely applied to adapt the transmission rate to varying network capacity. For rate adaption on variable bitrate (VBR) encoded video, it is still a challenge to properly identify and address the dynamics of bandwidth and segment bitrate. In this paper, the trend of client buffer level variation (TBLV) is analyzed to be a more effective metric for detecting the dynamics of bandwidth and segment bitrate compared to previous metrics. Then, a partial-linear trend prediction model is developed to accurately estimate TBLV. Finally, based on the prediction model, a novel simple rate adaption algorithm is designed to achieve efficient and smooth video quality level adjustment. Experimental results show that while maintaining similar average video quality, the proposed algorithm achieves up to 47.3% improvement in rate adaption smoothness compared to the existing work.
Yanping Zhou, Yizhou Duan, Jun Sun 0012, Zongming Guo
VCIP3
2014 Novel Efficient HEVC Decoding Solution on General-Purpose Processors
abstract
Although the emerging video coding standard High Efficiency Video Coding (HEVC) successfully doubles the compression efficiency of H.264/AVC, its growing computational complexity makes real-time decoding of high-definition HEVC videos a very challenging issue for the existing personal computers and mobile devices. In this paper, a systematical, efficient HEVC decoding solution on general processors is provided, consisting of structure-level, data-level, and task-level approaches. First, a redesigned overall structure of a HEVC decoder with data redundancy reduction mechanism is introduced, which cuts down basic data operation cost and achieves an average decoding speedup of 2.37 × compared to the HM 10.0 decoder. On this basis, novel single-instruction multiple-data (SIMD) algorithms such as low-complexity motion compensation, transpose-free transform, symmetric deblocking filter, and parallel-index sample adaptive offset are developed, which further parallelize the data operations of each decoding task and bring another 2.67 × decoding speedup. Finally, a frame-based task-level parallel framework is employed with a flexible entry scheme to efficiently support the simultaneous processing of multiple decoding tasks for different HEVC parallel strategies. The overall solution achieves decoding fps of 40-75 for 4k HEVC videos on the Intel i7-2600 3.4 GHz quad-core processor (4-thread decoding) and 35-55 for 720p videos on the ARM Cortex-A9 1.2 GHz duo-core processor (2-thread decoding). This proposal is the recommended cross-platform HEVC decoding solution of Intel, AMD, and Cisco, and has provided HEVC service to over 1500 million people in China via the Xunlei Kankan video client.
Yizhou Duan, Jun Sun 0012, Leju Yan, Keji Chen, Zongming Guo
IEEE Trans. Multim.2
2013 Rate-Quantization and Distortion-Quantization Models of Dead-Zone Plus Uniform Threshold Scalar Quantizers for Generalized Gaussian Random Variables
Yizhou Duan, Jun Sun 0012, Zongming Guo
MMM (1)2
2013 Rate-Distortion Analysis of Dead-Zone Plus Uniform Threshold Scalar Quantization and Its Application - Part I: Fundamental Theory
abstract
This paper provides a systematic rate-distortion (R-D) analysis of the dead-zone plus uniform threshold scalar quantization (DZ+UTSQ) with nearly uniform reconstruction quantization (NURQ) for generalized Gaussian distribution (GGD), which consists of two aspects: R-D performance analysis and R-D modeling. In R-D performance analysis, we first derive the preliminary constraint of optimum entropy-constrained DZ+UTSQ/NURQ for GGD, under which the property of the GGD distortion-rate (D-R) function is elucidated. Then for the GGD source of actual transform coefficients, the refined constraint and precise conditions of optimum DZ+UTSQ/NURQ are rigorously deduced in the real coding bit rate range, and efficient DZ+UTSQ/NURQ design criteria are proposed to reasonably simplify the utilization of effective quantizers in practice. In R-D modeling, inspired by R-D performance analysis, the D-R function is first developed, followed by the novel rate-quantization (R-Q) and distortion-quantization (D-Q) models derived using analytical and heuristic methods. The D-R, R-Q, and D-Q models form the source model describing the relationship between the rate, distortion, and quantization steps. One application of the proposed source model is the effective two-pass VBR coding algorithm design on an encoder of H.264/AVC reference software, which achieves constant video quality and desirable rate control accuracy.
Jun Sun 0012, Yizhou Duan, Jiaying Liu 0001, Zongming Guo
IEEE Trans. Image Process.1
2013 Rate-Distortion Analysis of Dead-Zone Plus Uniform Threshold Scalar Quantization and Its Application - Part II: Two-Pass VBR Coding for H.264/AVC
abstract
In the first part of this paper, we derive a source model describing the relationship between the rate, distortion, and quantization steps of the dead-zone plus uniform threshold scalar quantizers with nearly uniform reconstruction quantizers for generalized Gaussian distribution. This source model consists of rate-quantization, distortion-quantization (D-Q), and distortion-rate (D-R) models. In this part, we first rigorously confirm the accuracy of the proposed source model by comparing the calculated results with the coding data of JM 16.0. Efficient parameter estimation strategies are then developed to better employ this source model in our two-pass rate control method for H.264 variable bit rate coding. Based on our D-Q and D-R models, the proposed method is of high stability, low complexity and is easy to implement. Extensive experiments demonstrate that the proposed method achieves: 1) average peak signal-to-noise ratio variance of only 0.0658 dB, compared to 1.8758 dB of JM 16.0's method, with an average rate control error of 1.95% and 2) significant improvement in smoothing the video quality compared with the latest two-pass rate control method.
Jun Sun 0012, Yizhou Duan, Jiaying Liu 0001, Zongming Guo
IEEE Trans. Image Process.1
2012 Rate-Distortion Analysis and Modeling of Dead-Zone Plus Uniform Threshold Scalar Quantization for Generalized Gaussian Random Variables
abstract
This paper provides a systematical rate-distortion (R-D) analysis and modeling of generalized Gaussian distribution (GGD) under the dead-zone plus uniform threshold scalar quantization (DZ+UTSQ) and nearly-uniform reconstruction quantization (NURQ). In R-D analysis, we clearly explain the property of GGD source under efficient DZ+UTSQ/NURQ. On this basis, in R-D modeling, the heuristic D-R model is proposed.
Yizhou Duan, Jun Sun 0012, Zongming Guo
DCC2
2012 Novel rate-distortion modeling for H.264/AVC and its application in two-pass VBR coding
abstract
In this paper, we first employ rate-distortion (R-D) modeling to derive the source model for H.264 video coding. This source model consists of the distortion-quantization (D-Q) model and distortion-rate (D-R) model, which describe the relationship between rate, distortion and quantization step for generalized Gaussian distribution (GGD) under the quantization scheme of H.264/AVC. The accuracy of the source model is confirmed by comparing the estimated results with the actual data of JM16.0. And the parameter selection criteria are developed to better employ the source model in our two-pass rate control algorithm for H.264 VBR coding. Based on the D-Q and D-R models, the proposed algorithm is of high stability, low complexity and is easy to implement. Extensive experiments covering different resolutions and target rates demonstrates: 1) about 96.7% reduction in PSNR variance compared to the method of JM16.0 with the average rate control error of 1.95%, and 2) significant improvement in smoothing the video quality compared with the latest two-pass rate control method.
Yizhou Duan, Jun Sun 0012, Zongming Guo
ISCAS2
2012 Optimized bit extraction of SVC exploiting linear error model
abstract
The Scalable Video Coding (SVC) extension of the H.264/AVC video coding standard supports fidelity or quality (SNR) scalability. The quality enhancement packets would be discarded in case of limited network capacity, which calls for an optimized bit extraction strategy. In this paper, we first analyze the linear feature in H.264/AVC video coding. A linear error model is also constructed using this feature in case of SVC quality scalability. Then based on the linear error model, the rate and distortion (R-D) impact of each quality enhancement packet over the whole sequence is obtained. Finally a new priority assigning algorithm is designed for a more efficient extraction, giving high rank to those with great R-D impacts. Extensive experiments are presented to demonstrate the accuracy of the linear error model and the validity of the priority assigning algorithm. Tests on the set of eight standard video sequences show the quality promotion under any bitrate constraint, and a fidelity gain up to 0.4 dB PSNR is achieved by the proposed strategy, compared to the JSVM reference software with Quality Layer information.
Jun Sun 0012, Jiaying Liu 0001, Zongming Guo
ISCAS2
2012 Implementation of HEVC decoder on x86 processors with SIMD optimization
abstract
High Efficient Video Coding (HEVC) is the next generation video coding standard in progress. Based on the traditional hybrid coding framework, HEVC implements enhanced tools to improve compression efficiency at the cost of far more computational payload than the capacity of real-time video applications. In this paper, we focus on the software implementation of a real-time HEVC decoder over modern Intel x86 processors. First, we identify the most time-consuming modules of HM 4.0 decoder, represented by motion compensation, adaptive loopfilter, deblocking filter and integer transform. Then the single-execution-multiple-data (SIMD) methods are proposed to optimize the computational performance of these modules. Experimental results show that the optimized decoder is more than 4 times faster than the HM 4.0 decoder, with decoding speed of over 40 frames per second for 1920×1080 resolution videos on Intel i5-2400 processor.
Leju Yan, Yizhou Duan, Jun Sun 0012, Zongming Guo
VCIP3
2012 An optimized real-time multi-thread HEVC decoder
abstract
This demonstration illustrates an optimized HEVC decoder, which is compliant with the reference software HM 4.0. The optimized decoder is more than 4 times faster than the reference decoder, and can well meet the real-time decoding demands of the 1080p high-definition (HD) videos. In addition, based on the frame-level multi-thread framework, the optimized decoder can even achieve up to 13.2 times speedup with 4-thread parallel decoding over Intel i5-2400 processor.
Leju Yan, Yizhou Duan, Jun Sun 0012, Zongming Guo
VCIP3
2011 Efficient dead-zone plus uniform threshold scalar quantization of generalized Gaussian random variables
abstract
This paper studies the rate-distortion (R-D) performance of entropy-constrained dead-zone plus uniform threshold scalar quantization (DZ+UTSQ) and nearly-uniform reconstruction quantization (NURQ) for generalized Gaussian distribution (GGD). We first derive the preliminary constraint of R-D optimized DZ+UTSQ/NURQ for GGD. Then for GGD source of actual DCT coefficients, the refined constraint and precise conditions of optimum DZ+UTSQ/NURQ are rigorously deduced in the real coding bit rate range. Based on above analysis, efficient DZ+UTSQ/NURQ design criteria are proposed to reasonably simplify the implementation of effective quantizer in practice.
Yizhou Duan, Jun Sun 0012, Jiaying Liu 0001, Zongming Guo
VCIP2
2011 A novel parallel encoding framework for scalable video coding
abstract
In this paper, we first propose a new parallel video coding framework, considering three important factors: parallel strategy, computational complexity and task scheduling. Then combining the characteristics of scalable video coding (SVC), a novel parallel encoding structure for temporal and quality scalabilities is introduced to obtain a high speedup of parallel SVC. Since the data dependencies in SVC are complex and time variant, the scheduling of parallel SVC is extremely difficult. In order to find the optimal scheduling solution, directed acyclic graph (DAG) is exploited to model the dependencies of encoding tasks, and the complexities of the encoding tasks which are accurately estimated by the Kalman filter to weight the scheduling tasks. Finally, two heuristic scheduling algorithms are also proposed to achieve a high encoding speed of parallel SVC. Experimental results show that the speedup of our parallel method was higher (about 60%) than previous work. Using the proposed method, high definition (HD) SVC videos can be encoded in real time.
Jun Sun 0012, Jiaying Liu 0001, Zongming Guo, Longshe Huo
VCIP2
2010 Fast weighted algorithms for bitstream extraction of SVC Medium-Grain scalable video coding
abstract
The Medium-Granular scalable (MGS) technologies in H.264/AVC-based scalable video coding (SVC) provide a flexible foundation to accommodate different network capacities. In order to make use of MGS in multiple environments conveniently, we need to obtain the Rate Distortion (R-D) function of SVC and design efficient bitstream extractions. In this paper, we proposed a simple and effective distortion model to estimate the reconstruction distortion with drift error. Based on the model, a simple and fast weight-based priority setting algorithm is designed to achieve optimal R-D performance in MGS bitstream extractions. A smooth extraction algorithm is also provided for smooth quality substream extraction. Extensive experiments show the accuracy of the new R-D models and the effectiveness of the proposed extraction algorithms.
Ruiheng Li, Jun Sun 0012, Wen Gao 0001
ICME2
2010 Novel Statistical Modeling, Analysis and Implementation of Rate-Distortion Estimation for H.264/AVC Coders
abstract
In H.264/advanced video coding, the encoder employs the rate-distortion optimization (RDO) to select the optimal coding mode of each block. Although it is effective to employ the RDO technique for mode decision, the computation load increases drastically. To reduce the computation complexity of the RDO technique, in this paper, we propose efficient algorithms for the estimation of block-level rate and distortion. For rate estimation, we model the transform coefficients with accurate generalized Gaussian distributions, and the weighted sum of absolute quantized transform coefficients is proposed as an efficient rate estimator, where the weights provide an implicit mechanism for evaluating different contributions of different frequency components to the coding bits. For distortion estimation, we first analyze the origins of distortion thoroughly. Then a direct relationship between the discarded bits in quantization and the distortion is explored. According to this investigation, a simple and efficient algorithm is proposed for the distortion estimation. With above proposed algorithms, the RDO technique can be efficiently implemented in a low-complexity way. Extensive experimental results demonstrate that, compared with the original RDO implementation, the proposed algorithms achieve about 32% reduced total encoding time with ignorable coding performance degradation.
Xin Zhao 0003, Jun Sun 0012, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2009 On Rate-Distortion Modeling and Extraction of H.264/SVC Fine-Granular Scalable Video
abstract
Fine-granular scalable (FGS) technologies in H.264/AVC-based scalable video coding (SVC) provide a flexible foundation to accommodate different network capacities. To support efficient quality extraction, it is important to obtain the rate-distortion (R-D) or Distortion-Rate (D-R) function of each individual picture or a group of pictures (GOP). In this paper, firstly, the R-D function of SVC FGS pictures is analyzed with generalized Gaussian model and the D-R curve is proved to be a concave function overall. Considering the current sub-bitplane technology, the D-R function is revisited and inferred to be linear under MSE criterion within an FGS level, which also explains why the observed D-R curve with PSNR criterion is a piece-wise convex function. Secondly, the drift issue of SVC is analyzed, and a simple and effective distortion model is proposed to estimate the reconstruction distortion with drift error. Thirdly, with the above analysis and models, a virtual GOP concept is introduced, and a new priority setting algorithm is designed to achieve the optimal R-D performance in a virtual GOP. The D-R slope of each FGS packet and the D-R function of each virtual GOP are also obtained during the process. Finally, the D-R slopes of FGS levels are used in quality layer assignment to achieve equivalent coding efficiency to the SVC test model but with significantly reduced complexity. The D-R functions of virtual GOPs are utilized to design a practical method for smooth quality reconstruction. Compared to the prior methods, the smoothed video quality is improved not only objectively but also subjectively.
Jun Sun 0012, Wen Gao 0001, Debin Zhao, Weiping Li 0002
IEEE Trans. Circuits Syst. Video Technol.1
2007 Direct Mode Coding for B Pictures using Virtual Reference Picture
abstract
The direct mode used in the bi-predictive pictures (B-pictures) can efficiently improve the coding performance of B pictures, because it has small overhead and obtains a predictive picture from two reference pictures. The traditional temporal direct mode (TDM) derives the motion vector of the current block by scaling the motion vector of the co-located block in the backward reference picture. However, when the current block and its co-located block in backward reference picture belong to different objects with different motion directions, the prediction efficiency of TDM is drastically reduced. In this paper, we propose an improved direct mode prediction method. In the method, a virtual reference picture is generated using the pixel projection technique. Then the virtual reference picture is used to predict the direct mode blocks in B pictures. The proposed method can enhance the prediction performance of the direct mode blocks and achieve a higher coding efficiency.
Debin Zhao, Jun Sun 0012, Wen Gao 0001
ICME3
2007 Statistical Analysis and Modeling of Rate-Distortion Function in SVC Fine-Granular SNR Scalable Videos
abstract
In this paper, we propose a simple and effective rate-distortion model of fine-granular SNR scalable (FGS) layer in JVT scalable video coding (SVC). First, we introduce generalized Gaussian distributions (GGD) to model the distributions of the 16 (4*4) integer transform coefficients of the SVC FGS residue picture. Then we analyze the quantization scheme of SVC FGS coding, under which we analyze the distortion-rate function of generalized Gaussian model and conclude that the derivative of the distortion-rate function usually decreases to the traditional number 6.02 dB/bit as the coding rate increases. For some special cases where the shape of GGD is smooth, the derivative of the distortion-rate function increases to 6.02 dB/bit as the rate increases in the actual coding rate range. Guided by the observations, an effective and flexible rate-distortion model is proposed to approximate the actual rate-distortion function of SVC FGS layer. The average estimation error is only 0.07 dB in our extensive experiments. And our analysis and models bring us much insight into the SVC FGS coding and its rate-distortion function.
Jun Sun 0012, Wen Gao 0001, Debin Zhao
ICME1
2006 Statistical model, analysis and approximation of rate-distortion function in MPEG-4 FGS videos
abstract
In this paper, the generalized Gaussian distribution is employed first to model the DCT coefficients of image data from MPEG-4 fine-granularity scalability (FGS) frame. Then, according to the quantization theory, the distortion-rate function of the generalized Gaussian model is analyzed and it is concluded that the derivative of the distortion-rate function first decreases, and then increases up to the boundary of 6.02 as the bit rate increases. For actual FGS coding, the derivative of actual distortion-rate function usually decreases as the rate increases, and then begins to increase slowly at a comparatively high bit rate. Finally, based on above observations, a rate-distortion (R-D) model is proposed to approximate the actual distortion-rate function. Experiments show that the proposed R-D model is accurate and flexible.
Jun Sun 0012, Wen Gao 0001, Debin Zhao, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.1
2004 A novel FGS base-layer encoding model and weight-based rate adaptation for constant-quality streaming
abstract
We examine the coding problem of base layer (BL) in fine granularity scalable videos and rate adaptation of enhancement layer (EL) during streaming in order to support constant quality streaming. The problem arises from the facts that different frames of BL often exhibit significant quality variation in the usual FGS BL encoding methods, and consequently the actual EL rate-distortion (R-D) curves of different frames present significant difference. Under this condition it would be computationally exhaustive to scale the EL to flatten out the fluctuating BL quality and to attain constant quality. We propose, in this paper, an accurate constant quality encoding method of BL, under which we investigate the similarity of actual EL R-D curves, and then a simple but effective weight-based EL rate allocation algorithm is introduced. Experiments show, compared with default methods, our fast and simple methods not only provide constant-quality streaming, but also decrease the average MSE of videos.
Jun Sun 0012, Wen Gao 0001, Qingming Huang
ICIG1