Yan Huang 0033

dblp:75/6434-33 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
7since 2021 · last 2022
0000-0002-5548-2727ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2022 CNN-Based Fast CU Partitioning Algorithm for VVC Intra Coding
abstract
Over a year has passed since the finalization of Versatile Video Coding (H.266/VVC), yet it is still far from practical deployment, a major reason being the excessive complexity. The flexible and sophisticated quad-tree with nested multi-type tree partitioning structure in VVC provides considerable performance gains while bringing about an exponential increase in encoding time. To reduce the coding complexity, this paper proposes a Convolutional Neural Network (CNN) based fast Coding Unit (CU) partitioning algorithm for intra coding, which accelerates CU partition through predicting the partition modes with texture information and terminating redundant modes in advance. Corresponding classifiers are designed for different CU sizes to improve prediction accuracy. Low rate-distortion performance degradation is guaranteed by introducing performance loss due to misclassification into the loss function. Experiments show that the proposed method can save encoding time ranging from 38.39% to 62.33% with 0.92% to 2.36% bit rate increase.
Jun Xu 0040, Yan Huang 0033, Li Song 0001
ICIP4
2022 Generative Compression for Face Video: A Hybrid Scheme
abstract
As the latest video coding standard, versatile video coding (VVC) has shown its ability in retaining pixel quality. To excavate more compression potential for video conference scenarios under ultra-low bitrate, this paper proposes a bitrate-adjustable hybrid compression scheme for face video. This hybrid scheme combines the pixel-level precise recovery capability of traditional coding with the generation capability of deep learning based on abridged information, where Pixel-wise Bi-Prediction, Low-Bitrate-FOM and Lossless Keypoint Encoder collaborate to achieve PSNR up to 36.23 dB at a low bitrate of 1.47 KB/s. Without introducing any additional bi-trate, our method has a clear advantage over VVC under a completely fair comparative experiment, which proves the effectiveness of our proposed scheme. Moreover, our scheme can adapt to any existing encoder/configuration to deal with different encoding requirements, and the bitrate can be dynamically adjusted according to the network condition.
Anni Tang, Yan Huang 0033, Jun Ling, Zhiyu Zhang 0010, Rong Xie 0004, Li Song 0001
ICME2
2022 Intra Encoding Complexity Control with a Time-Cost Model for Versatile Video Coding
abstract
For the latest video coding standard Versatile Video Coding (VVC), the encoding complexity is much higher than previous video coding standards to achieve a better coding efficiency, especially for intra coding. The complexity becomes a major barrier of its deployment and use. Even with many fast encoding algorithms, it is still practically important to control the encoding complexity to a given level. Inspired by rate control algorithms, we propose a scheme to precisely control the intra encoding complexity of VVC. In the proposed scheme, a Time-PlanarCost (viz. Time-Cost, or T-C) model is utilized for CTU encoding time estimation. By combining a set of predefined parameters and the T-C model, CTU-level complexity can be roughly controlled. Then to achieve a precise picture-level complexity control, a framework is constructed including uneven complexity pre-allocation, preset selection and feedback. Experimental results show that, for the challenging intra coding scenario, the complexity error quickly converges to under 3.21%, while keeping a reasonable time saving and rate-distortion (RD) performance. This proves the efficiency of the proposed methods.
Yan Huang 0033, Jizheng Xu, Yan Zhao 0041, Li Song 0001
ISCAS1
2021 SVM Based Fast CU Partitioning Algorithm for VVC Intra Coding
abstract
Recently, Joint Video Experts Team (JVET) has completed the new Versatile Video Coding (H.266/VVC) standard. VVC employs a new block partition structure named quad-tree with nested multi-type tree (QTMT) to improve coding efficiency. However, the new block partition structure increases huge encoding time compared with HEVC for brute-force ratedistortion (RD) optimization. To reduce encoding complexity, we propose a Support Vector Machine (SVM) based fast CU partitioning algorithm for VVC intra coding in this paper which terminates redundant partitions early by predicting the partition of CU using texture information. We trained classifiers for CUs of different sizes to improve accuracy and control the complexity of the classifiers themselves. Different thresholds are set for each classifier to achieve a trade-off between encoding complexity and RD performance. Experimental results show that the proposed method can save encoder time ranging from 30.78% to 63.16% with 1.10% to 2.71% BD-BR increase.
Yan Huang 0033, Li Song 0001, Wenjun Zhang 0001
ISCAS2
2021 HEVC VMAF-oriented Perceptual Rate Distortion Optimization using CNN
abstract
Video coding standards like HEVC and VVC have achieved significant coding performance. However, the RDO module in coding framework ignores the characteristics of human visual system (HVS), which leads to insufficiency for perceptual video coding. Recently, learning-based objective assessment metric VMAF is developed and has been demonstrated higher quality assessment accuracy than conventional metrics. To incorporate VMAF into RDO aiming at improving perceptual coding efficiency, in this paper, a perceptual RDO scheme is proposed. A CNN-based on-line training method is first explored to determine the VMAF-related distortion estimation coefficient. Based on the VMAF-related coefficient and R-D model, a VMAF-based Lagrangian multiplier is proposed to adjust the R-D performance of each coding block. Experiments demonstrate that the proposed method can achieve an average -2.80% VMAF-based BD-Rate compared with the original HEVC, which effectively improves the coding performance.
Yan Huang 0033, Rong Xie 0004, Li Song 0001
PCS2
2021 Modeling Acceleration Properties for Flexible INTRA HEVC Complexity Control
abstract
It is a very well-known fact, that the high complexity of the High Efficiency Video Coding standard (HEVC) is the main hurdle for its wide deployment and use. To tackle this problem, a number of recent research outcomes exploit heuristic algorithms and machine learning, including deep learning, to reduce the coding complexity. However, in most cases, each encoder module, i.e., encoding process, is first accelerated individually, and then different acceleration algorithms are manually combined. Without a holistic strategy, the acceleration potential of multi-module combination is not exploited and the Rate-Distortion (RD) loss is generally not well controlled. To tackle these shortcomings, this paper exploits the acceleration properties of different modules, i.e., the numerical representation of potential time saving and possible RD loss, from which a heuristic model is explored. Then a Heuristic Model Oriented Framework (HMOF) is proposed which adapts the properties of modules to underlying acceleration algorithms. In the framework, two advanced acceleration algorithms, including Border Considered CNN (BC-CNN)-based Coding Unit (CU) partition and Naive Bayes-based Prediction Unit (PU) partition, are proposed for the CU and PU modules, respectively. Further, by leveraging the heuristic model as the guidance to combine the proposed acceleration algorithms, HMOF is globally optimized, where different time saving budgets are wisely allocated to different modules and a theoretically minimal RD loss is achieved. According to the experimental results, through fusing a suitable deep learning technique and a Bayes-Based prediction, the proposed acceleration framework HMOF enable multiple acceleration choices. Here the proposed joint optimization strategy help to make a choice leading to the best cost-performance. Furthermore, within the proposed framework, intra coding time can be precisely controlled with negligible Bjøntegaard delta bit-rate (BDBR) loss. In this context, as a complexity control method, HMOF outperforms the state-of-the-art complexity reduction algorithms under a similar complexity reduction ratio. These results partially demonstrate the superiority of the proposed technique.
Yan Huang 0033, Li Song 0001, Rong Xie 0004, Ebroul Izquierdo, Wenjun Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2021 VMAF Oriented Perceptual Coding Based on Piecewise Metric Coupling
abstract
It has been recognized that videos have to be encoded in a rate-distortion optimized manner for high coding performance. Therefore, operational coding methods have been developed for conventional distortion metrics such as Sum of Squared Error (SSE). Nowadays, with the rapid development of machine learning, the state-of-the-art learning based metric Video Multimethod Assessment Fusion (VMAF) has been proven to outperform conventional ones in terms of the correlation with human perception, and thus deserves integration into the coding framework. However, unlike conventional metrics, VMAF has no specific computational formulas and may be frequently updated by new training data, which invalidates the existing coding methods and makes it highly desired to develop a rate-distortion optimized method for VMAF. Moreover, VMAF is designed to operate at the frame level, which leads to further difficulties in its application to today's block based coding. In this paper, we propose a VMAF oriented perceptual coding method based on piecewise metric coupling. Firstly, we explore the correlation between VMAF and SSE in the neighborhood of a benchmark distortion. Then a rate-distortion optimization model is formulated based on the correlation, and an optimized block based coding method is presented for VMAF. Experimental results show that 3.61% and 2.67% bit saving on average can be achieved for VMAF under the low_delay_p and the random_access_main configurations of HEVC coding respectively.
Zhengyi Luo 0001, Yan Huang 0033, Rong Xie 0004, Li Song 0001, C.-C. Jay Kuo
IEEE Trans. Image Process.3
2020 Learning-Based Quality Enhancement For Scalable Coded Video Over Packet Lossy Networks
abstract
The layered feature of scalable video coding (SVC) offers a sufficient adaptation to unreliable transmission. When network condition drops sharply, enhancement layers will be abandoned, and only base layers are delivered. However, this will cause noticeable visual artifacts due to quality differences between different layers. To alleviate this problem, we novelly introduce a deep learning-based method into video reconstruction phase of scalable bitstreams. A super-resolution motivated recurrent network is proposed to extract and fuse features from both previous high-resolution frames and the current low-resolution frame. To the best of our knowledge, this is the first attempt to improve the performance of scalable bitstreams reconstruction by a specifically designed super-resolution network. By seamlessly integrating the accessible features, significant video quality improvements in terms of PSNR, SSIM, and VMAF are achieved. At the same time, the improvement of overall visual quality stability is apparent under packet lossy networks, indicating both efficiency and robustness of our approach.
Shengwei Yu, Xun Tong, Yan Huang 0033, Rong Xie 0004, Li Song 0001
ICME3
2019 VMAF Oriented Perceptual Optimization for Video Coding
abstract
In the light of low costs and automatic assessment, objective visual quality metrics enjoy many important applications such as perceptual coding. Recently multiple metrics obtain further improvement by means of machine learning. However, due to the absence of specific formulas, it's often hard to incorporate learning based metrics into video coding. In this paper, taking the state-of-the-art learning based metric VMAF for example, we propose a method of perceptual coding in an inferential manner for learning based metrics. The rate distortion optimization is adapted during coding as well. Experimental results show that compared with conventional methods, the proposed method can achieve obvious bitrate saving under HEVC coding.
Zhengyi Luo 0001, Yan Huang 0033, Rong Xie 0004, Li Song 0001
ISCAS2
2019 CNN Accelerated Intra Video Coding, Where Is the Upper Bound?
abstract
The very high complexity of the High Efficiency Video Coding standard (HEVC) is the main hurdle for its wide deployment and use. To tackle this problem, a number of recent research outcomes exploit Convolutional Neural Network (CNN) in each HEVC module for reducing the coding complexity. In this paper an effective method to analyse the potential of CNN techniques to reduce the computational cost of HEVC is proposed. A theoretical upper bound for the effectiveness of this approach in common HEVC modules is investigated. The theoretical maximum of learning-based complexity reduction in HEVC and possible reasons for Rate-Distortion (RD) loss are investigated. On the basis of this analysis, an Intra Video Coding Acceleration (IVCA) scheme is proposed, where Border Considered CNN (BC-CNN) based Coding Unit (CU) partition and heuristic Prediction Unit (PU) partition are seamlessly integrated. According to the experimental results, 66.7% of intra coding time can be saved with negligible 1.71% Bjøntegaard delta bit-rate (BDBR) loss. These results partially demonstrate the superiority of the proposed technique against other state-of-the-art approaches aiming at reducing HEVC complexity in intra mode.
Yan Huang 0033, Li Song 0001, Ebroul Izquierdo
PCS1
2018 GPU Based Motion-Compensated Frame Interpolation Acceleration for Future Video Coding
abstract
Being developed by Joint Video Exploration Team (JVET), Future Video Coding (FVC) aims at higher resolutions and higher compression performance than the state-of-the-art HEVC standard, undoubtedly at the cost of further computing increases. As an efficient computing platform, Graphics Processing Unit (GPU) is often used to accelerate encoding. But with the adoption of instruction set acceleration in the reference software of FVC, previous methods often become less efficient or even lead to a lower speed. In this paper, based on the comparative analysis of the time consumption between HEVC and FVC, we propose a GPU based acceleration method for the most computation-intensive step - frame interpolation of FVC, where frame caching strategy and a multi-stream mechanism is designed to make the best of GPU resources. Experimental results show that compared with the instruction set accelerated reference software of FVC, our method could achieve average 67.12% speed-up gains on the interpolation module and average 6.35% speed-up gains on overall encoding with exactly the same performance as before.
Jianlun Tang, Yan Huang 0033, Rong Xie 0004, Zhengyi Luo 0001, Li Song 0001
ICIP2
2018 An MCMC based Efficient Parameter Selection Model for x265 Encoder
abstract
As an open-source and computationally efficient High Efficiency Video Coding (HEVC) encoder, x265 has been gaining increasing popularity in video applications. x265 provides numerous encoding parameters in view of flexibility. However, proper and efficient setting of parameters often becomes a great challenge in practice. In this paper, we deeply investigate the influence of x265 parameters based on the Slow preset and pick out important parameters in terms of efficiency and complexity. Then a Markov Chain Monte Carlo (MCMC) based algorithm is proposed for efficient parameter adaptation at the target encoding time. This paper shows that carefully selected low-complexity encoding configurations can achieve the coding efficiency comparable to that of high-complexity ones. Specifically, average 26.72% encoding time reduction can be achieved while maintaining similar Rate Distortion (RD) performance to x265 presets using the proposed algorithm.
Yan Huang 0033, Li Song 0001, Rong Xie 0004, Zhengyi Luo 0001
ISCAS1