VLDB 2026 Research / reviewers in the wild / expert
Hao Yang 0008
dblp:54/4089-8
· DBLP profile ↗
14ranked-venue papers
3as first author
7since 2021 · last 2024
0009-0008-9115-4174ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | QVD: Post-training Quantization for Video Diffusion ModelsabstractRecently, video diffusion models (VDMs) have garnered significant attention due to their notable advancements in generating coherent and realistic video content. However, processing multiple frame features concurrently, coupled with the considerable model size, results in high latency and extensive memory consumption, hindering their broader application. Post-training quantization (PTQ) is an effective technique to reduce memory footprint and improve computational efficiency. Unlike image diffusion, we observe that the temporal features, which are integrated into all frame features, exhibit pronounced skewness. Furthermore, we investigate significant inter-channel disparities and asymmetries in the activation of video diffusion models, resulting in low coverage of quantization levels by individual channels and increasing the challenge of quantization. To address these issues, we introduce the first PTQ strategy tailored for video diffusion models, dubbed QVD. Specifically, we propose the High Temporal Discriminability Quantization (HTDQ) method, designed for temporal features, which retains the high discriminability of quantized features, providing precise temporal guidance for all video frames. In addition, we present the Scattered Channel Range Integration (SCRI) method which aims to improve the coverage of quantization levels across individual channels. Experimental validations across various models, datasets, and bit-width settings demonstrate the effectiveness of our QVD in terms of diverse metrics. In particular, we achieve near-lossless performance degradation on W8A8, outperforming the current methods by 205.12 in FVD. Shilong Tian, Hong Chen 0014, Chengtao Lv, Yu Liu 0031, Jinyang Guo 0002, Xianglong Liu 0001, Shengxi Li, Hao Yang 0008 |
ACM Multimedia | 8 |
| 2024 | Blind Quality Enhancement for Compressed VideoabstractDeep convolutional neural networks (CNNs) have achieved impressive success in enhancing the quality of compressed images/videos. These approaches mostly obtain the noise level in advance and train multiple architecture-identical models for enhancement on images/videos of known levels of noise. It largely hinders their practical applications where the noise level is unknown and resource is limited. To practically perform quality enhancement, we propose a novel blind quality enhancement framework for compressed video (BQEV), which utilizes a single network to conduct enhancement on videos compressed at various and unknown quality parameters (QPs). Since there exists feature similarity and difference among videos compressed at multiple QPs, BQEV utilizes this prior to efficiently handle enhancement on videos compressed at blind QPs, which consists of progressive feature extraction and QP-adaptive feature fusion subnets. They utilize temporal information and feature similarity to progressively extract valuable features and further employ the feature difference to conduct reasonable QP-adaptive feature fusion and quality enhancement, respectively. In the progressive feature extraction subnet, we first design a quality rank module to assign more attention to higher-quality frames for efficient utilization of temporal information, then propose a progressive extraction module to further extract features from different QPs. In the QP-adaptive feature fusion subnet, we develop a quality estimation module to guide reasonable feature fusion of these extracted progressive features for stable and promising enhancement results on multiple QPs. Experimental results demonstrate that BQEV achieves 0.31–0.69 dB PSNR improvement compared with videos compressed at various QPs, outperforming state-of-the-art approaches. Liquan Shen, Liangwei Yu, Hao Yang 0008, Mai Xu |
IEEE Trans. Multim. | 4 |
| 2023 | E-Net: a novel deep learning framework integrating expert knowledge for glaucoma optic disc hemorrhage segmentation
Yong-Li Xu, Hao Yang 0008, Shuai Lu 0003, Haihui Wang, Man Hu 0003 |
Multim. Tools Appl. | 3 |
| 2022 | Effective QTMT Partition Decision Algorithm for VVC IntercodingabstractAiming at the problem of high complexity of Versatile Video Coding (VVC) inter coding, this paper proposes a QTMT partition decision algorithm based on a multi-level decision framework. Specifically, the multi-level decision framework decomposes the multi-mode partition decision problem into multiple independent single-mode partition decision problems and then adopts a classification method based on machine learning to predict result of each single mode partition. To design more efficient classification features, fast motion estimation on 4×4 blocks is first performed to construct a motion field, and features on global/local motion and global/local consistency of its corresponding residuals are designed to measure motion activity and texture homogeneity. Furthermore, a misclassification protection mechanism is designed to decrease influences of misclassification on coding performance loss. Experimental results show that the proposed effective QTMT partition decision algorithm achieves a computational complexity reduction more than 51%, while incurring 1.65% BDBR increase compared with that of the original coding in the test model of VVC(VTM). Liquan Shen, Hao Yang 0008, Shiwei Wang 0005 |
MMSP | 2 |
| 2022 | Fast Intra Mode Decision Algorithm for Versatile Video CodingabstractTo achieve higher coding efficiency, the latest Versatile Video Coding (VVC) standard adopts a series of new intra coding techniques, including the quadtree plus multi-type tree (QTMT), intra sub-partitions (ISP) and intra block copy (IBC). However, this makes the intra coding more complicated, as VVC needs to traverse all prediction modes and partition types of QTMT to find the optimal combination. In this paper, we propose a fast algorithm for VVC from two aspects of mode selection and prediction terminating to reduce coding complexity. For the mode selection, adaptive mode pruning (AMP) is proposed to remove non-promising modes. First, since the newly introduced modes (IBC and ISP) are not effective for all blocks, learning-based classifiers are designed to remove them intelligently. Second, for normal modes, an ensemble decision strategy is proposed to sort the candidate modes and increase the probability of being the optimal mode for the first few candidates; thus, we can remove redundant candidates more efficiently. In terms of prediction terminating, we find that different optimal modes of current depth level lead to different termination probabilities of remaining intra predictions. Therefore, mode-dependent termination (MDT) is proposed to select an appropriate model through the optimal mode and terminate unnecessary intra predictions of remaining depth levels. The proposed algorithm is implemented on VVC test model, and simulation results show that it can achieve 51%$\sim$53% time savings with only 0.93%$\sim$1.08% BDBR increases. Xinchao Dong, Liquan Shen, Mei Yu 0001, Hao Yang 0008 |
IEEE Trans. Multim. | 4 |
| 2021 | A Distortion-Aware Multi-Task Learning Framework for Fractional Interpolation in Video CodingabstractMotion-compensated prediction adopts fractional-pixel interpolation to obtain the best motion vector. Traditional fixed interpolation filters cannot handle various content and structures well, and existing convolutional neural network based methods cannot fully exploit the distortion characteristics for fractional interpolation. Therefore, this paper proposes a distortion-aware multi-task learning framework (DA-MLF) to perform fractional interpolation. First, a multi-task training framework is proposed to provide the distortion characteristics as complementary information for improving the performance of subsequent interpolation. Then, a uniform interpolation sub-network is proposed to accomplish fractional interpolation, which utilizes the feature fusion module to fuse abundant local features, and the distortion awareness module to capture the multi-scale information of compression artifacts. Furthermore, DA-MLF is integrated into High Efficiency Video Coding (HEVC) test model, and multiple experiments are performed to evaluate the effectiveness of our method. On HEVC testing sequences, DA-MLF achieves 5.0%, 4.0% and 1.7% BD-rate reduction on average compared to the HEVC baseline, under low-delay P, low-delay B and random-access configurations, respectively. The experimental results validate that our framework not only achieves the best interpolation performance but also has the lowest computational complexity compared with state-of-the-art methods. Liangwei Yu, Liquan Shen, Hao Yang 0008, Xuhao Jiang, Bo Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Patch-Wise Spatial-Temporal Quality Enhancement for HEVC Compressed VideoabstractRecently, many deep learning based researches are conducted to explore the potential quality improvement of compressed videos. These methods mostly utilize either the spatial or temporal information to perform frame-level video enhancement. However, they fail in combining different spatial-temporal information to adaptively utilize adjacent patches to enhance the current patch and achieve limited enhancement performance especially on scene-changing and strong-motion videos. To overcome these limitations, we propose a patch-wise spatial-temporal quality enhancement network which firstly extracts spatial and temporal features, then recalibrates and fuses the obtained spatial and temporal features. Specifically, we design a temporal and spatial-wise attention-based feature distillation structure to adaptively utilize the adjacent patches for distilling patch-wise temporal features. For adaptively enhancing different patch with spatial and temporal information, a channel and spatial-wise attention fusion block is proposed to achieve patch-wise recalibration and fusion of spatial and temporal features. Experimental results demonstrate our network achieves peak signal-to-noise ratio improvement, 0.55 - 0.69 dB compared with the compressed videos at different quantization parameters, outperforming state-of-the-art approach. Liquan Shen, Liangwei Yu, Hao Yang 0008, Mai Xu |
IEEE Trans. Image Process. | 4 |
| 2020 | VRFCNN: Virtual Reference Frame Generation Network for Quality SHVCabstractFor more efficient inter prediction in quality scalable high efficiency video coding (SHVC), a learning-based framework for generating a virtual reference frame (VRF) is proposed in this letter. In our method, reconstructed base layer (BL) and enhancement layer (EL) frames are employed to make the generated VRF and current EL frame as same as possible. To this end, this letter proposes a novel VRF generation convolutional neural network (VRFCNN) to jointly handle enhancement of corresponding BL frame and compensation of previous EL frame. Specifically, the VRFCNN consists of BL enhancement, EL compensation and feature fusion subnets. Previous EL frames are firstly compensated with the learned coarse flow between two adjacent BL frames and then adopted to provide interlayer information for corresponding BL enhancement. The learned finer flow between the EL and enhanced BL features is adopted to provide temporal information for previous EL compensation. For efficiently handling slow- and fast-motion videos, the enhanced BL and compensated EL features are fused to generate a VRF. Experimental results show that VRFCNN averagely achieves 11.8% BD-rate reduction under low delay P configuration, which outperforms other methods. The code of our VRFCNN approach is available at https://github.com/dq0309/VRFCNN. Liquan Shen, Hao Yang 0008, Xinchao Dong, Mai Xu |
IEEE Signal Process. Lett. | 3 |
| 2020 | Low-Complexity CTU Partition Structure Decision and Fast Intra Mode Decision for Versatile Video CodingabstractQuadtree with nested multi-type tree (QTMT) partition structure is an efficient improvement in versatile video coding (VVC) over the quadtree (QT) structure in the advanced high-efficiency video coding (HEVC) standard. With the exception of the recursive QT partition structure, recursive multi-type tree partition is applied to each leaf node, which generates more flexible block sizes. Besides, intra prediction modes are extended from 35 to 67 so as to satisfy various texture patterns. These newly developed techniques achieve high coding efficiency but also result in very high computational complexity. To tackle this problem, we propose a fast intra-coding algorithm consisting of low-complexity coding tree units (CTU) structure decision and fast intra mode decision in this paper. The contributions of the proposed algorithm lie in the following aspects: 1) the new block size and coding mode distribution features are first explored for a reasonable fast coding scheme; 2) a novel fast QTMT partition decision framework is developed, which can determine the partition decision on both QT and multi-type tree with a novel cascade decision structure; and 3) fast intra mode decision with gradient descent search is introduced, while the best initial search point and search step are also investigated in this paper. The simulation results show that the complexity reduction of the proposed algorithm is up to 70% compared to VVC reference software (VTM), and averagely 63% encoding time saving is achieved with 1.93% BDBR increasing. Such results demonstrate that our method yields a superior performance in terms of computational complexity and compression quality compared to the state-of-the-art methods. Hao Yang 0008, Liquan Shen, Xinchao Dong, Ping An 0001, Gangyi Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | A content-based rate control algorithm for screen content video coding
Liquan Shen, Hao Yang 0008, Ping An 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Quality Enhancement Network via Multi-Reconstruction Recursive Residual Learning for Video CodingabstractLossy compression algorithms introduce multiple compression artifacts that severely decrease visual quality. These compression artifacts are highly related to texture contents, and the hierarchical coding units decision structure also brings multi-scale similarity to these artifacts. Current loop filters fail to utilize these characteristics to comprehensively remove compression artifacts. To this end, this letter proposes a novel quality enhancement method by adopting a multi-reconstruction recurrent residual network (MRRN). In particular, a modified recursive residual structure is designed to capture the multi-scale similarity of compression artifact. To effectively enhance frames with uneven noise, a multi-reconstruction structure is proposed, which outputs images with different denoise ratios and adaptively fuses them. Experimental results show that the proposed MRRN can improve coding efficiency up to 15.1% compared with the original loop filter in high-efficiency video coding. Averagely, 6.7%, 7.8%, 7.6% BD-rate reduction is achieved for all intra, low-delay P, and low-delay B, respectively. Meanwhile, as a quality enhancement method performed at encoder side, MRRN also achieves a good balance between coding performance and computational complexity compared to the state-of-the-art methods. Liangwei Yu, Liquan Shen, Hao Yang 0008, Ping An 0001 |
IEEE Signal Process. Lett. | 3 |
| 2018 | Efficient screen content intra coding based on statistical learning
Hao Yang 0008, Liquan Shen, Ping An 0001 |
Signal Process. Image Commun. | 1 |
| 2018 | Fast Intra Coding of High Dynamic Range Videos in SHVCabstractCompared with the conventional standard dynamic range (SDR) content, high dynamic range (HDR) content supplies viewers with more immersive experience by offering a much higher range of luminance. Most of current consumer devices cannot afford to this emerging technology, and content providers decide to create both an HDR version and an SDR version of the same video. In this letter, scalable high efficiency video coding (HEVC) scalable extension of HEVC (SHVC) serves as the coding framework where the base layer (BL) is an 8-b SDR version and the enhancement layer (EL) is a 12-b HDR version. Recently, many fast coding algorithms for SDR videos are proposed, and there is an urgent demand for fast coding algorithms for EL HDR videos. With the coding information of the BL SDR videos, this letter proposes a fast algorithm to reduce the complexity of intra coding for EL HDR videos. First, depth information of neighboring coding tree units (CTUs) in the HDR version and the colocated CTU in the SDR version is used for early coding unit (CU) depth determination. Moreover, four classifiers are trained to predict the CTU depth range. Two classifiers are trained for CTUs in frames with a high average luma, and another two classifiers are used for CTUs in frames with a low average luma. Experimental results show that the proposed algorithm achieves 43% encoding time saving on average, with only a 0.54% Bjøntegaard delta bit rate (BDBR) increase compared to the original SHVC test model. Guoliang Fu, Liquan Shen, Hao Yang 0008, Xiangyu Hu 0003, Ping An 0001 |
IEEE Signal Process. Lett. | 3 |
| 2017 | An efficient intra coding algorithm based on statistical learning for screen content codingabstractScreen content has different characteristics compared with natural content captured by cameras. To achieve more efficient compression, some new coding tools have been developed in the High Efficiency Video Coding (HEVC) Screen Content Coding (SCC) Extension, which also increase the computational complexity of encoder. In this paper, complexity analysis are first conducted to explore the distribution of complexities. Then, two classification trees, including early coding units (CU) partition tree (EPT) and CU content classification tree (CCT), are designed based on statistical characteristics and coding information. EPT is used to decide whether the CU skip the mode decision process of current depth level and CCT is used to classify the blocks into either natural blocks or screen blocks. Natural blocks will skip screen coding modes and screen blocks skip normal intra modes. Experimental results show the proposed algorithm can save 49% encoding time with 2.7% BD-rate increase on average for All Intra configuration under the SCC common test condition. Hao Yang 0008, Liquan Shen, Ping An 0001 |
ICIP | 1 |