VLDB 2026 Research / reviewers in the wild / expert
Changsheng Gao
dblp:133/3541
· DBLP profile ↗
24ranked-venue papers
9as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 9 first-author · 21 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Transform-Free Feature Coding via Entropy-Constrained Vector QuantizationabstractFeature coding has recently emerged as a key technique for efficient transmission of intermediate representations in distributed AI systems. Existing approaches largely follow a transform-based pipeline inherited from image and video coding, where the transform module is used to remove spatial structural redundancies in visual signals. However, our analysis indicates that such redundancies have already been largely removed during feature extraction, which reduces the necessity of the transform module. Building on this insight, we propose a new transform-free pipeline that directly encodes the extracted features via a vector quantization module and an entropy model. The proposed transform‑free framework jointly learns the quantization codebook and entropy model, enabling end‑to‑end optimization tailored to the inherent feature characteristics. Furthermore, the proposed method inherently avoids the computational complexity of the transform module. Experiments on features from diverse architectures and tasks demonstrate that our method achieves superior rate-distortion performance compared to transform-based baselines, while significantly reducing the encoding and decoding complexity. Qiaoxi Chen, Changsheng Gao, Li Li 0040, Dong Liu 0002 |
AAAI | 2 |
| 2026 | Topic-Guided and Context-Accumulated Conditions for Prompt Compression
Yenan Xu, Changsheng Gao, Li Li 0040, Dong Liu 0002 |
ISCAS | 2 |
| 2026 | Temporal Quality Aggregation for VQA: Benchmark and Psychology-Inspired Model
Baoliang Chen, Changsheng Gao, Lingyu Zhu 0006, Liang Xie 0013, Hanwei Zhu, Zhijian Hao |
QoMEX | 2 |
| 2026 | Token-Wise Attention-Guided Semantic Quality Assessment for Compressed Visual Features
Shien Ke, Changsheng Gao, Hadi Amirpour, Zhihua Wang 0002, Xiaoyan Sun 0001 |
QoMEX | 2 |
| 2026 | QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos: Methods and Results
Yixu Chen, Hai Wei, Pierre R. Lebreton, Patrick Le Callet, Alexander Kopte, Amritha Premkumar, Anna Meyer, Baojun Li, Changsheng Gao, Christian Herglotz, Christian Timmerer, Dandan Zhu 0001, Diwakara Reddy, Dong Liu 0002, Dounia Hammou, Guangtao Zhai, Hadi Amirpour, Hao Cheng 0015, Hichem Faraoun, Jonas Janzen, Krishna Srikar Durbha, Li Li 0040, Marc Windsheimer, MohammadAli Hamidi, Mykyta Skipenko, Paul Wawerek-Lopez, Pragyadipta Adhya, Prajit T. Rajendran, Rafal Mantiuk, Shien Ke, Sid Ahmed Fezza, Simon Deniffel, Wei Sun 0029, Weixia Zhang, Xiangguang Chen, Zuowei Cao, Minhao Tang, Xiaoyan Sun 0001, Xingwei Liu, Yeganeh Chatri, Yenan Xu |
QoMEX | 10 |
| 2026 | USTC-TD: A Test Dataset and Benchmark for Image and Video Coding in 2020sabstractImage/video coding has been a remarkable research area for both academia and industry for many years. Testing datasets, especially high-quality image/video datasets, are desirable for the justified evaluation of coding-related research, practical applications, and standardization activities. We put forward a test dataset, namely USTC-TD, which has been successfully adopted in the practical end-to-end image/video coding challenge ofIEEE International Conference on Visual Communications and Image Processing (VCIP)in 2022 and 2023. USTC-TD contains 40 images at 4K spatial resolution and 10 video sequences at 1080p spatial resolution, featuring various content due to the diverse environmental factors (e.g., scene type, texture, motion, view) and the designed imaging factors (e.g., illumination, lens, shadow). We quantitatively evaluate USTC-TD on different image/video features (spatial, temporal, color, lightness), and compare it with the previous image/video test datasets, which verifies its excellent compensation for the shortcomings of existing datasets. We also evaluate both classic standardized and recently learned image/video coding schemes on USTC-TD using objective quality metrics (PSNR, MS-SSIM, VMAF) and subjective quality metric (MOS), providing an extensive benchmark for these evaluated schemes. Based on the characteristics and specific design of the proposed test dataset, we analyze the benchmark performance and shed light on the future research and development of image/video coding. All the data are released online:https://esakak.github.io/USTC-TD. Zhuoyuan Li 0001, Junqi Liao, Chuanbo Tang, Haotian Zhang 0009, Yifan Bian, Xihua Sheng, Xinmin Feng, Yao Li 0016, Changsheng Gao, Li Li 0040, Dong Liu 0002, Feng Wu 0005 |
IEEE Trans. Multim. | 10 |
| 2025 | Feature Coding in the Era of Large Models: Dataset, Test Conditions, and BenchmarkabstractLarge models have achieved remarkable performance across various tasks, yet they incur significant computational costs and privacy concerns during both training and inference. Distributed deployment has emerged as a potential solution, but it necessitates the exchange of intermediate information between model segments, with feature representations serving as crucial information carriers. To optimize information exchange, feature coding is required to reduce transmission and storage overhead. Despite its importance, feature coding for large models remains an under-explored area. In this paper, we draw attention to large model feature coding and make three fundamental contributions. First, we introduce a comprehensive dataset encompassing diverse features generated by three representative types of large models. Second, we establish unified test conditions, enabling standardized evaluation pipelines and fair comparisons across future feature coding studies. Third, we introduce two baseline methods derived from widely used image coding techniques and benchmark their performance on the proposed dataset. These contributions aim to provide a foundation for future research and inspire broader engagement in this field. To support a long-term study, all source code and the dataset are made available at \href{https://github.com/chansongoal/LaMoFC}{https://github.com/chansongoal/LaMoFC}. Changsheng Gao, Qiaoxi Chen, Yenan Xu, Dong Liu 0002, Weisi Lin |
ICCV | 1 |
| 2025 | Rethinking Joint Optimization in Feature Compression: Insights from Person Re-IdentificationabstractJoint optimization, which jointly optimizes compression and machine vision algorithms, is widely regarded as an effective strategy for enhancing compression performance in the field of coding for machines. However, existing joint optimization methods usually incorporate a semantics parsing module at the end of the pipeline, raising a critical question: Does the performance improvement stem from the joint optimization itself, or is it primarily driven by the tailed semantics parsing module? To address this, we disentangle the tailed semantics parsing module from the joint optimization pipeline by leveraging the simplicity of the person re-identification task, where semantics parsing involves deterministic feature matching rather than a learned neural network. First, we propose a separate optimization pipeline and two joint optimization pipelines to systematically investigate the effectiveness of joint optimization. Our findings reveal that joint optimization alone does not necessarily guarantee performance improvement. Second, we evaluate the influence of the tailed semantics parsing module by equipping it with varying capabilities, demonstrating that higher parsing capability directly correlates with better machine vision performance. These findings underscore the pivotal role of tailed semantics parsing in enhancing machine vision performance and challenge the assumption that joint optimization alone drives improvement. This work offers new insights for designing effective coding methods, emphasizing the interplay between optimization strategies and tailed semantics parsing. Changsheng Gao, Zhuoyuan Li 0001, Li Li 0040, Dong Liu 0002, Feng Wu 0001, Weisi Lin |
ICME | 1 |
| 2025 | Compressed Feature Quality Assessment: Dataset and BaselinesabstractThe widespread deployment of large models in resource-constrained environments has underscored the need for efficient transmission of intermediate feature representations. In this context, feature coding, which compresses features into compact bitstreams, becomes a critical component for scenarios involving feature transmission, storage, and reuse. However, this compression process inevitably introduces semantic degradation that is difficult to quantify with traditional metrics. To address this, we formalize the research problem of Compressed Feature Quality Assessment (CFQA), aiming to evaluate the semantic fidelity of compressed features. To advance CFQA research, we propose the first benchmark dataset, comprising 300 original features and 12000 compressed features derived from three vision tasks and four feature codecs. Task-specific performance degradation is provided as true semantic distortion for evaluating CFQA metrics. We systematically assess three widely used metrics -- MSE, cosine similarity, and Centered Kernel Alignment (CKA) -- in terms of their ability to capture semantic degradation. Our findings demonstrate the representativeness of the proposed dataset while underscoring the need for more sophisticated metrics capable of measuring semantic distortion in compressed features. This work advances the field by establishing a foundational benchmark and providing a critical resource for the community to explore CFQA. To foster further research, we release the dataset and all associated source code at https://github.com/chansongoal/Compressed-Feature-Quality-Assessment. Changsheng Gao, Wei Zhou 0021, Guosheng Lin, Weisi Lin |
ACM Multimedia | 1 |
| 2025 | DT-UFC: Universal Large Model Feature Coding via Peaky-to-Balanced Distribution TransformationabstractLike image coding in visual data transmission, feature coding is essential for the distributed deployment of large models by significantly reducing transmission and storage burden. However, prior studies have mostly targeted task- or model-specific scenarios, leaving the challenge of universal feature coding across diverse large models largely unexplored. In this paper, we present the first systematic study on universal feature coding for large models. The key challenge lies in the inherently diverse and distributionally incompatible nature of features extracted from different models. For example, features from DINOv2 exhibit highly peaky, concentrated distributions, while those from Stable Diffusion 3 (SD3) are more dispersed and uniform. This distributional heterogeneity severely hampers both compression efficiency and cross-model generalization. To address this, we propose a learned peaky-to-balanced distribution transformation, which reshapes highly skewed feature distributions into a common, balanced target space. This transformation is non-uniform, data-driven, and plug-and-play, enabling effective alignment of heterogeneous distributions without modifying downstream codecs. With this alignment, a universal codec trained on the balanced target distribution can effectively generalize to features from different models and tasks. We validate our approach on three representative large models (LLaMA3, DINOv2, and SD3) across multiple tasks and modalities. Extensive experiments show that our method achieves notable improvements in both compression efficiency and cross-model generalization over task-specific baselines. All source code has been made available at https://github.com/chansongoal/DT-UFC. Changsheng Gao, Li Li 0040, Dong Liu 0002, Xiaoyan Sun 0001, Weisi Lin |
ACM Multimedia | 1 |
| 2025 | IMOFC: Identity-Level Metric Optimized Feature Compression for Identification TasksabstractFeature compression has attracted much attention in recent years due to its promising applications in scenarios where features are transmitted and analyzed by machine vision. However, existing research mainly focuses on coarse-grained features extracted from recognition tasks such as classification and detection, neglecting fine-grained features extracted from identification tasks. In this paper, we make a pioneering attempt to study fine-grained feature compression in the context of identification tasks. Our main focus is on the distortion metric, given its critical importance in optimizing the performance of a compression network. We initiate our discussion by reviewing the instance-level metrics in existing literature, highlighting their oversight of the inter-feature relationships. The inter-feature relationships are especially important for identification tasks as they involve similarity comparison among different identities. To address this problem, we propose to consider inter-feature relationships from the perspective of identity information. Specifically, we propose an identity-level metric to incorporate both intra-identity similarity and inter-identity discriminability. The intra-identity similarity constraint aims to cluster features from the same identity, while the inter-identity discriminability constraint ensures that features from different identities deviate from each other. We implement the identity-level metric on four different feature compression networks designed based on feature characteristics. Experimental results show the effectiveness of the proposed identity-level metric on person re-identification and face verification tasks. Changsheng Gao, Yiheng Jiang, Li Li 0040, Dong Liu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | End-to-End Learned Scalable Multilayer Feature Compression for Machine Vision TasksabstractWe propose an end-to-end learned scalable multilayer feature compression method. Our proposed method is illustrated in Figure 1 , where f 1 ,… ,f n stand for deep features at different layers. The deep feature f n denoting base layer is first transformed and quantized into the latent ${\hat y_n}$ . The latent ${\hat y_n}$ is then inversely transformed to reconstruct the feature as f ˆ n . In addition, the latent ${\hat y_n}$ is also fed into the entropy model of the previous-layer feature f n−1 as conditional information for the enhancement layer. The entropy model of the feature f n−1 takes both ${\hat y_{n - 1}}$ and ${\hat y_n}$ as inputs to improve compression efficiency. Qiaoxi Chen, Changsheng Gao, Dong Liu 0002 |
DCC | 2 |
| 2024 | Rethinking the Joint Optimization in Video Coding for Machines: A Case StudyabstractIn this work, we investigate the joint optimization strategy in the scenario of video coding for machines (VCM). We formulated two kinds of joint optimization strategies, Opt_JA and Opt_JH , and compared them with the separate optimization strategy Opt_S. The three optimization strategies are illustrated in Fig. 1 . In Opt_S , we separately train the feature compression network with mean squared error (MSE). In Opt_JA , we optimize all modules jointly toward the person re-identification task. In Opt_JH , only the aggregation module and feature compression module are jointly optimized. The feature compression consists of two fully-connected (FC) layers and two batch normalization (BN) layers. Specifically, we set five compression ratios (CR): 256, 128, 64, 32, and 16. Changsheng Gao, Zhuoyuan Li 0001, Li Li 0040, Dong Liu 0002, Feng Wu 0001 |
DCC | 1 |
| 2024 | End-to-End Learned Scalable Multilayer Feature Compression For Machine Vision TasksabstractIn the field of Video Coding for Machines (VCM), scalable feature compression has attracted attention for its potential to support a variety of machine vision tasks. However, the existing scalable feature compression methods exhibit limited performance. To address this problem, we propose an end-to-end learned scalable multilayer feature compression method in this paper. First, we propose to leverage an end-to-end feature compression method, which can efficiently exploit redundancy among features through a learning approach, to improve compression efficiency. Second, we introduce a novel strategy involving the use of the transformed latent of the base layer as the conditional information for the enhancement layer. Given the learnable nature of our compression method, we propose to optimize the base layer and the enhancement layer jointly. The joint optimization encourages the base layer to produce more suitable conditional information for the enhancement layer. Comparative experiments against existing feature compression and image compression methods verify our approach’s remarkable performance improvements. Qiaoxi Chen, Changsheng Gao |
ICIP | 2 |
| 2024 | DMOFC: Discrimination Metric-Optimized Feature CompressionabstractFeature compression, as an important branch of video coding for machines (VCM), has attracted significant attention and exploration. However, the existing methods mainly focus on intra-feature similarity, such as the Mean Squared Error (MSE) between the reconstructed and original features, while neglecting the importance of inter-feature relationships. In this paper, we analyze the inter-feature relationships, focusing on feature discriminability in machine vision and underscoring its significance in feature compression. To maintain the feature discriminability of reconstructed features, we introduce a dis-crimination metric for feature compression. The discrimination metric is designed to ensure that the distance between features of the same category is smaller than the distance between features of different categories. Furthermore, we explore the relationship between the discrimination metric and the discriminability of the original features. Experimental results confirm the effectiveness of the proposed discrimination metric and reveal there exists a tradeoff between the discrimination metric and the discrim-inability of the original features. Changsheng Gao, Yiheng Jiang, Li Li 0040, Dong Liu 0002 |
PCS | 1 |
| 2024 | Feature Compression With 3D Sparse ConvolutionabstractFeature compression is an important branch of video coding for machines (VCM). While existing methods draw inspiration from image compression, they have not fully utilized the unique characteristics of features. In this paper, we investigate feature characteristics in two key aspects: dimensionality and sparsity. Our analysis reveals that the low spatial dimensionality and high channel dimensionality of features make traditional 2D convolution-based methods, which usually downsample along spatial dimensions while increasing channels, unsuitable for feature compression. To address this, we propose compressing features using 3D convolution. Additionally, considering the sparsity characteristic, we propose applying sparse convolution to reduce model complexity. To thoroughly investigate the proposed 3D sparse convolution-based method, we verify it with various network structures and input features. Experimental results demonstrate the superiority of our proposed method over traditional 2D convolution-based approaches, highlighting its potential for effective feature compression. Changsheng Gao, Qiaoxi Chen, Li Li 0040, Dong Liu 0002, Xiaoyan Sun 0001 |
VCIP | 2 |
| 2024 | Perceptual Image Compression With Conditional Diffusion TransformersabstractGenerative models have significantly advanced generative AI, particularly in image and video generation. Recognizing their potential, researchers have begun exploring their application in image compression. However, existing methods face two primary challenges: limited performance improvement and high model complexity. In this paper, to address these two challenges, we propose a perceptual image compression solution by introducing a conditional diffusion model. Given that compression performance heavily depends on the decoder’s generative capability, we base our decoder on the diffusion transformer architecture. To address the model complexity problem, we implement the diffusion transformer architecture with Swin transformer. Equipped with enhanced generative capability, we further augment the decoder with informative features using a multi-scale feature fusion module. Experimental results demonstrate that our approach surpasses existing perceptual image compression methods while achieving lower model complexity. Rui Mao 0018, Xinmin Feng, Changsheng Gao, Li Li 0040, Dong Liu 0002, Xiaoyan Sun 0001 |
VCIP | 3 |
| 2024 | Learned Rate-Distortion Cost Prediction for Ultrafast Screen Content Intra CodingabstractAs online collaborations become more prevalent, screen content has become increasingly important in real-time video communications. To reduce communication costs, the H.265/HEVC standard introduced the Screen Content Coding (SCC) extension, which achieves significant bits savings but comes with a higher encoding complexity. There is a need for ultrafast SCC encoding to meet the demands of real-time applications. Our key idea is to predict the rate-distortion (RD) cost of each possible coding unit under each possible mode, rather than performing actual coding to obtain the RD cost. Specifically, we construct neural networks to predict RD costs for intra prediction, palette, and normal intra block copy (IBC) modes. For IBC merge mode, we conduct motion compensation trials and use a linear regression network for prediction. Using the predicted RD costs, we create a partition-mode map set that determines not only block partitioning but also optimal modes, significantly reducing encoding complexity. Our experimental results demonstrate that our method achieves a more than 90% reduction in encoding time with an average 9.4% BD-rate increase compared to the HEVC-SCC reference software in the all-intra configuration. Yanchen Zuo, Changsheng Gao, Dong Liu 0002, Li Li 0040, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Research on Electromagnetic Effect Generated by DC Converter on Human Body in Electric VehicleabstractThe field distribution in an electric vehicle is important characteristics to arrange the sensitive equipment. The high-power DC converter is one of the main interference sources. Power cables connecting to DC converter radiate the interference into the space inside vehicle. This paper builds DC converter system model, shielded cable model, whole electric vehicle model and human model respectively to analyze the field distribution and the specific absorption rate distribution inside electric vehicle. Jianjun Xiao 0002, Changsheng Gao, Zhichun Li, Dan Zhang 0019 |
VTC2023-Spring | 2 |
| 2023 | Towards Task-Generic Image Compression: A Study of Semantics-Oriented MetricsabstractInstead of being observed by human, multimedia data are now more and more fed into machines to perform different kinds of semantic analysis. One image may be analyzed multiple times by different machine vision algorithms for different purposes. While machine vision-oriented image compression has been studied, the existing methods are usually driven by a specific machine vision task, and may not be applicable for the other tasks. We address thetask-genericimage compression, in the hope that an image is compressed once but used multiple times for different tasks, all with satisfactory performance. Our study is based on the end-to-end learned image compression. We focus ourselves on the distortion metric, i.e., finding out a task-agnostic metric to estimate the quality of reconstructed images. On the one hand, we study deep feature distance as the metric, which transforms images into a latent space by a pretrained convolutional network—the latent space is believed to be more aligned to semantics—and calculates distance in the latent space. On the other hand, inspired by the saliency mechanism, we study an importance-weighted pixel distance as the metric, where the weights are generated to reflect the importance of the pixels to semantics. Moreover, we combine the two distances into one metric to investigate their complementary nature. An extensive set of experiments are performed to evaluate these metrics. Experimental results show that using the combined metric performs the best, and leads to 20.79%$\sim$42.69% bits saving under the same semantic analysis performance, compared to using signal fidelity metrics. Interestingly, we observe that using the combined metric also improves the visual quality of the reconstructed images. Changsheng Gao, Dong Liu 0002, Li Li 0040, Feng Wu 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Two-Step Fast Mode Decision for Intra Coding of Screen ContentabstractWith the rapid development of screen content video applications, screen content coding (SCC) is urgently needed to be used in commercial codecs. However, the extra encoding complexity introduced by the new SCC tools has posed a great challenge for its practical deployment. In this paper, motivated by our observations that there should be a fine-grained mapping between image content and candidate modes, we propose a two-step fast mode decision method to reduce the encoding complexity. First, we propose to use a convolution neural network (CNN) to automatically extract useful features for fine-grained content classification. Second, we build a precise and concise mapping from CUs to candidate modes by simultaneously considering CU content type, CU size, and mode complexity. Note that the spatial correlations between neighboring CUs and current CU are also utilized in candidate modes derivation. In addition to the two-step fast mode decision method, a content-aware early termination algorithm is further proposed to reduce the encoding complexity. Extensive experiments demonstrate that our method achieves better performance compared with state-of-the-art ones, with 50.13% total encoding complexity reduction and only 0.92% BD-rate increase. Changsheng Gao, Li Li 0040, Dong Liu 0002, Zhibo Chen 0001, Weiping Li 0003, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Cnn-Based Depth Map Prediction for Fast Block Partitioning in HEVC Intra CodingabstractHigh Efficiency Video Coding (HEVC) achieves significant improvement in compression efficiency by introducing quadtree-based block partition. However, in the HEVC reference software–HM, the optimal partition is found by a recursive rate-distortion optimization (RDO) process, which is computationally expensive and not friendly to hardware implementation. We propose a fast block partitioning algorithm using convolutional neural network (CNN) based depth map prediction for HEVC intra coding. We use the depth map to represent the block partition of a coding tree unit (CTU). Then, we design a CNN to predict the depth map for a CTU, and we construct a large-scale dataset to train the CNN. Through the depth map prediction, we obtain a block partitioning structure for the entire CTU, and then we could directly compress each coding unit, getting rid of the recursive RDO process for partitioning. Experimental results show that our proposed method reduces 65.55% encoding time of HM at the cost of 2.02% Bjøntegaard Delta rate (BD-rate) increase on the common test sequences. For 4K sequences, our method achieves 76.97% time saving with 2.89% BD-rate increase. Aolin Feng, Changsheng Gao, Li Li 0040, Dong Liu 0002, Feng Wu 0001 |
ICME | 2 |
| 2021 | SSSIC: Semantics-to-Signal Scalable Image Coding With Learned Structural RepresentationsabstractWe address the requirement of image coding for joint human-machine vision, i.e., the decoded image serves both human observation and machine analysis/understanding. Previously, human vision and machine vision have been extensively studied by image (signal) compression and (image) feature compression, respectively. Recently, for joint human-machine vision, several studies have been devoted to joint compression of images and features, but the correlation between images and features is still unclear. We identify the deep network as a powerful toolkit for generating structural image representations. From the perspective of information theory, the deep features of an image naturally form an entropy decreasing series: a scalable bitstream is achieved by compressing the features backward from a deeper layer to a shallower layer until culminating with the image signal. Moreover, we can obtain learned representations by training the deep network for a given semantic analysis task or multiple tasks and acquire deep features that are related to semantics. With the learned structural representations, we propose SSSIC, a framework to obtain an embedded bitstream that can be either partially decoded for semantic analysis or fully decoded for human vision. We implement an exemplar SSSIC scheme using coarse-to-fine image classification as the driven semantic analysis task. We also extend the scheme for object detection and instance segmentation tasks. The experimental results demonstrate the effectiveness of the proposed SSSIC framework and establish that the exemplar scheme achieves higher compression efficiency than separate compression of images and features. Ning Yan 0001, Changsheng Gao, Dong Liu 0002, Houqiang Li, Li Li 0040, Feng Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Asymmetric-Kernel CNN Based Fast CTU Partition for HEVC Intra CodingabstractHigh Efficiency Video Coding (HEVC) has higher encoding complexity due to sophisticated coding tree unit (CTU) partition with recursive rate-distortion optimization (RDO) procedures. In this paper, we propose a specified Asymmetric-Kernel CNN (AK-CNN) for fast CTU and PU (prediction unit) partition prediction. Shallow network structures with asymmetric horizontal and vertical convolution kernels are designed to precisely extract the texture features of each block with much lower complexity. We establish our own dataset with complete CTU partition patterns together with their RD-cost for network training. The confidence threshold decision scheme is designed in the PU partition part to achieve the best trade-off between the coding performance and complexity reduction. Experimental results demonstrate that our approach achieves 69.8% intra mode encoding complexity reduction with negligible rate-distortion performance degradation, superior to the existing fast partition algorithms. Jun Shi 0004, Changsheng Gao, Zhibo Chen 0001 |
ISCAS | 2 |