Guoqing Xiang

dblp:152/7961 · DBLP profile ↗
← Back
28ranked-venue papers
4as first author
16since 2021 · last 2025
0000-0001-5831-1897ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 12 since 2021Systems, architecture and hardware · 7 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Decouple Distortion from Perception: Region Adaptive Diffusion for Extreme-low Bitrate Perception Image Compression
abstract
Leveraging the generative power of diffusion models, generative image compression has achieved impressive perceptual fidelity even at extremely low bitrates. However, current methods often neglect the non-uniform complexity of images, limiting their ability to balance global perceptual quality with local texture consistency and to allocate coding resources efficiently. To address this, we introduce the Map-guided Masking Realism Image Diffusion Codec (MRIDC), designed to optimize the trade- off between local distortion and global perceptual quality in extreme-low bitrate compression. MRIDC integrates a vector-quantized image encoder with a diffusion-based decoder. On the encoding side, we propose a Map-guided Latent Masking (MLM) module, which selectively masks elements in the latent space based on prior information, allowing adaptive resource allocation aligned with image complexity. On the decoding side, masked latents are completed using the Bidirectional Prediction Controllable Generation (BPCG) module, which guides the constrained generation process within the diffusion model to reconstruct the image. Experimental results show that MRIDC achieves state-of-the-art perceptual compression quality at extremely low bitrates, effectively preserving feature consistency in key regions and advancing the rate-distortion-perception performance curve, establishing new benchmarks in balancing compression efficiency with visual fidelity. Our code can be found at https://github.com/xjc97/mridc.
Jinchang Xu, Zhe Li 0081, Peidong Jia, Guoqing Xiang, Zhijian Hao, Shanghang Zhang
CVPR7
2025 Three-Stage Progressive Pre-Analysis Framework for VMAF Controllable Image Coding
abstract
To achieve controllable subjective quality in image coding, this paper proposes a Video Multi-method Assessment Fusion (VMAF)-oriented image coding pre-analysis algorithm, enabling the adaptive derivation of quantization parameters corresponding to a specified quality target. First, a$Q-\mathcal{V}$model is constructed to describe the relationship between encoding quantization and VMAF distortion. Then, a Three-stage Progressive Control (TPC) algorithm, shown in Fig. 1(a), is designed to adapt quantization parameters using the discovered$Q-\mathcal{V}$model. The first two stages, based on lightweight feature extraction, iteratively fit the distortion metrics intrinsically calculated by VMAF to predict the VMAF value for a given sample under specified distortion conditions. The final stage fits the$Q-\mathcal{V}$model parameters using multi-point VMAF distortion data and outputs the corresponding quantization step for encoder control. A two-pass refinement algorithm, depicted in Fig. 1(b), further adjusts the quantization parameters based on the first encoding pass, improving quality control accuracy and framework robustness. Experiments on four datasets show that the quality control error remains below 1.293% for various VMAF targets, and the two-pass refinement reduces it further to 0.710%, outperforming existing methods.
Guoqing Xiang, Wenzhao Li, Mingyuan Yang, Fan Yang 0053, Shanghang Zhang, Huizhu Jia
DCC2
2025 Efficient Quality Controllable Neural Image Compression based on QD-Model
abstract
Neural image compression has achieved significant advancements, consistently outperforming traditional codecs in terms of performance. However, research on quality control algorithms for neural image compression is still lacking. In this paper, we propose a framework designed to control the quality of compressed images through a one-pass pre-analysis. First, we construct a foundational relationship between the quantization factor and compression distortion, utilizing variable rate neural image compression as the basis for quality control. Second, we introduce the image Content-Compression features-based Distortion Estimation Network (C2DEN) to efficiently fit the sample-adaptive Quantization-Distortion (QD) model. Leveraging the QD model, we convert the target quality into a quantization factor to control the compression model, enabling quality-controllable compression of samples. Experimental results show that the average quality errors on four different datasets are only 0.89%, 1.79%, 1.73%, and 1.61%. Compared with existing control methods, our method reduces the algorithm time complexity by 98.58%, 98.52%, 98.85%, and 98.50% while ensuring accuracy, which further demonstrates the superiority of our method.
Guoqing Xiang, Jinchang Xu, Shanghang Zhang
ICASSP2
2025 Adaptive Semantic Compression: Compatible Bitstream for Scalable Human-Machine Perception Sample Adaption
abstract
With the development of visual analysis models, collaborative image compression for machine and human perception has brought new challenges to the optimization of algorithms. Existing optimization algorithms achieve this target through meticulously designed model structures and bitstream design. However, the difference in bitstream design makes it incompatible with trained and existing decoders, hindering its practicality. In this paper, we proposed the Adaptive Semantic Compression (ASC) framework to fine-tune pre-trained codec on individual samples to obtain scalable bitstreams in an intuitive yet effective way. First, to improve the efficiency of application in machine perception, we proposed the Latent Semantic Contraction (LSC) method to fine-tune the latent code while preserving the machine task performance of the decoded image. Second, to further optimize human perception, we proposed the Spatial-frequency Decoder Adaptation (SFDA) module. By compensating for distortion in the spatial and frequency domains, SFDA improves the humane perception quality of the reconstructed image. The bitstreams composed of LSC and SFDA can be decoded by existing decoders to reconstruct images, thus fully exploiting the performance of the existing model. We implemented our algorithm on different pre-trained compression models and verified the flexibility and compatibility on various test images. Experimental results show that the LSC module can save 24.97% to 29.10% of bitrates with machine perception performance. Furthermore, the application of SFDA brings a 3.16% gain in the BD-Rate with PSNR, up to 15.69%, compared to LSC.
Dingquan Li, Guoqing Xiang, Jinchang Xu, Shanghang Zhang
ICME3
2025 Hardware-friendly rate estimation algorithm and architecture design for AVS3
Yunyao Yan, Guoqing Xiang, Jie Chen 0001, Xiaofeng Huang, Peng Zhang 0007, Huizhu Jia
Multim. Tools Appl.2
2024 Joint Frame-Level and Block-Level Rate-Perception Optimized Preprocessing for Video Coding
Huajie Tan, Guoqing Xiang, Huizhu Jia
MMAsia2
2024 A hardware-friendly algorithm for LCU-level pipe-lined integer motion estimation
Xizhong Zhu, Guoqing Xiang, Peng Zhang 0007
Multim. Tools Appl.2
2024 Two-Stage Perceptual Quality Oriented Rate Control Algorithm for HEVC
abstract
As a practical technique in mainstream video coding applications, rate control dominates important to ensure compression quality with limited bitrates constraints. However, most rate control methods mainly focus on objective quality while ignoring the perceptual quality improvement for human eyes. In this paper, we propose a two-stage rate control algorithm to optimize the perceptual quality at the frame encoding stage and the coding tree unit (CTU) encoding stage for high efficiency video coding (HEVC), respectively. Firstly, for the frame encoding stage, with inter-frame distortion dependency consideration, a frame-level rate control method is presented by adjusting the frame-level Lagrange multiplier adaptively with a preprocessing method. Secondly, for the CTU encoding stage, we propose a saliency-based CTU-level perceptual quality rate control algorithm, which employs CTU-level saliency weight to adjust the perceptual rate-distortion (R-D) model. We conduct the CTU-level rate control by an optimized Lagrange multiplier and quantization parameter (QP) to achieve perceptual quality optimization. Extensive experimental results reveal that, compared with state-of-the-art rate control methods on HEVC, our algorithm achieves significant perceptual coding performance with improved subjective visual quality.
Yunyao Yan, Guoqing Xiang, Huizhu Jia, Jie Chen 0001, Xiaofeng Huang
ACM Trans. Multim. Comput. Commun. Appl.2
2023 A Hardware-friendly CTU-level IME Algorithm for VVC
abstract
The new coding tools improved the performance for H.266/VVC but also brought challenges for hardware integer motion estimation (IME). First, the data dependency in deriving a predicted motion vector (PMV) is more severe. Second, the overhead of IME is increased by the complex partition mechanism. The challenges are tougher for IME in coding tree unit (CTU) level pipelined encoder. In this paper, we propose a hardware-friendly CTU-level IME algorithm with three innovative designs. First, a PMV prediction is proposed to derive PMVs in advance. Second, all divided blocks are categorized into either binary/quadra tree (BTQT) or ternary tree (TT) blocks. The motion vectors (MVs) of BTQT blocks are estimated with a multi-resolution search. The MVs of TT blocks are inferred from the estimated MVs with an inference algorithm. The proposed algorithm suffers $ 1.20\%$ degradation but reduced the complexity by $ 80\%$ compared to the reference software.
Xizhong Zhu, Guoqing Xiang, Xiaofeng Huang, Yunyao Yan, Huizhu Jia
DCC2
2023 An Efficient Real-Time Hardware Architecture for Deblocking Filter in AVS3
abstract
To achieve higher video compression efficiency to cope with the demand for ultra high definition video applications, the AVS3 standard has been proposed recently. As a block partition-based coding standard, AVS3 suffers from the blocking artifact problem especially at low bitrates, which can be alleviated by deblocking filter. This paper presents an efficient hardware architecture for the deblocking filter for AVS3 with high-throughput. First, a fast and hardware-friendly algorithm is proposed to localize the blocking artifact boundaries. Then, a buffer organization method and a data caching strategy are proposed to solve the problem of data dependency among coding units. Based on the proposed optimized algorithm and caching strategy, a four-stage pipelined deblocking filter module architecture is designed. The experimental results show that the proposed architecture can achieve 4K@120fps video processing at 100MHz which is sufficient for real-time application.
Xiaofeng Huang, Guoqing Xiang, Xizhong Zhu, Jiaojiao Yang, Peng Zhang 0007, Huizhu Jia
ICME3
2023 A Hardware-efficient Unified Motion Estimation for Video Coding
abstract
Motion estimation (ME) is one of the most critical tools in video coding and consumes the majority of the encoding complexity. Three types of ME are utilized in the latest video coding standards, namely integer, fractional, and affine MEs. They are implemented as three searches for the integer motion vector (IMV), fractional motion vector (FMV), and control point motion vectors (CPMVs). Many algorithms were proposed to reduce the complexity for them individually, but the overall overhead of three searches is still challenging for hardware implementations. Therefore, we propose a hardware-efficient Unified Motion Estimation (UME) to derive three types of MVs with only one search. An IME with sub-block refinement is performed to collect extra motion information while searching for the IMV. The FMV and CPMVs are then derived from the collected information using a mixed error surface and an overdetermined system. Compared to the default ME algorithms in VVC, the time cost for ME is reduced by 41.63% with a coding loss of only 0.87% under LDB configuration. For hardware implementations, the minimum required resources and corresponding latency are significantly reduced by 75.35% and 69.17%, respectively.
Xizhong Zhu, Guoqing Xiang, Peng Zhang 0007, Huizhu Jia
ACM Multimedia2
2023 A Reconfigurable Multiple Transform Selection Architecture for VVC
abstract
Video coding plays an important role in the highly information-based world as videos contribute the largest part of network traffic. The latest video coding standard Versatile Video Coding (VVC) introduces a new transform scheme multiple transform selection (MTS), which brings considerable coding gains at the expense of high coding complexity. In this article, we propose a reconfigurable MTS architecture that supports all transform types in VVC with square and rectangular sizes ranging from$4\times $4 to 32$\times32$. Firstly, we explore the features of three types of transform matrices and extract the features that are beneficial to designing a unified architecture. Then, we present an improved calculation scheme for general transforms, where the transform matrix is decomposed into two simpler matrices to increase the similarity and decrease the complexity of matrices involved in three types of transform operations. Thanks to the improved calculated scheme, a unified shift-adder unit (SAU) is designed and highly reused by different types. Moreover, we provide a twirling two-point splicing (T2S) scheme to improve reusability and deal with issues of data mismatch when conducting discrete cosine transform (DCT)-II of different sizes. As a consequence, an architecture with constant throughput of 32 pixels/cycle is implemented and specified in Verilog HDL. The synthesis results indicate that the application specific integrated circuit (ASIC)-based and field-programmable gate array (FPGA)-based hardware architectures achieve significant advantages both in area reduction and power consumption compared to existing methods in the literature.
Zhijian Hao, Heming Sun, Guoqing Xiang, Peng Zhang 0007, Xiaoyang Zeng, Yibo Fan
IEEE Trans. Very Large Scale Integr. Syst.3
2022 Efficient Algorithm and Hardware Architecture for Rate Estimation in Mode Decision of AVS3
abstract
Towards enabling advanced video coding for emerging ap-plications, the AVS3 standard has been developed recently, achieving twice the coding efficiency of the AVS2 stan-dard through complex coding tools including advanced rate-distortion optimization (RDO) to select the best mode. The bit-rates are produced with the Advanced-Entropy-Coding (AEC) in the RDO process of AVS3. However, AEC dom-inates the time complexity of RDO and among the steps, con-text updating and interval subdivision are performed recur-sively, which is not conducive to real-time application, espe-cially for the hardware implementation. Thus this paper pro-poses an adaptive rate estimation algorithm with a piece-wise linear function that is very friendly to hardware implemen-tation to accelerate the rate estimation process in the RDO for AVS3 practical applications. The proposed architecture can meet the requirement of 4K@120fps ultra-high-definition videos at 200 MHz, whereas the BD-Rate increases only by 0.67% under the All-Intra (AI) configuration.
Yunyao Yan, Guoqing Xiang, Huizhu Jia, Yuan Li 0014, Peng Zhang 0007, Jie Chen 0001
ICME2
2022 A 3.1 Gbin/s advanced entropy coding hardware design for AVS3
abstract
AVS3 is a newly proposed video coding standard by the Audio Video coding Standard Workgroup, demonstrating higher compression efficiency than the High Efficiency Video Coding standard. Advanced entropy coding is one of the performance bottlenecks of the AVS3 standard video encoder due to the strong data dependency in its arithmetic coding process. A novel arithmetic encoding hardware structure is presented in this paper, and as we know, this is the first paper on AVS3 AEC hardware implementation. Firstly, we select and apply the typical optimization schemes adopted in HEVC context-based adaptive binary arithmetic coding designs. Secondly, Utilizing the unique characteristics of AVS3 AEC, we propose mathematical reordering, variable-clock-cycle range updating and variable-clock-cycle context modeling methods to optimize the critical path. Our design can encode 2.6457 bins per clock cycle, and the corresponding throughput is 3131 Mbin/s in Globalfoundries 28nm process. Compared with the basic anchor structure, it has obtained a performance improvement of 319% and can meet the 8k@l20fps ultra-high-definition video encoding requirements.
Yujie Cai, Xiaoyang Zeng, Yibo Fan, Peng Zhang 0007, Guoqing Xiang, Haibing Yin
ISCAS6
2022 An Area-efficient Unified Transform Architecture for VVC
abstract
The next-generation video coding standard Versatile Video Coding (VVC) adopts Multiple Transform Selection (MTS) to the transform module, improving coding efficiency at the expense of high computational complexity. Compared to High Efficiency Video Coding (HEVC), VVC supports larger sizes and extends the transform types to Discrete Cosine Transform (DCT)-II, Discrete Sine Transform (DST)-VII, and DCT-VIII. This paper presents an area-efficient unified architecture for VVC. To reduce the area consumption, we propose an optimized calculation scheme for general transformations where the transform matrix is decomposed into two simpler matrices named the Low-value matrix and the Error matrix. Based on the decomposition algorithm, Shift-Addition Units (SAUs)-based circuits are designed to conduct matrix multiplication and can be reused by three types. As a result, this unified architecture is capable of performing all types and sizes in VVC. The synthesis results indicate that this architecture achieves an area reduction of 37.9% $\sim$ 72.2% compared with related works for 32-point transforms.
Zhijian Hao, Qi Zheng 0004, Yibo Fan, Guoqing Xiang, Peng Zhang 0007, Heming Sun
ISCAS4
2021 Hardware-Friendly Coding Unit Decision Scheme for HEVC
abstract
Quad-tree based coding unit partition in High-Efficiency Video Coding (HEVC) achieved significant coding efficiency improvements, but also brought increasing computational complexity. Especially, design challenges like data dependence, large area cost, and imbalance of processing time of each coding tree unit (CTU), make it hard to achieve a real-time structure for real-time hardware encoder for all CTU sizes. To solve these problems, we proposed a hardware-friendly fast CU decision scheme with multi-stage algorithms for HEVC hardware encoder, aiming at the most complex modules: IME (Integer Motion Estimation), FME/IP (Fractional Motion Estimation/ Intra Prediction), and MD (Mode Decision). Firstly, in IME stage, a zero block detection method based fast CU and PU decision algorithm was presented. Secondly, we presented an estimated RDO (Rate-Distortion Optimization) based algorithm in the Hadamard domain for the early CU decision further in FME/IP stage. Finally, under the condition of hardware computing time limitation of several CU sizes, we proposed a computation time constraint CU fast decision algorithm for MD stage. Experiments demonstrated that, compared with the original HM13.0 implementation, the proposed scheme achieved about 53.9% encoding time saving with merely 2.3% coding performance degradation. What's more, significant area cost and data dependency have been alleviated, which will be more hardware-friendly for HEVC encoder design.
Ju Huang, Xiaofeng Huang, Guoqing Xiang, Yuan Li 0014, Huizhu Jia
ISCAS3
2020 Spatiotemporal Perception Aware Quantization Algorithm For Video Coding
abstract
Adaptive quantization (AQ) proves to be an effective tool to improve coding performance. In this paper, we propose an adaptively spatiotemporal perception aware quantization algorithm to increase subjective coding performance. First, the perceptual complexity models are conducted with spatial and temporal characteristics to measure the spatiotemporally perceptual redundancies, respectively. With the help of the models, the adaptively spatial and temporal quantization parameter (QP) offsets are then calculated for each coding tree unit (CTU), respectively. Finally, the perceptually optimal Lagrange multiplier of each CTU is determined with the spatial-temporal QP offset. Experimental results show that the proposed algorithm reduces 8.6% BD-Rate with SSIM (Structural Similarity Index Metric) in average over the AVS2 (the second generation of Audio Video Coding Standard) reference software RD17.0 in Low-Delay P (LDP) configurations. The subjective assessment proves the proposed algorithm can significantly reduce the bit rates with the same subjective quality.
Yunyao Yan, Guoqing Xiang, Yuan Li 0014, Wei Yan 0020, Yungang Bao
ICME2
2020 A Novel Quality Enhanced Low Complexity Rate Control Algorithm for HEVC
abstract
Rate control (RC) is a key technology in video coding which is mainly responsible for adapting the compressed video quality as much as possible under limited bandwidth. Typical RC methods consist of initial quantization parameters of I-frames decision and control parameters updating procedures. However, on the one hand, I-frames quantization parameter (QP) is decided by the sum of absolute transformed difference (SATD) computation, which is time consuming for low delay applications. On the other hand, the RC parameters are only updated ignoring the distortion characteristics for inter frame, which cannot achieve optimal rate distortion (RD) performance. Therefore, a novel RC method which considers distortion characteristics for model parameters updating for quality enhancement is proposed in this paper, including a low complexity based I-frame QP decision strategy of low delay applications. For which estimated distortion characteristics and the previous QPs are employed, respectively. Through the experimental results, with more accurate RC precision and negligible I-frame QP computation time, the proposed rate control scheme improves the performance gain of 2.6% bitrate savings for the whole test sequences on average in HM 16.9.
ChungWen Ku, Guoqing Xiang, Huizhu Jia, Yuehua Cui, Yuan Li 0014
VCIP3
2020 An adaptive spatio-temporal perception aware quantization algorithm for AVS2
Yunyao Yan, Guoqing Xiang, Yuan Li 0014, Huizhu Jia
J. Vis. Commun. Image Represent.2
2019 Bit Allocation based on Visual Saliency in HEVC
abstract
As one of the important part in the HEVC reference software, R-lambda model adopts mean absolute difference (MAD) for the coding unit tree (CTU) level bit allocation. However, this optimum method may neglect some important characteristics of human visual system (HVS). In this paper, we propose a novel bit allocation algorithm to process some salient visual information priority. Firstly, an improved video saliency detection algorithm is proposed, which induces temporal correlation into a 2D visual attention model. Secondly, the visual saliency based CTU level bit allocation algorithm is presented by allocating bits for CTUs with their saliency weights. What's more, with considerations of the temporal quality consistence among Saliency Areas (SAs), a window based weight smoothing model is proposed to achieve better subjective quality. Finally, several experiments are performed on the HEVC reference software, HM16.9, under the low delay P configuration, and the experimental results show that the average BD-Rate of the entire test sequences and of the SAs reduce 1.7% and 6.2%, respectively. The proposed algorithm can also improve subjective quality remarkably.
ChungWen Ku, Guoqing Xiang, Wei Yan 0020, Yuan Li 0014
VCIP2
2018 A perceptually temporal adaptive quantization algorithm for HEVC
Guoqing Xiang, Huizhu Jia, Mingyuan Yang, Xinfeng Zhang 0001, Xiaofeng Huang, Jie Liu 0035
J. Vis. Commun. Image Represent.1
2018 A novel adaptive quantization method for video coding
Guoqing Xiang, Huizhu Jia, Mingyuan Yang, Yuan Li 0014
Multim. Tools Appl.1
2017 Fast rate distortion optimized quantization method for HEVC
abstract
Rate-Distortion Optimized Quantization (RDOQ) brings significant improvement of coding performance in High Efficiency Video Coding (HEVC). However, it results in high computational complexity to determine the best quantization levels for each transform coefficient when applying rate and distortion optimization operation. In this paper, we firstly proposed an efficient way to skip the all zero coding units. Then we further propose a fast RDOQ scheme by skipping the optimization step of All Quantized Zero Blocks except DC Coefficient (AQZB-DC). Moreover, a rate difference model is established to select optimal quantization level for AQZB-DC blocks. Experiments are carried out on reference software HM 16.0 and the results show that the proposed method achieves 42.00% and 39.63% quantization time saving under Random Access (RA) and Low Delay (LD) configuration on average, while the BD-Rate loss is only 0.03% and 0.06%, respectively.
Meng Wang 0017, Hongfei Fan, Shanshe Wang, Shengfu Dong, Guoqing Xiang, Huizhu Jia
ISCAS7
2017 LLCNN: A convolutional neural network for low-light image enhancement
abstract
In this paper, we propose a CNN based method to perform low-light image enhancement. We design a special module to utilize multiscale feature maps, which can avoid gradient vanishing problem as well. In order to preserve image textures as much as possible, we use SSIM loss to train our model. The contrast of low-light images can be adaptively enhanced using our method. Results demonstrate that our CNN based method outperforms other contrast enhancement methods.
Chuang Zhu, Guoqing Xiang, Yuan Li 0014, Huizhu Jia
VCIP3
2016 Structure preserving single image super-resolution
abstract
In this paper, we present a novel structure preserving method for single image super-resolution to well construct edge structures and small detail structures. In our approach, the sharp edges are recovered via a novel edge preserving interpolation technique based on a well estimated gradient field and the edge preserving method, which incorporate the local and non-local structure information. The gradient of interpolated high-resolution(HR) image is then regarded as an edge preserving constraint to reconstruct the detail structures. Experimental results demonstrate that the new approach can reconstruct faithfully the HR images with sharp edges and texture structures, and annoying artifacts (blurring, jaggies, ringing, etc.) are greatly suppressed. It outperforms the state-of-the-art approaches, based on subjective and objective evaluations.
Fan Yang 0053, Don Xie, Huizhu Jia, Rui Chen 0006, Guoqing Xiang, Wen Gao 0001
ICIP5
2016 Adaptive perceptual preprocessing for video coding
abstract
The quantization of block DCT coefficients is too coarse, and will result in much non-uniform artifacts, therefore, blocking and ring artifacts are usually visible in reconstructed video frames especially when the bitrate is not sufficient. In this paper, we present an adaptive perceptual preprocessing (APP) method to reduce the probability of these artifacts generated in video encoding. The APP algorithm employs a novel filter based on just noticeable distortion filter (JNDF) and adaptive bilateral filter (ABF), and they can be adaptively chosen by block characteristics and quantization parameters. Experimental results demonstrate our proposed algorithm can significantly improve the subjective quality of reconstructed image.
Guoqing Xiang, Huizhu Jia, Jie Liu 0035, Binbin Cai, Yuan Li 0014
ISCAS1
2016 Hardware-oriented adaptive multi-resolution motion estimation algorithm and its VLSI architecture
abstract
In this paper, we propose a hardware architecture of an adaptive multi-resolution motion estimation algorithm (AMMEA) for high definition video encoder to reduce hardware cost. The texture-based search strategies are based on temporal stationarity and spatial homogeneity with Sobel edge operator. The proposed algorithm makes motion estimation more concise. We also propose Sobel edge operator hardware architecture. The four-pixel SAD unit which is the basic processing element (PE) in our proposed architecture is used for SAD calculation and Sobel edge operator computation. The hardware architecture achieves very high data utilization and data throughout. Using our proposed AMMEA with regular data flow, simulation results show that the proposed architecture can significantly reduce the hardware cost with a negligible PSNR loss of 0.03dB compared with the full-search. The design is implemented with SMIC 0.18μm CMOS technology and costs 950K gates count, and it supports the real-time encoding of 1080P@30fps with two reference frames under a clock frequency of 150MHz.
Guoqing Xiang, Huizhu Jia, Jie Liu 0035, Yuan Li 0014
ISCAS1
2015 An adaptive inter CU depth decision algorithm for HEVC
abstract
The emerging High-Efficiency Video Coding (HEVC) standard has introduced a number of new coding tools, such as a quad-tree based coding unit (CU). The quadtree-structured coding unit achieves significant coding efficiency improvements compared to H264/AVC. However, the complexity of CU depth decision associated with Rate-Distortion (R-D) cost computation dramatically increased. In order to alleviate the computational burden in HEVC inter coding, a fast CU depth decision algorithm is proposed in this paper. Firstly, zero CU detection method for HEVC is proposed as early termination algorithm. Secondly, the CU depth pruning strategies are adaptively determined according to standard deviation of statistic spatiotemporal depth information. Finally, when the neighbors are not available or have a very weak correlation, edge gradient of current coding tree unit (CTU) is considered as main factor for CU depth pruning method. Experimental results demonstrate that, compared with the original HM16.0 implementation, the proposed algorithm achieves about 40.5% encoding time saving with ignorable coding performance degradation.
Jie Liu 0035, Huizhu Jia, Guoqing Xiang, Xiaofeng Huang, Binbin Cai, Chuang Zhu, Don Xie
VCIP3