Zhenyu Liu 0001

dblp:74/4038-1 · DBLP profile ↗
← Back
49ranked-venue papers
19as first author
0since 2021 · last 2019
0000-0002-8816-766XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 15 first-authorSystems, architecture and hardware · 13 · 4 first-authorArtificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Hardware accelerators and domain-specific architectures · 50% Performance modeling and evaluation · 20% Electronic design automation · 15%
Computer graphics and multimedia
3 papers
Image and video coding · 100%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
video coding accelerator
0.522016
A High-Throughput and Multi-Parallel VLSI Architecture for HEVC Deblocking Filter · IEEE Trans. Multim. 2016
CU Partition Mode Decision for HEVC Hardwired Intra Encoder Using Convolution Neural Network · IEEE Trans. Image Process. 2016
Image and video coding
rate control
0.412019
Optimize x265 Rate Control: An Exploration of Lookahead in Frame Bit Allocation and Slice Type Decision · IEEE Trans. Image Process. 2019
Image and video coding
rate-distortion optimization
0.412019
Optimize x265 Rate Control: An Exploration of Lookahead in Frame Bit Allocation and Slice Type Decision · IEEE Trans. Image Process. 2019
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural network accelerator design
0.312018
Computation Error Analysis of Block Floating Point Arithmetic Oriented Convolution Neural Network Accelerator Design · AAAI 2018
Performance modeling and evaluation › numerical algorithms
numerical error analysis
0.312018
Computation Error Analysis of Block Floating Point Arithmetic Oriented Convolution Neural Network Accelerator Design · AAAI 2018
Image and video coding › video compression
coding unit partitioning
0.212016
CU Partition Mode Decision for HEVC Hardwired Intra Encoder Using Convolution Neural Network · IEEE Trans. Image Process. 2016
Image and video coding › video compression
intra prediction
0.212016
CU Partition Mode Decision for HEVC Hardwired Intra Encoder Using Convolution Neural Network · IEEE Trans. Image Process. 2016
Integrated circuit design
digital circuit design
0.212016
A High-Throughput and Multi-Parallel VLSI Architecture for HEVC Deblocking Filter · IEEE Trans. Multim. 2016
Electronic design automation
hardware/software co-design
0.212016
CU Partition Mode Decision for HEVC Hardwired Intra Encoder Using Convolution Neural Network · IEEE Trans. Image Process. 2016
Image and video coding › video compression › video codec
HEVC
0.112016
A High-Throughput and Multi-Parallel VLSI Architecture for HEVC Deblocking Filter · IEEE Trans. Multim. 2016
Image and video coding
video coding standards
0.112016
A High-Throughput and Multi-Parallel VLSI Architecture for HEVC Deblocking Filter · IEEE Trans. Multim. 2016

Methods — techniques the papers use, named apart from their topics

rate-distortion optimization · 0.5ping-pong buffer design · 0.5parallel filtering · 0.5convolutional neural network · 0.5VLSI accelerator design · 0.5quantization parameter offset · 0.4lookahead complexity estimation · 0.4noise-to-signal ratio analysis · 0.3block floating point arithmetic · 0.3
YearPublicationVenuePosition
2019 Optimize x265 Rate Control: An Exploration of Lookahead in Frame Bit Allocation and Slice Type Decision
abstract
To improve the Rate-Distortion (R-D) quality, x265 rate-control made a variety vital decisions, such as scene cut detection, slice type decision, and coding-unit quantization parameter (QP) offsets, leveraging on lookahead to evaluate the information propagation through the current and the near future consecutive frames. However, as the frame base QP that dominates the bit amount allocated to one frame was only determined by the long-term complexity history in the original algorithm, the frame bit allocation became insensitive to the recent scene changes with the growth of coding time. In addition, the specified threshold in slice type decision, which was compared with the estimated frame coding costs to detect the B-type slice, did not consider the impacts of quantization. The aforementioned irrational elements degraded the rate accuracy and the R-D performance of x265. In this paper, the frame base QP is determined with not only the coding complexity history, but also the complexity changes and the data dependencies between the current and the near future pictures by exploring lookahead. Moreover, the quantization scale is introduced to the threshold specification in slice type decision, which identifies more pictures as B-type properly when increasing QP. The proposed algorithms were conducted in x265 version 2.4. Experiments revealed that, under the default preset (-preset medium), 0.617dB and up to 1.705dB BDPSNR quality gains were achieved, while saving the encoding time by 1.13% and improving the rate accuracy by 4.2% on average.
Zhenyu Liu 0001, Xiangyang Ji
IEEE Trans. Image Process.1
2019 High-Performance FPGA-Based CNN Accelerator With Block-Floating-Point Arithmetic
abstract
Convolutional neural networks (CNNs) are widely used and have achieved great success in computer vision and speech processing applications. However, deploying the large-scale CNN model in the embedded system is subject to the constraints of computation and memory. An optimized block-floating-point (BFP) arithmetic is adopted in our accelerator for efficient inference of deep neural networks in this paper. The feature maps and model parameters are represented in 16-bit and 8-bit formats, respectively, in the off-chip memory, which can reduce memory and off-chip bandwidth requirements by 50% and 75% compared to the 32-bit FP counterpart. The proposed 8-bit BFP arithmetic with optimized rounding and shifting-operation-based quantization schemes improves the energy and hardware efficiency by three times. One CNN model can be deployed in our accelerator without retraining at the cost of an accuracy loss of not more than 0.12%. The proposed reconfigurable accelerator with three parallelism dimensions, ping-pong off-chip DDR3 memory access, and an optimized on-chip buffer group is implemented on the Xilinx VC709 evaluation board. Our accelerator achieves a performance of 760.83 GOP/s and 82.88 GOP/s/W under a 200-MHz working frequency, significantly outperforming previous accelerators.
Xiaocong Lian, Zhenyu Liu 0001, Zhourui Song, Jiwu Dai, Wei Zhou 0020, Xiangyang Ji
IEEE Trans. Very Large Scale Integr. Syst.2
2018 Computation Error Analysis of Block Floating Point Arithmetic Oriented Convolution Neural Network Accelerator Design
abstract
The heavy burdens of computation and off-chip traffic impede deploying the large scale convolution neural network on embedded platforms. As CNN is attributed to the strong endurance to computation errors, employing block floating point (BFP) arithmetics in CNN accelerators could save the hardware cost and data traffics efficiently, while maintaining the classification accuracy. In this paper, we verify the effects of word width definitions in BFP to the CNN performance without retraining. Several typical CNN models, including VGG16, ResNet-18, ResNet-50 and GoogLeNet, were tested in this paper. Experiments revealed that 8-bit mantissa, including sign bit, in BFP representation merely induced less than 0.3% accuracy loss. In addition, we investigate the computational errors in theory and develop the noise-to-signal ratio (NSR) upper bound, which provides the promising guidance for BFP based CNN engine design.
Zhourui Song, Zhenyu Liu 0001, Dongsheng Wang 0002
AAAI2
2018 Enhance the HEVC Fast Intra CU Mode Decision Based on Convolutional Neural Network by Corner Power Estimation
abstract
HEVC (high efficiency video coding) uses block coding regime to fulfill the goal that allocates coding unit (CU) size according to the image complexity. It maintains a competitive visual quality with low bitrate. The process to decide the CU size, which is called rate distortion optimization (RDO), consumes a large amount of calculations. In this paper, we embed CNN into HEVC codec on efficiently performing the CU size decision and decreasing the amount of CU for full RDO process. Firstly, a corner-edge power estimate algorithm is proposed for early CU size decision and preparing balanced data for CNN training and inference. Secondly, the CNN with a deeper structure and stronger expressive ability is used in our mission. We proposed a method to adapt the learning rate according to different CNN layers for more effective training. Our algorithm saves overall 59.7% Intra encoding time, whereas the average BDBR increasing is merely 2.4%.
Liangliang Chang, Zhenyu Liu 0001
DCC2
2018 Rate Control Optimization of X265 Using Information from Quarter-Resolution Pre-Motion-Estimation
abstract
The rate control of x265 capitalizes on the low-resolution frame pre-motion-estimation to analyze the information prorogation in consecutive frames, with which to detect the scene cut, to decide the slice type, and to adjust the quantization ( Q) offsets in coding-unit (CU)-level, respectively. In this paper, we increase the accuracy of the frame-level base Q calculation and the slice type decision in x265. We implemented the proposed algorithms in x265 version 2.4. Experiments revealed that, 0.5904dB BDPNSR on average and up to 1.7205dB coding quality improvement were achieved.
Zhenyu Liu 0001
ICIP1
2018 CNN Based CU Partition Mode Decision Algorithm for HEVC Inter Coding
abstract
As compared with the predecessors, the superior compression performance of HEVC mainly stems from the hierarchical quadtree coding scheme, which is composed of coding unit(CU), prediction unit(PU), and transform unit(TU). The best CU/PU/TU partition mode is chosen from plenty of candidate modes. This procedure is denoted as rate-distortion optimization(RDO) that consumed more than 90% computation resources in HEVC encoding. In this paper, we devise the convolutional neural network(CNN) based fast CU mode decision algorithm for HEVC inter prediction. The contributions of our proposals include: (1) Because the maximum number of CU/PU candidate mode in one CTU is reduced, the corresponding VLSI encoder hardware complexity is ameliorated; (2) With the CTU pipeline architecture, the parallelism of the RDO processing will not be deteriorated by our fast algorithm. Our experiments show that the proposed VLSI friendly algorithm speeds up the HEVC inter coding by 45.0% at the cost of averagely 2.91 % Bjontegaard Delta bit-rate(BDBR) increase in HEVC reference test model HM-15.0.
Zhenyu Liu 0001, Xiangyang Ji, Dongsheng Wang 0002
ICIP2
2018 Parallel Content-Aware Adaptive Quantization-Oriented Lossy Frame Memory Recompression for HEVC
abstract
Since the development of ultrahigh-definition video, the huge bandwidth and power requirements of external memory have hindered the development of video encoder applications. Power constraints have become a particularly serious problem for portable video codec systems. With high-rate configurations [quantization parameter (QP) ≤ 22 in HEVC test model (HM) reference software], the compression performance of the existing lossless compression algorithms noticeably degrades, because the reference frames are becoming rich of textures. On the other hand, the mathematical analysis of this paper revealed that more quantization noises can be endured by the texture-rich area. Therefore, we develop an adaptive quantization-oriented parallel lossy frame memory recompression algorithm. The contributions of this paper include the following. First, a content-aware adaptive quantization method is devised to achieve a stable high compression ratio that does not deteriorate for highly quality texture-rich pictures. When QP$\in $ [12, 22], a data reduction ratio improvement of up to 14% is obtained compared with the best lossless algorithm. Furthermore, it can reduce the quality loss by 0.49-3.36 dB in terms of Bjøntegaard delta peak signal-to-noise rate (BD-PSNR) compared with the fixed length quantization method. Second, to solve the low throughput problem caused by the pixel-grain prediction method, a parallel directional prediction scheme is developed. It can double or quadruple the throughput with a prediction accuracy loss of only 1.7% or 3.3%, respectively. Using the above-mentioned methods, bandwidth and memory requirements are reduced up to 70.6% and 41.0%, respectively, with a corresponding savings of 59.3% in the dynamic power consumption of the off-chip dynamic random access memory, while the BD-PSNR is -0.04 dB, or, equivalently, Bjøntegaard delta bit rate (BD-BR) is 1.27%. Using TSMC 65-nm CMOS technology, the proposed frame memory compressor and decompressor can achieve the throughputs of up to 2.89 and 2.26 Gpixels/s, respectively. It is applicable to a Super Hi-Vision(8K)@68-frames/s real-time encoding with a Level D reference data reuse scheme.
Xiaocong Lian, Zhenyu Liu 0001, Wei Zhou 0020, Zhemin Duan
IEEE Trans. Circuits Syst. Video Technol.2
2017 Coding sensitive based approximation algorithm for power efficient VBS-DCT VLSI design in HEVC hardwired Intra encoder
abstract
High Efficiency Video Coding (HEVC), emerging as the latest video coding standard, obtained a 50% bit-rate reduction while maintaining the competitive visual quality as H.264/AVC. Rate-Distortion Optimization (RDO) is a computation intensive module in HEVC encoding. In specific, during Intra coding, RDO accounts for 62% of the overall encoding time. The 2-dimensional DCT is the most area and power consuming component for VLSI implementation of RDO module. In this paper, we decompose the matrix multiplication of DCT into several sparse butterfly structures in series. In addition, the computation and the storage of 25% high frequency coefficients are dropped by our approximation algorithm. The proposed algorithms are integrated in HM15.0. It is verified that our methods could save 15.9% time with 1.03% BDBR augment. We further implement the DCT VLSI design using TSMC 90nm standard cell library. In worst conditions (125°C, 0.9V), the power dissipation of our DCT is 12.7mW at the 311MHz maximum clock speed. As compared to the primitive design, we achieved 71.9% of hardware and 70.2% of power reductions.
Liangliang Chang, Zhenyu Liu 0001, Xiangyang Ji, Dongsheng Wang 0002
ICIP2
2017 Data-centric computation mode for convolution in deep neural networks
abstract
Deep Convolutional Neural Network (CNN) based methods have shown outstanding performance in a wide range of applications. Nowadays neural networks become deeper, leading to demand of substantial computation and memory resources. Customized hardware is one option which maintains high performance in lower energy consume than general CPUs or GPUs. While hardware designing, we need to address the problem of massive data transmission, and ensure high throughput at the same time. Actually, substantial data transfer consumes more energy than computation. The larger scales of neural networks become, the harder to solve this problem. In this paper, we focus on convolution operation, which occupies nearly 90% computation and runtime in deep CNN. We propose a data-centric computation mode for convolution, which declines the total requirements of data transfer during convolution processing period efficiently, and utilizes data locality to achieve high throughput. Different from previous methods, which adopt efficient on-chip memory hierarchy or focus on partial results' movements, our proposed method concentrates on operands themselves in convolution, minimizing data transfer right from the start. Obviously, it can be combined with others to achieve higher energy efficient. Furthermore, we also simulate and analyse the hardware overhead of our data-centric convolution, corroborating its potentiality of performing high throughput in low energy consumption.
Peiqi Wang 0001, Zhenyu Liu 0001, Haixia Wang 0001, Dongsheng Wang 0002
IJCNN2
2016 CNN oriented fast HEVC intra CU mode decision
abstract
The real-time requirements of hardwired HEVC encoder demand that, at the grain of coding tree unit (CTU), the maximum computation should be reduced by a fast CU mode decision algorithm. In addition, to realize the parallel rate-distortion optimization (RDO) of different CU modes, the current CU mode decision should not use the auxiliary information from other CU modes. Considering the above constraints, we applied convolutional neural network (CNN) to analyze the textures of source picture blocks, and then reduce the maximum number of CU modes, which will undergo the exhaustive RDO. In the CNN architecture design, we introduced the quantization parameter by considering the effect of quantization to the coding costs. We further optimized the CNN training strategy to improve the prediction accuracy. Experimental results demonstrated that, the proposed algorithm can save 63% Intra encoding time at the cost of the averaged 2.66% BDBR increase.
Zhenyu Liu 0001, Xianyu Yu, Shaolin Chen, Dongsheng Wang 0002
ISCAS1
2016 Error models of finite word length arithmetic in CNN accelerator design
abstract
Convolution Neural Network (CNN) is a state of the art machine learning algorithm. For CNN accelerator implementations, fixed-point and floating-point are two typical numeric representations. Because of the effects of rounding, reducing the word length would save the hardware and the power overheads while sacrificing the computation accuracy. The inherent robustness of neural network makes it possible to maintain the classification accuracy with very limited word length. Therefore, for the CNN accelerator designs, the primary issue is to determine the optimal arithmetic and the associated word length. In this paper, we developed the analytical error models to investigate the finite word length impacts of fixed-point and floating-point arithmetic in the test phase of CNN, respectively. It is revealed that the rounding errors are accumulated during the layer-wise convolution. Therefore, with the augment of network scale, the required word lengths for fixed-point and floating-point are both increased.
Zhenyu Liu 0001, Dongsheng Wang 0002
VCIP2
2016 HEVC fast FME algorithm using IME RD-costs based error surface fitting scheme
abstract
Motion Estimation (ME), which is composed of integer motion estimation (IME) and fractional motion estimation (FME), is the most computational intensive module in HEVC encoding procedure. In this paper, a new fast fractional pixel motion search method, that is based on a six-parameter two-dimension error surface model, is proposed. In our proposal, by solving the over-determined equations, nine integer-pixel rate-distortion costs (RDC), including the best integer-pixel search candidate and its eight neighboring integer pixels, are used to estimate the six parameters in the model. Then, we can obtain the minimal position on the fitted error surface equation, which is the quarter-pixel accurate search center. We provided three kinds of search patterns in the quarter-pixel search stage, which could take a tradeoff between the computational complexity and the prediction accuracy. Experimental results demonstrate that, as compared with HM reference software (HM-15.0), the three proposed FME patterns could reduce 35.1%, 29.4%, and 22.5% encoding time, while the corresponding compression efficiency losses in terms of BDBR are 3.04%, 0.79%, and 0.43%, respectively.
Zhenyu Liu 0001, Xiangyang Ji, Dongsheng Wang 0002
VCIP2
2016 Lossless Frame Memory Compression Using Pixel-Grain Prediction and Dynamic Order Entropy Coding
abstract
Power constraints constitute a critical design issue for the portable video codec system, in which the external dynamic random access memory (DRAM) accounts for more than half of the overall system power requirements. With the ultrahigh-definition video specifications, the power consumed by accessing reference frames in the external DRAM has become the bottleneck for the portable video encoding system design. To relieve the dynamic power stresses introduced by the DRAM, a lossless compression algorithm is devised to reduce the external traffic and the memory requirements of reference frames. First, pixel-granularity directional prediction is adopted to decrease the prediction residual energy by 54.1% over the previous horizontal prediction. Second, the dynamickth-order unary/Exp-Golomb rice coding is applied to accommodate the large-valued prediction residues. With the aforementioned techniques, an average data traffic reduction of 68.5% for the off-chip reference frames is obtained, which consequently reduces the dynamic power requirements of the DRAM by 42.3%. Based on the high data reduction ratio of the proposed compression algorithm, a partition group table-based storage space reduction scheme is provided to improve the utilization of row buffers in the DRAM. Consequently, an additional 14.5% of the DRAM dynamic power can be saved by reducing the number of row buffer activations. In total, a 56.8% decrease in the dynamic power requirements of the external reference frame access can be obtained using our strategies. With TSMC 65-nm CMOS logic technology, our algorithm was implemented in a parallel VLSI architecture based on a compressor and decompressor at the cost of 36.5k and 34.7k, respectively, in terms of gate count. The throughputs of the proposed compressor and decompressor are 1.54 and 0.78 Gpixels/s, which are suitable for quad full high definition (4K) @ 94 frames/s real-time encoding with the level-D reference data reuse scheme.
Xiaocong Lian, Zhenyu Liu 0001, Wei Zhou 0020, Zhemin Duan
IEEE Trans. Circuits Syst. Video Technol.2
2016 CU Partition Mode Decision for HEVC Hardwired Intra Encoder Using Convolution Neural Network
abstract
The intensive computation of High Efficiency Video Coding (HEVC) engenders challenges for the hardwired encoder in terms of the hardware overhead and the power dissipation. On the other hand, the constrains in hardwired encoder design seriously degrade the efficiency of software oriented fast coding unit (CU) partition mode decision algorithms. A fast algorithm is attributed as VLSI friendly, when it possesses the following properties. First, the maximum complexity of encoding a coding tree unit (CTU) could be reduced. Second, the parallelism of the hardwired encoder should not be deteriorated. Third, the process engine of the fast algorithm must be of low hardware- and power-overhead. In this paper, we devise the convolution neural network based fast algorithm to decrease no less than two CU partition modes in each CTU for full rate-distortion optimization (RDO) processing, thereby reducing the encoder's hardware complexity. As our algorithm does not depend on the correlations among CU depths or spatially nearby CUs, it is friendly to the parallel processing and does not deteriorate the rhythm of RDO pipelining. Experiments illustrated that, an averaged 61.1% intraencoding time was saved, whereas the Bjøntegaard-Delta bit-rate augment is 2.67%. Capitalizing on the optimal arithmetic representation, we developed the high-speed [714 MHz in the worst conditions (125 °C, 0.9 V)] and low-cost (42.5k gate) accelerator for our fast algorithm by using TSMC 65-nm CMOS technology. One accelerator could support HD1080p at 55 frames/s real-time encoding. The corresponding power dissipation was 16.2 mW at 714 MHz. Finally, our accelerator is provided with good scalability. Four accelerators fulfill the throughput requirements of UltraHD-4K at 55 frames/s.
Zhenyu Liu 0001, Xianyu Yu, Shaolin Chen, Xiangyang Ji, Dongsheng Wang 0002
IEEE Trans. Image Process.1
2016 A High-Throughput and Multi-Parallel VLSI Architecture for HEVC Deblocking Filter
abstract
This paper presents a high-throughput and multi-parallel VLSI hardware architecture for the deblocking filter in the HEVC video coding standard. First, an implementation-friendly and fast boundary judgment method is proposed to avoid using the original recursion loop approach. Then a dedicated parallel VLSI architecture composed of four parallel filtering cores is presented based on the proposed boundary judgment method. With the parallel luma/chroma filtering and parallel vertical/horizontal edges filtering order, the proposed VLSI architecture can process filtering operations for one largest coding unit (LCU) with less filtering cycles than other conventional approaches. Furthermore, filtering efficiency is improved due to a novel ping-pang buffer architecture and the on-chip single-port SRAM with dedicated data arrangement in the memory modules. Experimental results demonstrate that the proposed deblocking filter architecture improves the performance by 28-89% at the expense of the slightly increased gate count compared to the previously known architecture in HEVC. The proposed architecture can reach a high operating clock frequency of 278 MHz with TSMC 90 nm library and meet the real time requirement of the deblocking filter for 8 K × 4 K video format at 123 frame/s.
Wei Zhou 0020, Jingzhi Zhang, Xin Zhou 0001, Zhenyu Liu 0001, Xiaoxiang Liu
IEEE Trans. Multim.4
2015 Pixel-grain prediction and K-order UEG-rice entropy coding oriented lossless frame memory compression for motion estimation in HEVC
abstract
With the resolution of video sequences increasing, the memory bandwidth and space requirement has become one bottleneck of a video coding system. Most of previous frame memory compression works merely focused on reducing the averaged bandwidth demand. In this paper, a lossless frame memory compression algorithm for motion estimation in HEVC and the corresponding hardware architecture are described. The compression algorithm is composed of the 64×4 partition compression algorithm and the partition group table based storage scheme. Experimental results demonstrate that, on average, the proposed algorithm reduces 70.1% memory bandwidth and 41% memory space, while saving 59.0% dynamic energy. With TSMC 90nm CMOS logic technology, our algorithm were implemented in a parallel VLSI architecture of compressor and decompressor at the cost of 48.1K and 44.2K respectively in gate-count. The achieved peak throughput is 1.51Gpixel/s, which is sufficient to handle the SHV(8K)@30fps real-time encoding with the level-D reference data reuse scheme.
Xiaocong Lian, Zhenyu Liu 0001, Wei Zhou 0020, Zhemin Duan
ICIP2
2015 VLSI friendly fast CU/PU mode decision for HEVC intra encoding: Leveraging convolution neural network
abstract
To alleviate the computational intensity of Intra encoding for High Efficiency Video Coding (HEVC), we introduce the convolution neural network to reduce the number of the promising CU/PU candidate modes to carry out the exhaustive RDO processing. The practical merits include: Firstly, the proposed algorithm reduces the maximum computational complexity at the grain of 64 × 64 coding tree unit(CTU), which makes it efficient to ameliorate the complexity of the real-time hardwired encoder implementation. Secondly, because the CU/PU mode decision is made based on the analysis of source block textures, our algorithm does not depend on intermediate results of encoding. That is, the proposed algorithm will not deteriorate the processing schedule of CTU encoding. Experimental results show that, when our algorithm is integrated with HM12.0, the 61.1% Intra encoding time was saved, whereas the averaging BDBR augment is merely 3.39%.
Xianyu Yu, Zhenyu Liu 0001, Dongsheng Wang 0002
ICIP2
2015 Deblocking strength prediction based CTU-level SAO category determination in HEVC encoder
abstract
High efficiency video coding (HEVC) is a video compression standard that outperforms the predecessor H.264/AVC by doubling the compression efficiency. To enhance the coding accuracy, HEVC adopts sample adaptive offset (SAO), which reduces the distortion of reconstructed pixels using classification based non-linear filtering. In the traditional coding tree unit (CTU) based VLSI encoder implementation, during the pixel classification stage, SAO cannot use the raw samples in the boundary of the current CTU because these pixels have not been processed by deblocking filter (DF). This paper proposes a category determination algorithm based on estimating the deblocking strengths on CTU boundaries and selectively adopting the promising samples in these areas during SAO classification. Compared with HEVC test mode (HM11.0), experimental results indicate that the proposed method achieves an average 0.15% BD-bitrate reduction (equivalent to 0.0084 dB increases in P-SNR).
Gaoxing Chen, Zhenyu Pei, Zhenyu Liu 0001, Takeshi Ikenaga
VCIP3
2014 HDTV1080p HEVC Intra encoder with source texture based CU/PU mode pre-decision
abstract
HEVC doubles the coding efficiency with more than 4x coding complexity as compared to H.264/AVC. To alleviate the burden of Intra encoder, we estimate the RD-cost from the source image textures, and dynamically select two promising CU/PU mode candidates to execute exhaustive RDO processing. As integrated in our hardwired encoder, the averaged 61.7% computation complexity was saved with 4.53% rate augment. With TSMC 90nm technology, the real-time encoder for HDTV1080p at 44fps is implemented with 2269k-gate at 357MHz operating frequency.
Zhenyu Liu 0001, Dongsheng Wang 0002, Qingrui Han, Yang Song 0002
ASP-DAC2
2014 Linear Rate Estimation Model for HEVC RDO Using Binary Classification Based Regression
abstract
Rate-Distortion Optimization in High Efficiency Video Coding promotes the coding efficiency, but also imposes intensive computation to the encoder, because the complex Syntax-based context-adaptive Binary Arithmetic Coding is performed for each candidate coding configuration. We develop the classification based regression method to derive the rate models, which fast estimate the bit cost of quantization coefficient block from its distribution features. Experiments demonstrate that, our method reduces the averaged 28.4% computation time in rate cost estimation, while the coding efficiency degradation is 0.0428dB.
Sanchuan Guo, Zhenyu Liu 0001, Dongsheng Wang 0002, Qingrui Han, Yang Song 0002
DCC2
2014 Linear adaptive search range model for uni-prediction and motion analysis for bi-prediction in HEVC
abstract
High Efficiency Video Coding (HEVC) is the up-to-date video coding standard. Compared to the predecessor H.264/AVC, HEVC can further reduce approximately 50% bit rate on average with the competing perceptual quality. On the other hand, experiment shows that HEVC requires more than 4 times computational complexity during the encoding procedure. In ours test, even using fast TZSearch, integer motion estimation (IME) still accounts for 20%-30% of encoding time. In this paper, we propose two adaptive search range (ASR) algorithms to address this problem in IME. First, we present an ASR algorithm based on linear adaptive search range model (LAM-ASR) for uni-prediction. This model considers the impacts of the motion consistency, PU size and the amplitude of motion vector predictor (MVP). In order to offer more flexibility, we introduce a scale factor to this model. Second, for bi-prediction, we propose another ASR algorithm based on motion analysis (MA-ASR), which assigns different search range to PU by making full use of the motion information obtained from uni-prediction. Experimental results show that when embedded into the fast TZSearch method of the reference software, the two proposed ASR algorithms can averagely save 42.0% of the IME time with 0.023dB BD-PSNR degradation or equally 0.7% BD-BR increase.
Longshan Du, Zhenyu Liu 0001, Takeshi Ikenaga, Dongsheng Wang 0002
ICIP2
2014 Binary classification based linear rate estimation model for HEVC RDO
abstract
Rate-Distortion Optimization in High Efficiency Video Coding promotes the coding efficiency, but also imposes intensive computation to the encoder, because the complex Syntax-based context-adaptive Binary Arithmetic Coding is performed for each candidate coding configuration. We develop the classification based regression method to derive the rate models, which fast estimate the bit cost of quantization coefficient block from its distribution features. Experiments demonstrate that, our method reduces the averaged 28.4% computation time in rate cost estimation, while the coding efficiency degradation is 0.0428dB.
Zhenyu Liu 0001, Sanchuan Guo, Dongsheng Wang 0002
ICIP1
2013 41.7BN-pixels/s reconfigurable intra prediction architecture for HEVC 2560×1600 encoder
abstract
The complexity of High Efficiency Video Coding (HEVC) intra prediction design mainly comes from two aspects. First, as compared with the predecessor H.264/AVC, HEVC increases the number of prediction angles from 9 up to 33. Second, HEVC employs 5 kinds of n×n prediction unit size, including 4×4, 8×8, 16×16, 32×32 and 64×64. The computation intensity of intra encoding is increased by one order. In this paper, we provide the high efficient reconfigurable VLSI architecture for all intra directional prediction modes. The proposed design possesses the following merits: (1) Our prediction engine is equipped with sixteen uniform modules, and can be configured to produce 2 · m number row-wise n prediction samples in each cycle, where n = {4, 8, 16, 32, 64} and m = 64/n; (2) As our design always produces the row-wise samples, the hardware consuming transpose register array between the prediction residue module and the following DCT engine is eliminated. This feature further avoids the bubble operations in the horizontal predictions. With TSMC 90nm CMOS technology, the proposed architecture achieves 357MHz operating frequency at the cost of 817.3k gates, and the corresponding power dissipation is 114mW. Our implementation can fulfill the throughput requirement of HD2560 × 1600@46fps real-time encoding.
Zhenyu Liu 0001, Dongsheng Wang 0002, Hongxiang Zhu
ICASSP1
2013 Bayesian theory oriented Optimal Data-Provider Selection for CMP
abstract
With the number of cores and working sets of parallel workloads soaring, shared L2 caches exhibit fewer misses than private L2 caches via making better use of the all available cache capacity. However, shared L2 caches induce higher overall L1 miss latencies because of longer average distance between requestor and home node, and potentially congestions at some nodes. We observe that there is a high probability that the requested data of an L1 miss resides in a neighbor node's L1 cache. In such cases, these long-distance accesses to the home nodes can be potentially avoided. In order to successfully leverage the aforementioned property, we propose Bayesian theory oriented Optimal Data-Provider Selection (ODPS). ODPS partitions the multi-core into clusters of 2×2 nodes, and introduces the Proximity Data Prober (PDP) to detect whether an L1 miss can be served by one L1 cache within the same cluster. Furthermore, we devise the Bayesian Decision Classifier (BDC) to intelligently and adaptively select a remote L2 cache or a neighboring L1 node as the data provider according to the minimal miss cost based on the Bayesian decision theory.
Guohong Li, Zhenyu Liu 0001, Sanchuan Guo, Chongmin Li, Dongsheng Wang 0002
ICCD2
2013 Fast prediction mode decision with hadamard transform based rate-distortion cost estimation for HEVC intra coding
abstract
The emerging video coding standard, High Efficiency Video Coding (HEVC), aims at doubling coding efficiency of H.264/AVC. In the intra encoding, Rate-Distortion Optimization (RDO) processes are employed to determine the best prediction mode. RDO based mode decision accounts for 35-39% coding time, because the large scale two-dimensional(2D) DCT/IDCT introduces a plethora of computation- and hardware-consuming multiplications. In this paper, the simplified Rate-Distortion (RD) cost estimation algorithms, which are based on the Hadamard transform and the distortion evaluation without the signal reconstruction, are proposed. When embedded to HEVC test Model(HM-5.2), our methods averagely achieve 16.1% time saving at the price of 0.055dB BD-PSNR loss, or equivalently 1.27% BD-BR increasing in intra coding. The corresponding VLSI design of the proposed algorithms is implemented with TSMC 90nm 1P9M technology. The maximum clock speed is 418 MHz under the worst work conditions (125°C, 0.9V). As compared with the primitive design, 64.9% hardware cost can be saved by our schemes. One proposed engine fulfills the throughput for 4K-UHDTV@28fps real-time encoding.
Zhenyu Liu 0001, Dongsheng Wang 0002, Qingrui Han, Yang Song 0002
ICIP2
2013 Fast HEVC intra mode decision using matching edge detector and kernel density estimation alike histogram generation
abstract
Intra coding algorithm in High Efficiency Video Coding employs up to 35 directional prediction modes. Upon the end of alleviating the intra encoding complexity, we proposed the candidate mode selection algorithm from analyzing the textures of the source image block. Considering the fine difference between the neighboring prediction directions, we devise the fix-point arithmetic based edge detector, which improves the direction detection accuracy as compared with the typical previous works while maintaining the low computational overhead. To improve the robustness of the edge direction statistics, we further introduce the conception of kernel density estimation into the histogram calculation. Our proposals is orthogonal to the published HEVC fast intra mode decision algorithms. Experimental results verified that, on average, the proposed methods reduced the encoding time by 25.21% in high efficiency mode, and 37.61% in low complexity mode, whereas the averaging BDPSNR losses are 0.0608dB and 0.0781dB, respectively.1.
Zhenyu Liu 0001, Takeshi Ikenaga, Dongsheng Wang 0002
ISCAS2
2013 Content-aware write reduction mechanism of phase-change RAM based Frame Store in H.264 Video codec system
abstract
H.264 video codec system requires big capacity of Frame Store (FS) for buffering reference frames. The up-to-date Phase-change Random Access Memory (PRAM) is the promising approach for on-chip caching the reference signals, as PRAM offers the advantages in terms of high density and low leakage power. However, the write endurance problem, that is a PRAM cell can only tolerant limited number of write operations, becomes the main barrier in practical applications. This paper studies the wear reduction techniques of PRAM based FS in H.264 codec system. On the basis of rate-distortion theory, the content oriented selective writing mechanisms are proposed to reduce bit updates in the reference frame buffers. Experiments demonstrate that, for typical video sequences with different frame sizes, our methods averagely achieve more than 30% reduction of bit updates, while introducing around 20% BDBR cost. The power consumption is reduced by 55% on average, and the estimated PRAM lifetime is extended by 61%.
Sanchuan Guo, Zhenyu Liu 0001, Guohong Li, Dongsheng Wang 0002
ISCAS2
2013 A mode-mapping and optimized MV conjunction based MGS-scalable SVC to AVC IPPP transcoder
abstract
Scalable Video Coding (SVC) is an extension of H.264/AVC, aiming to provide the ability to adapt to heterogeneous environments. It offers great flexibility for bitstream adaptation in multi-point applications such as videoconferencing. However, transcoding between SVC and AVC is necessary due to the existence of legacy AVC-based systems. This paper proposes a 3-stage fast SVC-to-AVC transcoder for medium-grain quality scalability (MGS). Hierarchical-P structured SVC bitstream is transcoded into IPPP structured AVC bitstream with multiple reference frames. In the first stage, mode decision is accelerated by proposed SVC-to-AVC mode mapping scheme. In the second stage, INTER motion estimation is accelerated by an optimized motion vector (MV) conjunction method to predict the MV with a reduced search range. In the last stage, Hadamard-based all zero block (AZB) detection is utilized for early termination. Simulation results show that proposed transcoder achieves very similar coding efficiency as the optimal result, but with averagely 92.3% computational time saving.
Lei Sun 0005, Zhenyu Liu 0001, Takeshi Ikenaga
ISCAS2
2013 Fully pipelined DCT/IDCT/Hadamard unified transform architecture for HEVC Codec
abstract
Great amount of two-dimensional (2D) discrete cosine transforms and Hadamard transforms are executed in HEVC. Upon the end of real-time UHDTV Codec, the full pipeline variable block size 2D transform engine with the efficient hardware utilization is proposed to handle the DCT/IDCT and Hadamard transforms. The efficiency comes from two aspects. First, the hardware for small-size transforms is fully reused by other larger-size transform processing. Second, we devise the unified architecture for IDCT and DCT through the algorithm optimization. The maximum clock speed of our design is 311MHz under 90nm technology. Experiments demonstrate that, at 47MHz clock frequency, one proposed engine provides the throughput for 8K-UHDTV real-time decoding, and it also fully supports the real-time encoding of HDTV1080p@20fps with 311MHz clock speed1.
Zhenyu Liu 0001, Dongsheng Wang 0002
ISCAS2
2013 A Low-Complexity Quantization-Domain H.264/SVC to H.264/AVC Transcoder with Medium-Grain Quality Scalability
Lei Sun 0005, Zhenyu Liu 0001, Takeshi Ikenaga
MMM (1)2
2013 Cluster Cache Monitor
abstract
As the number of cores and the working sets of parallel workloads increase, shared L2 caches exhibit fewer misses than private L2 caches by making a better use of the total available cache capacity, but they also induce higher overall L1 miss latencies because of the longer average distance between two nodes, and the potential congestions at certain nodes. One of the main causes of the long L1 miss latencies are accesses to home nodes of the directory. However, we have observed that there is a high probability that the target data of an L1 miss resides in the L1 cache of a neighbor node. In such cases, these long-distance accesses to the home nodes can be potentially avoided. We organize the multi-core into clusters of 2×2 nodes, and in order to leverage the aforementioned property, we introduce the Cluster Cache Monitor (CCM). The CCM is a hardware structure in charge of detecting whether an L1 miss can be served by one of the cluster L1 caches, and two cluster-related states in the coherence protocol in order to avoid long-distance accesses to home nodes upon hits in the cluster L1 caches. We evaluate this approach on a 64-node multi-core using SPLASH-2 and PARSEC benchmarks, and we find that the CCM can reduce the execution time by 15% and reduce the energy by 14%, while saving 28% of the directory storage area compared to a standard multi-core with a shared L2. We also show that the CCM outperforms recent mechanisms, such as ASR, DCC and RNUCA.
Guohong Li, Olivier Temam, Zhenyu Liu 0001, Dongsheng Wang 0002, Sanchuan Guo
SBAC-PAD3
2012 Lagrangian Multiplier Optimization Using Markov Chain Based Rate and Piecewise Approximated Distortion Models
abstract
The traditional Lagrangian RDO algorithm assumes the transformed residues as memo- ryless random variables, and then doesn't perform well when the prediction residues posses the strong temporal correlations. We extend the RDO by modeling the residues as the first-order Markov source and calibrating the distortion model with the piecewise approximation function.
Zhenyu Liu 0001, Dongsheng Wang 0002, Takeshi Ikenaga
DCC1
2012 Lagrangian multiplier optimization using correlations in residues
abstract
Rate distortion optimization (RDO) algorithm plays the vital role in the up to date hybrid video codec H.264/AVC. The RDO algorithm of H.264/AVC reference software is built up by assuming that the transformed residues are memoryless variables. However, our experiments reveal that, for some sequences, the strong temporal correlations exist in the prediction residues. This paper extends the Lagrangian optimization techniques by modeling the transformed residues as the first-order Markov source and calibrating the distortion model with the piecewise approximation function. The proposed algorithms adjust the Lagrangian multiplier dynamically to improve the overall coding quality. Comprehensive experiments testify that, as compared with the JM reference software, our optimizations can achieve up to 1.875dB coding gain. Moreover, our algorithms posses more robust coding performance and introduce less computational overhead than the Laplace distribution based methods. The inherent short process latency makes it possible to cooperate our algorithms with rate control operation. Last but not least, the proposed approach is also useful for the emerging standard, HEVC.
Zhenyu Liu 0001, Dongsheng Wang 0002, Takeshi Ikenaga
ICASSP1
2012 A pixel-domain mode-mapping based SVC-to-AVC transcoder with coarse grain quality scalability
Lei Sun 0005, Zhenyu Liu 0001, Takeshi Ikenaga
ICPR2
2012 Wear-Resistant Hybrid Cache Architecture with Phase Change Memory
abstract
Phase-change Random Access Memory (PRAM) is one of the most promising technologies among emerging non-volatile memory technologies, which provides many benefits, such as high density, non-volatility and low leakage power. However, the limited write endurance of PRAM prevents it from being used as a drop-in replacement of SRAM cache. Moreover, the inherent high latency and power dissipation of write operations are both hindrances that PRAM faces. In this paper, we study the L2 cache write operations incurred by different types of data, and accordingly, propose Wear-Resistant Hybrid Cache Architecture (WRHCA), in which the write access behavior of the hybrid L2 cache, that is composed of SRAM and PRAM, is optimized. Through the prediction of data access patterns, the proposed WRHCA prevents write-prone data from entering PRAM L2 cache, and consequently, the wear-out issue of PRAM is alleviated efficiently. Experimental results on the basis of the trace-driven simulator demonstrate that, as compared to the baseline system with pure PRAM L2 cache, our optimized WRHCA saved 85.5% write operations to PRAM on average, and boosted the performance by the averaged 6.4% CPI reduction. Last but not least, as compared with the primitive 3-level SRAM cache with the same chip area, our WRHCA achieved 60.9% reduction in terms of power consumption.
Sanchuan Guo, Zhenyu Liu 0001, Dongsheng Wang 0002, Haixia Wang 0001, Guohong Li
NAS2
2011 One-round renormalization based 2-bin/cycle H.264/AVC CABAC encoder
abstract
Context-based Adaptive Binary Arithmetic Coder (CABAC) is the advanced entropy coding tool employed by main and higher profiles of H.264/AVC. As compared with Context-based Adaptive Variable Length Coding (CAVLC), under the same bit rate, CABAC achieves up to 0.5dB PSNR gain. On the other hand, the high complexity of CABAC severely hinders the whole encoder throughput. To over- come the throughput bottleneck of CABAC, the authors devise the one-round renormalization and the associated VLSI architecture to omit the multiple-iteration operation of one bin's encoding. The pro- posed full-context CABAC hardwired encoder garners the constant 2-bin/cycle throughput. Using TSMC one-poly nine-metal 90 nm CMOS technology, the prototyping is implemented with 33.9k logic gates and 1562-bit on-chip SRAM. In the worst operating conditions (0.9V, 125°C), the operating frequency is 238.1MHz, which can support HDTV720p real-time encoding at 329 fps frame rate with the quantization parameter (QP) not less than 18.
Zhenyu Liu 0001, Dongsheng Wang 0002
ICIP1
2011 Register Length Analysis and VLSI Optimization of VBS Hadamard Transform in H.264/AVC
abstract
Fidelity range extensions of H.264/AVC adopt variable block size (VBS) transform techniques to employ 8 × 8/4 × 4 Hadamard transforms adaptively during the fractional motion estimation. In this literature, the hardwired VBS Hadamard transform accelerator is developed with the following contributions: 1) developed a hardware reusing scheme between 8 × 8 and 4 × 4 transforms within the architecture design; 2) devised the intermediate bit-truncation algorithm to reduce the hardware cost while maintaining the computational precision well; and 3) reduced the bit-width of sum of absolute transformed differences (SATD) value as compared to the primitive implementation, resulting in optimization in both power and hardware cost for the SATD generator implementation. With TSMC 0.18 μm CMOS technology, the experiments demonstrate that for each VBS Hadamard transform engine, 13.0-30.4% saving in hardware cost and 12.6-32.4% saving in power consumption are achieved, whereas the incurred coding quality loss is less than 0.2089 dB in terms of BDPSNR. From the aspect of the whole encoder implementation, and considering the parallelism in searching factional pixel candidates, the proposed strategies garner 2.0 3.9% overall gate count reduction.
Zhenyu Liu 0001, Dongsheng Wang 0002, Takeshi Ikenaga
IEEE Trans. Circuits Syst. Video Technol.1
2009 Hardware optimizations of variable block size Hadamard transform for H.264/AVC FRExt
abstract
Variable block size (VBS) transform technique is adopted in Fidelity Range Extensions (FRExt) of H.264/AVC, in which 8 × 8/4 × 4 Hadamard transforms are adaptively employed during the fractional motion estimation. The hardwired VBS Hadamard transform unit is developed by authors and the following contributions are described in this literature: (1) Hardware reusing scheme is adopted in the architecture design; (2) In the light of the noise analysis, the intermediate data bit-truncation scheme is developed to reduce the hardware cost while maintaining its computational precision well; (3) With mathematical analysis, the bit-width of SATD value is reduced as compared to the intuitive implementation, therefore, the power and hardware cost are both optimized for the SATD generator implementation; (4) Hybrid 4:2/3:2 compressor based CSA tree is analyzed in the circuits design of SATD generator; and (5) Clock-gating technique is employed to reduce the power dissipation of 4×4 transform operation. With TSMC 0.18 ¿m CMOS technology, experimental results reveal that 12.2-30.4% saving in hardware cost and 12.4-32.4% saving in power consumption are achieved by using our algorithms.
Zhenyu Liu 0001, Dongsheng Wang 0002, Takeshi Ikenaga
ICIP1
2009 Motion Estimation Optimization for H.264/AVC Using Source Image Edge Features
abstract
The H.264/AVC coding standard processes variable block size motion-compensated prediction with multiple reference frames to achieve a pronounced improvement in compression efficiency. Accordingly, the computation of motion estimation increases in proportion to the product of the number of reference frame and the number of intermode. The mathematical analysis in this paper illustrates that the motion-compensated prediction errors are mainly determined by the detailed textures in the source image. The image block being rich in textures contains numerous high-frequency signals, which make variable block size and multiple reference frame techniques essential. On the basis of rate-distortion theory, in this paper, the spatial homogeneity of an image block is made as a relative concept with respect to the current quantization step. For the homogenous block, its futile reference frames and intermodes can be eliminated efficiently. It is further revealed that the sum of absolute differences value of an image block is mainly determined by the sum of its edge gradient amplitude and the current quantization step. Consequently, the image content-based early termination algorithm is proposed, and it outperforms the original method adopted by JVT reference software. Moreover, the dynamic search range algorithm based on the edge gradient amplitude of source image block is analyzed. One eminent advantage of the proposed edge-based algorithms is their efficiency to the macroblock-pipelining architecture, and another desirable feature is their orthogonality to fast block-matching algorithms. Experimental results show that when these algorithms are integrated with hybrid unsymmetrical-cross multi-hexagongrid search, an averaged 31.4-60.0% motion estimation time can be saved, whereas the averaging BDPSNR loss is 0.0497 dB for all tested sequences.
Zhenyu Liu 0001, Satoshi Goto, Takeshi Ikenaga
IEEE Trans. Circuits Syst. Video Technol.1
2008 Optimization of Propagate Partial SAD and SAD tree motion estimation hardwired engine for H.264
abstract
Variable block size motion estimation algorithm is the effcient approach to reduce the temporal redundancies and it has been adopted by the latest video coding standard H.264/AVC. The computational complexity augment coming from the variable block size technique makes the hardwired accelerator essential, especially for real-time applications. In this paper, the authors apply the architecture level and the circuits level approaches to improve the performance of Propagate Partial SAD and SAD Tree hardwired engines, which outperform other counterparts when considering the impact of supporting the variable block size technique. Experiments demonstrate that by using the proposed approaches, compared with the original architectures, 14.7% and 18.0% hardware cost can be saved for Propagate Partial SAD architecture and SAD Tree architecture, respectively. With TSMC 0.18 mm 1P6M CMOS technology, the proposed Propagate Partial SAD architecture attains 231.6 MHz operating frequency at a cost of 84.1 k gates. Correspondingly, the execution speed of the optimized SAD Tree architecture is improved to 204.8 MHz with 88.5 k gate hardware overhead.
Zhenyu Liu 0001, Satoshi Goto, Takeshi Ikenaga
ICCD1
2008 Fast motion estimation for H.264/AVC using image edge features
abstract
The key to high performance in video coding lies on efficiently reducing the temporal redundancies. For this purpose, H.264/AVC coding standard has adopted variable block size motion estimation on multiple reference frames to improve the coding gain. However, the computational complexity of motion estimation is also increased in proportion to the product of the reference frame number and the inter mode number. The mathematical analysis reveals that the prediction errors mainly depend on the image edge gradient amplitude and quantization parameter. Consequently, this paper proposed the image content based early termination algorithm, which outperforms the method adopted by JVT reference software, especially at high and moderate bit rates. In light of rate-distortion theory, this paper also relates the homogeneity of image to the quantization parameter. For the homogenous block, its search computation for the futile reference frames and inter modes can be efficiently discarded. Then the computation saving performance increases with quantization parameter. These content based fast algorithms were integrated with Unsymmetrical-cross Multihexagon-grid Search (UMHexagonS) algorithm. Compared to the original UMHexagonS fast matching algorithm, 26.14-54.97% search time can be saved with an average of 0.0369dB coding quality degradation.
Zhenyu Liu 0001, Satoshi Goto, Takeshi Ikenaga
ICME1
2008 Hardware-oriented direction-based fast fractional motion estimation algorithm in H.264/AVC
abstract
In this paper, a hardware-oriented fast fractional motion estimation (FME) algorithm for H.264/AVC is proposed. For both 1/2-pixel and 1/4-pixel FME, only 4 points are searched, and thus the required processing unit (PU) number is reduced from 9 to 4. Experiments show that the proposed algorithm can provide very similar image quality as full search, and the average PSRN drop is 0.17dB or bitrate increase is 4.08%. Compared with the full search FME hardware architecture [5], the proposed algorithm can effectively reduce about 32.3% hardware cost and thus is very suitable for low-cost and low-power applications.
Yang Song 0002, Zhenyu Liu 0001, Takeshi Ikenaga, Satoshi Goto
ICME3
2008 Motion Feature and Hadamard Coefficient-Based Fast Multiple Reference Frame Motion Estimation for H.264
abstract
In the state-of-the-art video coding standard, H.264/AVC, the encoder is allowed to search for its prediction signals among a large number of reference pictures that have been decoded and stored in the decoder to enhance its coding efficiency. Therefore, the computation complexity of the motion estimation (ME) increases linearly with the number of reference picture. Many fast multiple reference frame ME algorithms have been proposed, whose performance, however, will be considerably degraded in the hardwired encoder design due to the macroblock (MB) pipelining architecture. Considering the limitations of the traditional four-stage MB pipelining architecture, two fast multiple reference frame ME algorithms are proposed here. First, on the basis of mathematical analysis, which reveals that the efficiency of multiple reference frames will be degraded by the relative motion between the camera and the objects, for the slow-moving MB, the authors adopt the multiple reference frames but reduce their search range. On the other hand, for the fast-moving MB, the first previous reference frame is used with the full search range during the ME processing. The mutually exclusive feature between the large search range and the multiple reference frames makes the computation saving performance of the proposed algorithm insensitive to the nature of video sequence. Second, following the Hadamard transform coefficient-based all_zeros block early detection algorithm, two early termination criteria are proposed. These methods ensure the pronounced computation saving efficiency when the encoded video has strong spatial homogeneity or temporal stationarity. Experimental results show that 72.7%-93.7% computation can be saved by the proposed fast algorithms with an average of 0.0899 dB coding quality degradation. Moreover, these fast algorithms can be combined with fast block matching algorithms to further improve their speedup performance.
Zhenyu Liu 0001, Yang Song 0002, Satoshi Goto, Takeshi Ikenaga
IEEE Trans. Circuits Syst. Video Technol.1
2007 Hardware-efficient propagate partial sad architecture for variable block size motion estimation in H.264/AVC
abstract
One hardware efficient and high speed architecture for variableblock size motion estimation in H.264 is presented in this paper. Through compressing the propagated data and optimizing theprocessing element and adder tree circuits in pipeline, this architecture gets more hardware efficient datapath logic. Compared with the original Propagate Partial SAD structure, 12.1% hardware cost can be saved. With TSMC 0.18μm CMOS 1P6M standard celllibrary, the maximum clock speed of this design is 227MHz in worstwork conditions (1.62V, 125°C). With the 48x32 search range, the maximum throughput of our design is 147786 MB/S, which can be used in the real-time encoding of VGA resolution frame with 4 reference frames at 30Hz.
Zhenyu Liu 0001, Yiqing Huang 0002, Yang Song 0002, Satoshi Goto, Takeshi Ikenaga
ACM Great Lakes Symposium on VLSI1
2007 VLSI Oriented Fast Multiple Reference Frame Motion Estimation Algorithm for H.264/AVC
abstract
In H.264/AVC standard, motion estimation can be processed on multiple reference frames (MRF) to improve the video coding performance. For the VLSI real-time encoder, the heavy computation of fractional motion estimation (FME) makes the integer motion estimation (IME) and FME must be scheduled in two macro block (MB) pipeline stages, which makes many fast MRF algorithms inefficient for the computation reduction. In this paper, two algorithms are provided to reduce the computation of FME and IME. First, through analyzing the block's Hadamard transform coefficients, all-zero case after quantization can be accurately detected. The FME processing in the remaining frames for the block, detected as all-zero one, can be eliminated. Second, because the fast motion object blurs its edges in image, the effect of MRF to aliasing is weakened. The first reference frame is enough for fast motion MBs and MRF is just processed on those slow motion MBs with a small search range. The computation of IME is also highly reduced with this algorithm. Experimental results show that 61.4%-76.7% computation can be saved with the similar coding quality as the reference software. Moreover, the provided fast algorithms can be combined with fast block matching algorithms to further improve the performance.
Zhenyu Liu 0001, Yang Song 0002, Takeshi Ikenaga, Satoshi Goto
ICME1
2007 Ultra Low-Complexity Fast Variable Block Size Motion Estimation Algorithm in H.264/AVC
abstract
Variable block size motion estimation (VBSME) is the most computation consuming part in the newest H.264 /AVC video coding standard. To reduce computation, an ultra low-complexity fast VBSME algorithm is proposed in this paper. Different to previous fast algorithms which separately pressed the 7 block modes, in the presented algorithm, all the block modes are simultaneously calculated for most of the search points, and thus the SADs can be reused by all blocks. Moreover, a low-pass filter based pixel decimation is also introduced to decrease the matching cost. Experiments show that the proposed fast VBSME algorithm can reduce the computation to about 0.2% of fast full search (FFS) with robust image quality.
Yang Song 0002, Zhenyu Liu 0001, Takeshi Ikenaga, Satoshi Goto
ICME2
2007 Enhanced Strict Multilevel Successive Elimination Algorithm for Fast Motion Estimation
abstract
This paper presents an enhanced strict multilevel successive elimination algorithm (EMSEA) for fast block-matching motion estimation, which is based on the previous strict multilevel successive elimination algorithm (SMSEA) (Song et al., 2006). Different to the SMSEA algorithm with fixed parameters, in EMSEA algorithm, the whole search area is divided into two regions and each region has its own parameters. Therefore, the computation complexity of SMSEA algorithm can be further decreased. Experiments show that the proposed EMSEA algorithm can reduce 16.2% of the SMSEA computation and maintain almost the same image quality, which is better than the TSS and DS algorithms.
Yang Song 0002, Zhenyu Liu 0001, Takeshi Ikenaga, Satoshi Goto
ISCAS2
2006 Low-Pass Filter Based Vlsi Oriented Variable Block Size Motion Estimation Algorithm for H.264
abstract
In this paper, a fast motion estimation algorithm, which is friendly to VLSI hardware implementation is proposed. This algorithm has such features: First, through "Haar" low-pass filter based subsampling, the computation complexity at each search position is reduced to about 25% of the original algorithm; Second, one modified motion vector prediction is provided to eliminate the data dependence among sub-partitions in the same macro block (MB). Based on this approach, parallel processing for variable block size motion estimation (VBSME) with integer pixel accuracy can be realized; Third, one "adaptive sub-search window" scheme is proposed to further reduce computation cost and it also can facilitate reference frame data reusing to reduce memory transfer from the external RAM to the on-chip SRAM. The proposed VBSME algorithm is very suitable for parallel VLSI implementation
Zhenyu Liu 0001, Yang Song 0002, Takeshi Ikenaga, Satoshi Goto
ICASSP (2)1
2005 A VLSI array processing oriented fast fourier transform algorithm and hardware implementation
abstract
Many parallel Fast Fourier Transform (FFT) algorithms adopt multiple stages architecture to increase performance. However, data permutation between stages consumes volume memory and processing time. An FFT array processing mapping algorithm is proposed in this paper to overcome this demerit. In this algorithm, arbitrary 2k butterfly units (BUs) could be scheduled to work in parallel on n=22 data (k=0, 1,..., s-1). Because no inter stage data transfer is required, memory consumption is reduced to 1/3 of the original algorithm. Moreover, with the increasing of BUs, not only does throughput increase linearly, system latency also decreases linearly. This array processing orientated architecture provides flexible tradeoff between hardware cost and system performance. An 18-bit word-length 1024-point FFT architecture with 4 BUs is given to demonstrate this mapping algorithm. The design is implemented with TSMC 0.18μm CMOS technology. The core area is 2.99x1.12mm 2 and clock frequency is 326MHz in typical condition (1.8V, 25°C). This processor could complete 1024 FFT calculation in 7.839μs.
Zhenyu Liu 0001, Yang Song 0002, Takeshi Ikenaga, Satoshi Goto
ACM Great Lakes Symposium on VLSI1