EDBT 2026 Demo / reviewers in the wild / expert
Dajiang Zhou
dblp:79/3501
· DBLP profile ↗
55ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0003-2500-8466ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 2 first-author · 4 since 2021Systems, architecture and hardware · 25 · 3 first-authorArtificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
5 papers |
Image and video processing · 57% Multimedia systems and quality of experience · 26% Image and video coding · 17% | |
| Theoretical computer science
1 paper |
Coding theory · 100% | |
| Artificial intelligence
2 papers |
Deep learning architectures and training · 77% Language models and text generation · 23% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Integrated circuit design · 72% Hardware accelerators and domain-specific architectures · 28% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing › video enhancement
compressed video quality enhancement |
1.7 | 2 | 2025 | RivuletMLP: An MLP-based Architecture for Efficient Compressed Video Quality Enhancement · CVPR 2025 Multi-Frame Deformable Look-Up Table for Compressed Video Quality Enhancement · AAAI 2025 |
Machine learning › Deep learning architectures and training › feedforward neural network
multilayer perceptron |
0.9 | 1 | 2025 | RivuletMLP: An MLP-based Architecture for Efficient Compressed Video Quality Enhancement · CVPR 2025 |
Coding theory
source coding |
0.9 | 1 | 2025 | L3TC: Leveraging RWKV for Learned Lossless Low-Complexity Text Compression · AAAI 2025 |
Coding theory › source coding › lossless compression
text compression |
0.9 | 1 | 2025 | L3TC: Leveraging RWKV for Learned Lossless Low-Complexity Text Compression · AAAI 2025 |
Integrated circuit design › digital circuit design
VLSI architecture |
0.5 | 2 | 2017 | Fast Algorithm and VLSI Architecture of Rate Distortion Optimization in H.265/HEVC · IEEE Trans. Multim. 2017 A New Reference Frame Recompression Algorithm and Its VLSI Architecture for UHDTV Video Codec · IEEE Trans. Multim. 2014 |
Image and video coding
rate-distortion optimization |
0.3 | 1 | 2017 | Fast Algorithm and VLSI Architecture of Rate Distortion Optimization in H.265/HEVC · IEEE Trans. Multim. 2017 |
Natural language and speech › Language models and text generation
efficient language model |
0.3 | 1 | 2025 | L3TC: Leveraging RWKV for Learned Lossless Low-Complexity Text Compression · AAAI 2025 |
Hardware accelerators and domain-specific architectures
video coding accelerator |
0.2 | 1 | 2014 | A New Reference Frame Recompression Algorithm and Its VLSI Architecture for UHDTV Video Codec · IEEE Trans. Multim. 2014 |
Image and video processing
motion estimation |
0.1 | 1 | 2012 | An Advanced Hierarchical Motion Estimation Scheme With Lossless Frame Recompression and Early-Level Termination for Beyond High-Definition Video Coding · IEEE Trans. Multim. 2012 |
Image and video coding › video compression › video codec
HEVC |
0.1 | 1 | 2014 | A New Reference Frame Recompression Algorithm and Its VLSI Architecture for UHDTV Video Codec · IEEE Trans. Multim. 2014 |
Methods — techniques the papers use, named apart from their topics
outlier-aware tokenization · 1.7high-rank reparameterization · 1.7entropy coding · 1.7RWKV · 1.7MLP · 1.7temporal alignment · 0.9multi-scale fusion · 0.9lookup table · 0.9convolutional neural network · 0.9quantization-free framework · 0.6hybrid spatial-domain prediction · 0.4semi-fixed-length coding · 0.2semi-fixed length coding · 0.2lossless frame recompression · 0.1hierarchical memory organization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Frame Deformable Look-Up Table for Compressed Video Quality EnhancementabstractThe rapid progress of multimedia technology has led to an increased focus on enhancing the quality of experience (QoE) for video. Specifically, the demand for low-latency and high-quality decoding has grown significantly. Compressed Video Quality Enhancement (CVQE) methods based on Deep Neural Networks (DNNs) have achieved remarkable success. However, most of the methods suffer from high computational complexity, thereby limiting their practicality in low-latency scenarios. Recently, Look-Up Table (LUT) methods have shown great efficiency, which makes them considerably promising in the field of low-latency CVQE. In this paper, we propose an efficient multi-frame deformable Look-Up Table structure for CVQE. Firstly, we design an efficient CNN to explore the inter-frame correlation and then predict the multi-scale convolution offsets. Secondly, we introduce a temporal feature extraction module and a multi-scale fusion module. We first exploit the predicted offsets to guide sampling for precise temporal alignment and extract multi-frame information. Then, higher quality frames are reconstructed from the fused multi-scale features. During the inference, we convert these two modules into LUTs to achieve a sound trade-off between model performance and computational complexity. Experiments demonstrate that our proposed method dramatically outperforms the state-of-the-art LUT-based methods, and obtains competitive performance compared to CNN-based methods with the capability to run in real-time(30fps) at 1080p resolution. Gang He 0002, Guancheng Quan, Chang Wu 0001, Dajiang Zhou, Yunsong Li 0001 |
AAAI | 5 |
| 2025 | L3TC: Leveraging RWKV for Learned Lossless Low-Complexity Text CompressionabstractLearning-based probabilistic models can be combined with an entropy coder for data compression. However, due to the high complexity of learning-based models, their practical application as text compressors has been largely overlooked. To address this issue, our work focuses on a low-complexity design while maintaining compression performance. We introduce a novel Learned Lossless Low-complexity Text Compression method (L3TC). Specifically, we conduct extensive experiments demonstrating that RWKV models achieve the fastest decoding speed with a moderate compression ratio, making it the most suitable backbone for our method. Second, we propose an outlier-aware tokenizer that uses a limited vocabulary to cover frequent tokens while allowing outliers to bypass the prediction and encoding. Third, we propose a novel high-rank reparameterization strategy that enhances the learning capability during training without increasing complexity during inference. Experimental results validate that our method achieves 48% bit saving compared to gzip compressor. Besides, L3TC offers compression performance comparable to other learned compressors, with a 50x reduction in model parameters. More importantly, L3TC is the fastest among all learned compressors, providing real-time decoding speeds up to megabytes per second. Junxuan Zhang, Zhengxue Cheng, Yan Zhao 0041, Dajiang Zhou, Guo Lu, Li Song 0001 |
AAAI | 5 |
| 2025 | RivuletMLP: An MLP-based Architecture for Efficient Compressed Video Quality Enhancement
Gang He 0002, Guancheng Quan, Dajiang Zhou, Yunsong Li 0001 |
CVPR | 5 |
| 2025 | PQNAS: Mixed-precision Quantization-aware Neural Architecture Search with Pseudo QuantizerabstractQuantization-aware neural architecture search is an efficient way to automatically search for the best quantized model that can meet the limited resource constraints on edge devices. Existing methods utilize the straight-through estimator for training the quantized supernet, but lead to oscillation problem of learned weights. To address this issue, we introduce pseudo quantization noise (PQN) in quantization-aware NAS. Accurate range of PQN is vital to ensuring the high accuracy of the quantized networks. However, producing accurate noise for activation quantization during supernet training is challenging, as it requires precise estimation of the quantization parameters for each subnet. Different distributions of different subnets will result in different noise range. To this end, we propose PQNAS, a mixed-precision quantization-aware NAS framework with Pseudo Quantizer. We propose adaptive quantization parameters (AQP) which are trained with the distribution of activation in the pseudo quantizer. With AQP, we can obtain accurate noise range for different subnets during training. Experimentally, comparing with the existing method, the proposed PQNAS achieves 0.99%∼6.48% Top-1 accuracy improvement on ImageNet 1K dataset and 1.1% mAP improvement on COCO dataset. Tianxiao Gao, Li Guo 0006, Shiai Zhu, Dajiang Zhou |
ICASSP | 5 |
| 2023 | A High Resolution SAR Imaging Method for Moving Target Based on Range Doppler and Particle Swarm Optimization AlgorithmabstractSynthetic aperture radar (SAR) imaging for moving target can obtain complete situational awareness information of the detection area, and can realize the monitoring and control for moving target in the region of interest, which has important military and civilian dual-use value. However, due to the complex motion of target, the processing results of the existing SAR imaging methods severly defocused. In this paper, a high resolution SAR imaging method for moving target is proposed. First, we eliminate the coupling induced by linear range cell migration (RCM) by keystone transform. Then, the particle swarm optimization algorithm (PSO) is utilized to estimate the Doppler frequency rate, which can solve the problem of Doppler frequency rate mismatching when azimuth compression. Simulation results verifies the effectiveness of the proposed method. Dajiang Zhou, Hanqing Zhu, Yulin Huang 0001, Yongchao Zhang 0001, Jianyu Yang 0001, Qingying Yi |
IGARSS | 1 |
| 2023 | Moving Target Detection Method for Passive Radar Using LEO Communication Satellite ConstellationabstractIn recent years, many countries are actively deploying Low-Earth-Orbit (LEO) communication satellite constellations, which have the advantages of both high power flux density (PFD) on the surface of the earth and large signal bandwidth. From the perspective of radar application, these new LEO constellations are very suitable as opportunity of illuminator for target detection in passive radar systems. In this paper, the echo signal using LEO communication satellite is analyzed, and a moving target detection method is proposed. Hanqing Zhu, Dajiang Zhou, Zhongyu Li 0001, Hongyang An, Jianyu Yang 0001 |
IGARSS | 2 |
| 2018 | Quad-multiplier packing based on customized floating point for convolutional neural networks on FPGAabstractDeep convolutional neural networks (CNNs) are widely used in many computer vision tasks. Since CNNs involve billions of computations, it is critical to reduce the resource /power consumption and improve parallelism. Compared with extensive researches on fixed point conversion for cost reduction, floating point customization has not been paid enough attention due to its higher cost than fixed point. This paper explores the customized floating point for both the training and inference of CNNs. 9-bit customized floating point is found sufficient for the training of ResNet-20 on CIFAR-10 dataset with less than 1% accuracy loss, which can also be applied to the inference of CNNs. With reduced bit-width, a computational unit (CU) based on Quad-Multiplier Packing is proposed to improve the resource efficiency of CNNs on FPGA. This design can save 87.5% DSP slices and 62.5% LUTs on Xilinx Kintex-7 platform compared to CU using 32-bit floating point. More CUs can be arranged on FPGA and higher throughput can be expected accordingly. Zhifeng Zhang 0003, Dajiang Zhou, Shinji Kimura |
ASP-DAC | 2 |
| 2018 | Embedded Frame Compression for Energy-Efficient Computer Vision SystemsabstractComputer vision applications are rapidly gaining popularity in embedded systems, which typically involve a difficult trade-off between vision performance and energy consumption under a constraint of real-time processing throughput. Recently, hardware (FPGA and ASIC-based) implementations have emerged that significantly improve the energy efficiency of vision computation. These implementations, however, often involve intensive memory traffic that retains a significant portion of energy consumption at the system level. To address this issue, we present a lossy embedded compression framework to exploit the trade-off between vision performance and memory traffic for input images. Differential pulse-code modulation-based gradient-oriented quantization is developed as the lossy compression algorithm. We also present its hardware design that supports up to 12-scale 1080p@60fps real-time processing. For histogram of oriented gradient-based deformable part models on VOC2007, the proposed framework achieved a 49.6%-60.5% memory traffic reduction at a detection rate degradation of 0.05%-0.34%. For AlexNet on ImageNet, memory traffic reduction achieved up to 60.8% with less than 0.61% classification rate degradation. Li Guo 0006, Dajiang Zhou, Jinjia Zhou, Shinji Kimura |
ISCAS | 2 |
| 2018 | Sparseness Ratio Allocation and Neuron Re-pruning for Neural Networks CompressionabstractConvolutional neural networks (CNNs) are rapidly gaining popularity in artificial intelligence applications and employed in mobile devices. However, this is challenging because of the high computational complexity of CNNs and the limited hardware resource in mobile devices. To address this issue, compressing the CNN model is an efficient solution. This work presents a new framework of model compression, with the sparseness ratio allocation (SRA) and the neuron re-pruning (NRP). To achieve a higher overall spareness ratio, SRA is exploited to determine pruned weight percentage for each layer. NRP is performed after the usual weight pruning to further reduce the relative redundant neurons in the meanwhile of guaranteeing the accuracy. From experimental results, with a slight accuracy drop of 0.1%, the proposed framework achieves 149.3× compression on lenet-5. The storage size can be reduced by about 50% relative to previous works. 8-45.2% computational energy and 11.5-48.2% memory traffic energy are saved. Li Guo 0006, Dajiang Zhou, Jinjia Zhou, Shinji Kimura |
ISCAS | 2 |
| 2018 | A Variable-Clock-Cycle-Path VLSI Design of Binary Arithmetic Decoder for H.265/HEVCabstractThe next-generation 8K ultra-high-definition video format involves an extremely high bit rate, which imposes a high throughput requirement on the entropy decoder component of a video decoder. Context adaptive binary arithmetic coding (CABAC) is the entropy coding tool in the latest video coding standards including H.265/High Efficiency Video Coding and H.264/Advanced Video Coding. Due to critical data dependencies at the algorithm level, a CABAC decoder is difficult to be accelerated by simply leveraging parallelism and pipelining. This letter presents a new very-large-scale integration arithmetic decoder, which is the most critical bottleneck in CABAC decoding. Our design features a variable-clock-cycle-path architecture that exploits the differences in critical path delay and in probability of occurrence between various types of binary symbols (bins). The proposed design also incorporates a novel data-forwarding technique (rLPS forwarding) and a fast path-selection technique (coarse bin type decision), and is enhanced with the capability of processing additional bypass bins. As a result, its maximum throughput achieves 1010 Mbins/s in 90-nm CMOS, when decoding 0.96 bin per clock cycle at a maximum clock rate of 1053 MHz, which outperforms previous works by 19.1%. Jinjia Zhou, Dajiang Zhou, Shuping Zhang, Shinji Kimura, Satoshi Goto |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Chain-NN: An energy-efficient 1D chain architecture for accelerating deep convolutional neural networksabstractDeep convolutional neural networks (CNN) have shown their good performances in many computer vision tasks. However, the high computational complexity of CNN involves a huge amount of data movements between the computational processor core and memory hierarchy which occupies the major of the power consumption. This paper presents Chain-NN, a novel energy-efficient 1D chain architecture for accelerating deep CNNs. Chain-NN consists of the dedicated dual-channel process engines (PE). In Chain-NN, convolutions are done by the 1D systolic primitives composed of a group of adjacent PEs. These systolic primitives, together with the proposed column-wise scan input pattern, can fully reuse input operand to reduce the memory bandwidth requirement for energy saving. Moreover, the 1D chain architecture allows the systolic primitives to be easily reconfigured according to specific CNN parameters with fewer design complexity. The synthesis and layout of Chain-NN is under TSMC 28nm process. It costs 3751k logic gates and 352KB on-chip memory. The results show a 576-PE Chain-NN can be scaled up to 700MHz. This achieves a peak throughput of 806.4GOPS with 567.5mW and is able to accelerate the five convolutional layers in AlexNet at a frame rate of 326.2fps. 1421.0GOPS/W power efficiency is at least 2.5 to 4.1x times better than the state-of-the-art works. Dajiang Zhou, Xushen Han, Takeshi Yoshimura |
DATE | 2 |
| 2017 | Measurement-domain intra prediction framework for compressively sensed imagesabstractThis paper presents a measurement-domain intra prediction coding framework that is compatible with compressive sensing (CS) based image sensors. In this framework, we propose a low-complexity intra prediction algorithm that can be directly applied to the measurements captured by the image sensor. Moreover, we propose a structural random 0/1 measurement matrix, embedding the block boundary information that can be extracted from the measurements for intra prediction. Experiment results show that our proposed framework can compress the measurements and increase coding efficiency, with 30% BD-rate reduction compared to the direct output of CS based sensors. This can significantly save both the energy consumption and the bandwidth in communication of wireless camera systems to be massively deployed in the era of IoT. Jian-Bin Zhou, Dajiang Zhou, Li Guo 0006, Takeshi Yoshimura, Satoshi Goto |
ISCAS | 2 |
| 2017 | Approximate-DCT-derived measurement matrices for compressed sensingabstractThe paper presents an algorithm to generate a series of deterministic ternary matrices, which are compatible with the new generation of CMOS image sensor (CIS), Compressed Sensing (CS) based CIS (CS-CIS). The proposed matrices are derived from the approximate DCT, hence preserving the energy compaction property as DCT does. With the proposed 2D-zigzag trimming, the limited measurements could compact most of the energy. The experiment result shows that our proposed matrix contributes a significant improvement in image quality and bitrate saving of 5.7 dB BD-PSNR increase on average comparing with the random binary matrix used in the-state-of-art CS-CIS. Jian-Bin Zhou, Dajiang Zhou, Takeshi Yoshimura, Satoshi Goto |
ISCAS | 2 |
| 2017 | 100x Evolution of Video Codec ChipsabstractIn the past two decades, there has been tremendous progress in video compression technologies. Meanwhile, the use of these technologies, along with the ever-increasing demand for emerging ultra-high-definition applications greatly challenges the design of video codec chips, with the extensive requirements on both memory (DRAM) bandwidth and computation power. Besides, the high data dependencies of video coding algorithms restrict the degree of efficient hardware parallelism and pipelining. This paper describes the techniques to realize high-performance video codec chips. Firstly, we introduce various optimization techniques to solve the DRAM traffic issue. Furthermore, the techniques to reduce the computational complexity and alleviate data dependencies are described. The proposed techniques have been implemented in several ASIC video codecs. Experiments show that the DRAM traffic and DRAM access time are reduced by 80% and 90% respectively. The performance of the video codec chips can achieve [email protected], which is more than 100x better than previous works. Jinjia Zhou, Dajiang Zhou, Satoshi Goto |
ISPD | 2 |
| 2017 | VLSI Implementation of HEVC Motion Compensation With Distance Biased Direct Cache Mapping for 8K UHDTV ApplicationsabstractUltrahigh definition television is becoming increasingly attractive and practical with the doubled compression performance delivered by High Efficiency Video Coding (H.265/HEVC). Meanwhile, implementation of real-time video codecs is challenged by not only the huge throughput and memory bandwidth requirements but also the increased complexity of new algorithms. For motion compensation (MC) that is a known bottleneck in video decoding, the enlarged and diversified prediction unit sizes impose notably higher difficulties in trading off area, power, and memory traffic. This paper presents a very large scale integration implementation of HEVC MC that supports 7680 × 4320@60 frames/s bidirectional prediction. The MC design incorporates a highly efficient cache realized by novel architecture optimizations including distance biased directing mapping, eight-bank memory structure, row-based miss information compression, and mask-based block conflict checking. As a result, the proposed design not only achieves 8× throughput enhancement but also improves hardware efficiency by at least 2.01 times, in comparison with prior arts. Dajiang Zhou, Jian-Bin Zhou, Takeshi Yoshimura, Satoshi Goto |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Fast Algorithm and VLSI Architecture of Rate Distortion Optimization in H.265/HEVCabstractIn H.265/high efficiency video coding (HEVC) encoding, rate distortion optimization (RDO) is an important cost function for mode decision and coding structure decision. Despite being near-optimum in terms of coding efficiency, RDO suffers from a high complexity. To address this problem, this paper presents a fast RDO algorithm and its very large scale implementation (VLSI) for both intra- and inter-frame coding. The proposed algorithm employs a quantization-free framework that significantly reduces the complexity for rate and distortion optimization. Meanwhile, it maintains a low degradation of coding efficiency by taking the syntax element organization and probability model of HEVC into consideration. The algorithm is also designed with hardware architecture in mind to support an efficient VLSI implementation. When implemented in the HEVC test model, the proposed algorithm achieves 62% RDO time reduction with 1.85% coding efficiency loss for the “all-intra” configuration. The hardware implementation achieves 1.6 × higher normalized throughput relative to previous works, and it can support a throughput of 8k@30fps (for four fine-processed modes per prediction unit) with 256 k logic gates when working at 200 MHz. Heming Sun, Dajiang Zhou, Landan Hu, Shinji Kimura, Satoshi Goto |
IEEE Trans. Multim. | 2 |
| 2017 | A Dual-Clock VLSI Design of H.265 Sample Adaptive Offset Estimation for 8k Ultra-HD TV EncodingabstractSample adaptive offset (SAO) is a newly introduced in-loop filtering component in H.265/High Efficiency Video Coding (HEVC). While SAO contributes to a notable coding efficiency improvement, the estimation of SAO parameters dominates the complexity of in-loop filtering in HEVC encoding. This paper presents an efficient VLSI design for SAO estimation. Our design features a dual-clock architecture that processes statistics collection (SC) and parameter decision (PD), the two main functional blocks of SAO estimation, at high- and low-speed clocks, respectively. Such a strategy reduces the overall area by 56% by addressing the heterogeneous data flows of SC and PD. To further improve the area and power efficiency, algorithm-architecture co-optimizations are applied, including a coarse range selection (CRS) and an accumulator bit width reduction (ABR). CRS shrinks the range of fine processed bands for the band offset estimation. ABR further reduces the area by narrowing the accumulators of SC. They together achieve another 25% area reduction. The proposed VLSI design is capable of processing 8k at 120-frames/s encoding. It occupies 51k logic gates, only one-third of the circuit area of the state-of-the-art implementations. Jian-Bin Zhou, Dajiang Zhou, Shuping Zhang, Takeshi Yoshimura, Satoshi Goto |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | CNN-MERP: An FPGA-based memory-efficient reconfigurable processor for forward and backward propagation of convolutional neural networksabstractLarge-scale deep convolutional neural networks (CNNs) are widely used in machine learning applications. While CNNs involve huge complexity, VLSI (ASIC and FPGA) chips that deliver high-density integration of computational resources are regarded as a promising platform for CNN's implementation. At massive parallelism of computational units, however, the external memory bandwidth, which is constrained by the pin count of the VLSI chip, becomes the system bottleneck. Moreover, VLSI solutions are usually regarded as a lack of the flexibility to be reconfigured for the various parameters of CNNs. This paper presents CNN-MERP to address these issues. CNN-MERP incorporates an efficient memory hierarchy that significantly reduces the bandwidth requirements from multiple optimizations including on/off-chip data allocation, data flow optimization and data reuse. The proposed 2-level reconfigurability is utilized to enable fast and efficient reconfiguration, which is based on the control logic and the multiboot feature of FPGA. As a result, an external memory bandwidth requirement of 1.94MB/GFlop is achieved, which is 55% lower than prior arts. Under limited DRAM bandwidth, a system throughput of 1244GFlop/s is achieved at the Vertex UltraScale platform, which is 5.48 times higher than the state-of-the-art FPGA implementations. Xushen Han, Dajiang Zhou, Shinji Kimura |
ICCD | 2 |
| 2015 | A fixed-complexity HEVC inter mode filtering algorithm based on distribution of IME-FME cost ratioabstractMotion Estimation (ME) is a key performance bottleneck of H.265/HEVC encoding. In implementing ME, a fixed or bounded complexity is often desired for efficient pipelining and parallelism. We present a fixed-complexity mode filtering algorithm to reduce the complexity of fractional ME (FME). Firstly, a layer filtering is applied to select the mode candidates. Secondly, the ratio of integer-fractional ME costs (Rf/i) is defined and the confidence interval of Rf/iis utilized for further mode selection. Finally, an algorithm to fix the complexity is applied. Experimental results show that the FME encoding time is reduced by 83.3% with slightly BD-rate increase. Jinjia Zhou, Yizhou Zou, Dajiang Zhou, Satoshi Goto |
ISCAS | 3 |
| 2015 | An independent bandwidth reduction device for HEVC VLSI video systemabstractFRC (frame re-compression) is a kind of widely used technique in reducing the SDRAM (synchronous dynamic random access memory) bandwidth of HEVC video system. However, in previous research works, FRC imposes requirements on accessing pattern and hence its usage are only limited in HEVC video codecs. While in a typical HEVC VLSI video system, there exists many other video IPs with high bandwidth requirements. Therefore, in this article, we propose a new FRC architecture to overcome the limitation and make it applicable to all the video IPs in a HEVC VLSI video system, which raises the overall bandwidth reduction rate of the whole video system. Our proposal has two points: firstly we propose a system internal bus based FRC architecture, which is independent, transparent, and easily connected to all other video IPs. Secondly, we propose a FA (freely access) scheme to remove the requirements on access pattern in previous work. By using this proposal, the bandwidth reduction rate in our VLSI video system model is raised from 92.4% to 69.6%. Jiayi Zhu 0001, Li Guo 0006, Dajiang Zhou, Shinji Kimura, Satoshi Goto |
ISCAS | 3 |
| 2015 | Merge mode based fast inter prediction for HEVCabstractThe latest High Efficiency Video Coding (HEVC/H.265) obtains 50% bit rate reduction than H.264/AVC standard with comparable quality, but at the cost of high computational complexity. Inter prediction accounts for large complexity and merge mode is one of the most important new features introduced in HEVC. To address this issue, this paper utilizes the merge mode to accelerate inter prediction by three fast mode decision methods. 1) A merge candidate decision is proposed to select the best merge mode by Sum of Absolute Transformed Difference (SATD) cost to reduce the merge time. 2) An early merge termination is presented still based on SATD cost with more than 90% accuracy. 3) Based on efficient merge mode, symmetric motion partition (SMP) modes can be disabled for non-8 × 8 code units (CUs). Experimental results demonstrate that our work can achieve 53.1%-54.2% time reduction on average with 1.57%-2.30% BD-rate increment. Besides, our method achieves an improvement of 18%-30% time reduction with 0.89%-2.85% BD-rate increment when combined with other existing approaches. Zhengxue Cheng, Heming Sun, Dajiang Zhou, Shinji Kimura |
VCIP | 3 |
| 2015 | Ultra-High-Throughput VLSI Architecture of H.265/HEVC CABAC Encoder for UHDTV ApplicationsabstractUltra high definition television (UHDTV) imposes extremely high throughput requirement on video encoders based on High Efficiency Video Coding (H.265/HEVC) and Advanced Video Coding (H.264/AVC) standards. Context-adaptive binary arithmetic coding (CABAC) is the entropy coding component of these standards. In very-large-scale integration implementation, CABAC has known difficulties in being effectively pipelined and parallelized, due to the critical bin-to-bin data dependencies in its algorithm. This paper addresses the throughput requirement of CABAC encoding for UHDTV applications. The proposed optimizations including prenormalization, hybrid path coverage and lookahead rLPS to reduce the critical path delay of binary arithmetic encoding (BAE) by exploiting the incompleteness of data dependencies in rLPS updating. Meanwhile, the number of bins BAE delivers per clock cycle is increased by the proposed bypass bin splitting technique. The context modeling and binarization components are also optimized. As a result, our CABAC encoder delivers an average of 4.37 bins per clock cycle. Its maximum clock frequency reaches 420 MHz when synthesized in 90 nm. The corresponding overall throughput is 1836 Mbin/s that is 62.5% higher than the state-of-the-art architecture. Dajiang Zhou, Jinjia Zhou, Wei Fei, Satoshi Goto |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | High-Throughput Power-Efficient VLSI Architecture of Fractional Motion Estimation for Ultra-HD HEVC Video EncodingabstractFractional motion estimation (FME) significantly enhances video compression efficiency, but its high computational complexity also limits the real-time processing capability. In this brief, we present a VLSI implementation of FME design in High Efficiency Video Coding for ultrahigh definition video applications. We first propose a bilinear quarter pixel approximation, together with a search pattern based on it to reduce the complexity of interpolation and fractional search process. Furthermore, a data reuse strategy is exploited to reduce the hardware cost of transform. In addition, using the considered pixel parallelism and dedicated access pattern for memory, we fully pipeline the computation and achieve high hardware utilization. This design has been implemented as a 65-nm CMOS chip and verified. The measured throughput reaches 995 Mpixels/s for 7680 × 4320 30 frames/s at 188 MHz, at least 4.7 times faster than prior arts. The corresponding power dissipation is 198.6 mW, with a power efficiency of 0.2 nJ/pixel. Due to the optimization, our work achieves more than 52% improvement on power efficiency, relative to previous works in H.264. Gang He 0002, Dajiang Zhou, Yunsong Li 0001, Zhixiang Chen 0002, Tianruo Zhang, Satoshi Goto |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | A Frame-Parallel 2 Gpixel/s Video Decoder Chip for UHDTV and 3-DTV/FTV ApplicationsabstractThe first single-chip design that supports real-time H.264/Advanced Video Coding decoding of 8k (7680×4320) 60 frames/s is realized. It also supports multiview decoding for up to 32720p views or 161080p views. To significantly improve the throughput and reduce the memory bandwidth requirement, frame-level parallelism is exploited for the proposed design. First, a frame dependency protection scheme enables frame-parallel decoding, by reusing multiple replicas of an existing design. This results in a system throughput of 2 Gpixels/s, at least 3.75 times better than previous chips. Moreover, a reference window synchronization scheme and a 2-level hybrid caching structure are proposed to achieve 44% memory bandwidth reduction of motion compensation, by utilizing frame-level data reuse. The bandwidth reduction results in 22% Dynamic Random-Access Memory power saving of the whole decoder. Jinjia Zhou, Dajiang Zhou, Jiayi Zhu 0001, Satoshi Goto |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | OpenCL based high-quality HEVC motion estimation on GPUabstractThis paper presents a high quality H.265/HEVC motion estimation implementation with the cooperation of CPU and GPU. The data dependency from MVP (Motion Vector Predictor) restricts the degree of parallelism on GPU. To overcome the constraint from MVP, we propose to use an estimated MVP on GPU and the accurate MVP to refine the motion vector on CPU. GPU fully utilizes its tremendous parallel computing ability without the restriction from MVP. CPU makes up for the deviation from GPU with a small range refinement. Encoding speed benefits from the high degree of parallelism and compression performance is maintained by the CPU refinement. Experimental result shows that the speedup achieves 2.39 times and 32.77 times in the whole ×265 encoder with CPU SIMD (Single Instruction Multiple Data) on and off, respectively. On the other hand, the quality degradation is negligible with only 0.05% increase of BD-rate. Dajiang Zhou, Satoshi Goto |
ICIP | 2 |
| 2014 | A 610 Mbin/s CABAC decoder for H.265/HEVC level 6.1 applicationsabstractThis paper presents a high-throughput decoder of HEVC context-based adaptive binary arithmetic coding (CABAC). A multi-sub-engine arithmetic decoder (MSE-AD) design is proposed to increase the average number of bins delivered per clock cycle by adaptively processing different patterns of upcoming bins with balanced critical path delay. A syntax element (SE) grouping scheme is proposed to maximize the utilization of MSE-AD under the SE parsing order specified in the standard. We also employ a prediction-based pipeline to alleviate the data hazard problem. The proposed CABAC decoder delivers 2.36 bins per clock cycle and achieves a maximum clock frequency of 258MHz in 90nm technology. The resulting performance is 610 Mbin/s which is enough for H.265/HEVC level 6.1 (8K×4K@60fps) applications. Yijin Zhao, Jinjia Zhou, Dajiang Zhou, Satoshi Goto |
ICIP | 3 |
| 2014 | Reducing power consumption of HEVC codec with lossless reference frame recompressionabstractMotion estimation and motion compensation in HEVC and similar video codecs involve huge memory traffic in storing and loading reference frames. The resulting memory power composes a significant portion of system energy consumption. This paper presents a memory power reduction framework that losslessly compresses and decompresses reference frames on-the-fly. We first present the architecture that supports the random access of frame data compressed in variable ratios. The latest recompression algorithms and the corresponding VLSI implementation are also introduced. The framework has been implemented in several UHDTV ASIC and FPGA codecs. Experiments under HEVC common test conditions show an average memory traffic reduction of near 60%. Dajiang Zhou, Li Guo 0006, Jinjia Zhou, Satoshi Goto |
ICIP | 1 |
| 2014 | VLSI architecture of HEVC intra prediction for 8K UHDTV applicationsabstractThis paper presents an efficient VLSI architecture of intra prediction for 8K×4K HEVC decoder. It supports all 35 intra prediction modes and prediction sizes ranging from 4×4 to 64×64. This works proposed a Cyclic SRAM Banks based Parallel Reference Sample Fetching (CSB-PRSF), which guarantees enough reference samples for prediction and reduces the number of registers used for storing reference samples. To guarantee high throughput, 16 pixels are predicted by 4×4 Block Based Pipelining, and dependency between neighboring blocks is eliminated by Hybrid Data Forwarding and Block Reordering. This architecture is synthesized using 90nm technology and the maximum working frequency is 469 MHz, with 72.1K gates area. Running at 397MHz, the architecture can support 4320p@120fps HEVC intra decoding, with full modes and full sizes. Jian-Bin Zhou, Dajiang Zhou, Heming Sun, Satoshi Goto |
ICIP | 2 |
| 2014 | Fast SAO estimation algorithm and its VLSI architectureabstractSAO estimation is the process of determining SAO parameters in video encoding. There are two difficulties for VLSI implementation of SAO estimation. The first is that there are huge amount of samples to deal with in statistic collection phase. The other is that the complexity of RDO in parameters determination phase is very high. In this article, a fast SAO estimation algorithm and its corresponding VLSI architecture are proposed. For the first difficulty, we use bitmaps to collect statistic of all the 16 samples in one 4×4 block simultaneously. For the second difficulty, we simplify a series of complicated procedures in HM to balance the complexity and BD-rate performance. Experimental results show that the proposed algorithm maintains the picture quality improvement. The VLSI design based on this algorithm can be implemented by 156.32K gates, 8832 bits SPRAM, 400MHz @ 65nm technology and is capable of 8Kx4K @ 120fps encoding. Jiayi Zhu 0001, Dajiang Zhou, Shinji Kimura, Satoshi Goto |
ICIP | 2 |
| 2014 | Motion compensation architecture for 8K UHDTV HEVC decoderabstractThis paper presents a motion compensation (MC) architecture for 8K UHDTV HEVC video decoder. UHDTV's high resolution significantly increases throughput and memory traffic. Moreover, HEVC supports new coding tools like various sizes of coding unit ranging from 8 to 64. To solve these problems, we propose three optimization schemes. Firstly, four-bank parallel 2D cache organization is proposed to reduce 61.86% memory traffic and support higher interpolator throughput for HEVC. Secondly, we propose pipelined Write-Through mechanism (WTM) to achieve conflict-free performance. Moreover, WTM scheme contributes to around 50% reduction on both memory area and logic gate. Finally, highly parallel interpolator with proposed cache forms integral structure supporting UHDTV. In 90nm process, our design cost 103.6k logic gates with 12kB cache memory. The proposed architecture can support real-time decoding 7680×4320@30fps at 280MHz. Dajiang Zhou, Satoshi Goto |
ICME | 2 |
| 2014 | Low-Complexity Rate-Distortion Optimization Algorithms for HEVC Intra Prediction
Zhe Sheng, Dajiang Zhou, Heming Sun, Satoshi Goto |
MMM (1) | 2 |
| 2014 | An area-efficient 4/8/16/32-point inverse DCT architecture for UHDTV HEVC decoderabstractThis paper presents a new VLSI architecture for HEVC inverse discrete cosine transform (TDCT). Compared to prior arts, this work reduces hardware cost by: reducing computational logic of 1-D IDCTs with a reordered parallel-in serial-out (RPISO) scheme that shares the inputs of the butterfly structure; and reducing the area of the transpose buffer with a cyclic memory organization that achieves 100% I/O utilization of the SRAMs. In the implementation of a unified 4/8/16/32-point IDCT, the proposed schemes demonstrate 35% and 62% reduction of logic and memory costs, respectively. The IDCT implementation can support real-time decoding of 4K×2K 60fps video with a total hardware cost of 357,250um2on 2-D IDCT and 80,988um2on transpose memory in 90nm process. Heming Sun, Dajiang Zhou, Jiayi Zhu 0001, Shinji Kimura, Satoshi Goto |
VCIP | 2 |
| 2014 | A low power 720p motion estimation processor with 3D stacked memoryabstractIn this paper, a motion estimation processor (MEP) with 3D stacked memory architecture is proposed to 1) reduce the memory and core power consumption; 2) provide higher bandwidth. Firstly, a memory die is designed and staked with MEP die. By adding face-to-face (F2F) pad and through silicon vias (TSV) definitions, 2D electronic design automation (EDA) tools are extended to support the proposed 3D stacking architecture. Moreover, a novel memory controller is applied to control the data transmission and the timing between memory die and MEP die. Finally, 3D physical design is completed for the whole system including TSV/F2F placement, floor plan optimization, power network generation, etc. Comparing with 2D technology, the number of IO pins is reduced by 77%. After optimizing the floor plan of the MEP die and memory die, the routing wire length is reduced by 13.4% and 50% respectively. The simulation results show that the max bandwidth is more than 14GB/s and whole design can support real-time 720p@60fps encoding at 8MHz with less than 65mW, which is only one sixth of the state-of-the-art MEP. Shuping Zhang, Jinjia Zhou, Dajiang Zhou, Satoshi Goto |
VLSI-SoC | 3 |
| 2014 | Alternating asymmetric search range assignment for bidirectional motion estimation in H.265/HEVC and H.264/AVC
Jinjia Zhou, Dajiang Zhou, Satoshi Goto |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | A New Reference Frame Recompression Algorithm and Its VLSI Architecture for UHDTV Video CodecabstractVideo encoders and decoders for HEVC-like compression standards require huge external memory bandwidth, which occupies a significant portion of the codec power consumption. To reduce the memory bandwidth, this paper presents a new lossless reference frame recompression algorithm along with a high-throughput hardware architecture. Firstly, hybrid spatial -domain prediction is proposed to combine the merits of DPCM scanning and averaging. The prediction is then enhanced with multiple modes to accommodate various image characteristics. Finally, efficient residual regrouping based on semi-fixed-length (SFL) coding is used to improve the compression performance. Compared to no compression, the proposed scheme can reduce data traffic by an average of 57.6% with no image quality degradation. The average compression ratio is 2.49, an improvement of at least 12.2-13.2%, relative to the state-of-the-art algorithms. By applying a reordered two-step architecture and the two optimizations, residual reuse and simplified coding mode decision, the hardware cost is similar to that of previous reference frame recompression architectures. The computational complexity increase caused by multi-mode prediction affects the HW cost slightly. This work can be implemented with 45.1 k gates for the compressor and 34.5 k gates for the decompressor at 300 MHz, enough to support a 3840 ×2160@60fps video encoder and decoder. Li Guo 0006, Dajiang Zhou, Satoshi Goto |
IEEE Trans. Multim. | 2 |
| 2014 | High-Performance H.264/AVC Intra-Prediction Architecture for Ultra High Definition Video ApplicationsabstractThis paper presents an H.264/AVC intra-prediction design for ultrahigh definition (ultra-HD) video. Due to the huge throughput requirements of ultra-HD, design challenges such as complexity and data dependency, which currently exist for lower resolutions, become even more critical. To solve these problems, we first propose an interlaced block reordering scheme together with a preliminary mode decision (PMD) strategy to resolve the data dependency between intra mode decision and reconstruction. In the meantime, hardware cost is reduced by PMD. We also propose a probability-based reconstruction scheme to solve the problem of long pipeline latency. In addition, hardware reuse strategies including a shared fine decision module and processing element-reusable prediction generator, are applied to further optimize the design. As a result, the hardware complexity is reduced by 77% in terms of area and frequency, and it takes an average of 33 cycles to process a macroblock. The implementation result demonstrates that our design can support up to the specification of 7680 × 4320p 60 f/s when running at 273 MHz. The design is implemented with 451.5 k gates in 65-nm CMOS. Gang He 0002, Dajiang Zhou, Wei Fei, Zhixiang Chen 0002, Jinjia Zhou, Satoshi Goto |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | A 6.72-Gb/s, 8pJ/bit/iteration WPAN LDPC decoder in 65nm CMOSabstractAn LDPC decoder in 65nm CMOS targeting WPAN (IEEE 802.15.3c) is presented with measurement results. A modified-PCM based message permutation strategy with compatible data flow is proposed to solve the network problem raised by high parallelism LDPC decoding. Compared to the state-of-art, decoder chip achieves 17.7%, 33.5% and 49% improvements in chip density, gate count and energy efficiency, respectively. Zhixiang Chen 0002, Xiao Peng 0002, Xiongxin Zhao, Leona Okamura, Dajiang Zhou, Satoshi Goto |
ASP-DAC | 5 |
| 2013 | A 24.5-53.6pJ/pixel 4320p 60fps H.264/AVC intra-frame video encoder chip in 65nm CMOSabstractAn H.264/AVC intra-frame video encoder is implemented in 65nm CMOS. With an efficient intra prediction design, its maximum throughput reaches 1991Mpixels/s for 7680×4320p 60fps video, 9.4x to 32x faster than previous designs. The encoder also incorporates a 1.41Gbins/s CABAC architecture that has been enhanced by 31%. Moreover, low energy consumption is achieved by the high parallelism and hardware efficiency of this design. 1080p 30fps encoding dissipates only 2mW at 0.8V and 9MHz. Dajiang Zhou, Gang He 0002, Wei Fei, Zhixiang Chen 0002, Jinjia Zhou, Satoshi Goto |
ASP-DAC | 1 |
| 2013 | A high-performance CABAC encoder architecture for HEVC and H.264/AVCabstractThis paper presents a high-performance context adaptive binary arithmetic coding (CABAC) architecture for the next-generation UHDTV applications. Its maximum throughput has been enhanced by 31%~34% with the proposed pre-normalization (prenorm.), hybrid path coverage (HPC), bypass bin splitting (BPBS) and state dual-transition (SDT) schemes. Both the HEVC and H.264/AVC formats can be supported with our architecture by applying a dualstandard binarization design. The proposed CABAC architecture has been silicon proven in a 65nm video encoder chip. It delivers 4.27~4.40 bins/cycle with synthesized and measured clock rates of 401.5MHz and 330MHz, respectively. Therefore a high performance of 1.452Gbin/s is achieved for real-time UHDTV encoding. Jinjia Zhou, Dajiang Zhou, Wei Fei, Satoshi Goto |
ICIP | 2 |
| 2013 | A combined SAO and de-blocking filter architecture for HEVC video decoderabstractThe up-coming video compression standard, high efficiency video coding (HEVC), reduces 50% bit rates in encoding video sequences with same picture quality compared to H.264/AVC. In the in-loop filter (LF) part of HEVC, sample adaptive offset (SAO) is newly added and de-blocking filter (DBF) has been changed a lot. Thus how to construct a high speed and low cost VLSI architecture for HEVC SAO and de-blocking filter is a challenge. In this article, we propose a HEVC LF architecture composed of fully utilized de-blocking filter and SAO. Block based SAO and DBF are employed in this architecture to achieve seamless pipeline between them. The implementation results show that it can be synthesized to 240MHz with 65nm technology. Thus this solution can process 3.84G pixels/s and support 4320p(7680×4320)@120fps decoding. Jiayi Zhu 0001, Dajiang Zhou, Gang He 0002, Satoshi Goto |
ICIP | 2 |
| 2012 | An optimized MC interpolation architecture for HEVCabstractIn the latest draft video compression standard, HEVC, a new 8-tap MC interpolation filter is adopted. For this component, we propose an efficient VLSI design which is composed of a reconfigurable filter, an optimized pipeline engine organization, and a filter reuse scheme. This results in 30% area saving from a non-optimized design. The proposed implementation supports a maximal throughput of QFHD@60fps. Our results also demonstrate the implementation cost of a well optimized HEVC interpolation component can be comparable to that of H.264, despite of the enhanced coding performance. Zhengyan Guo, Dajiang Zhou, Satoshi Goto |
ICASSP | 2 |
| 2012 | Interlaced asymmetric search range assignment for bidirectional motion estimationabstractBidirectional motion estimation significantly enhances video coding efficiency, but its huge complexity is also a critical problem for implementation. This paper presents an interlaced asymmetric search range assignment (IASRA) algorithm. By applying a large and a small search ranges to two reference directions and switching the assignment of these two search ranges once per macroblock, total complexity for bidirectional motion estimation can be reduced by near half with slight coding efficiency drop. IASRA also has the flexibility to be combined with existing fast algorithms and architectures for further complexity saving. We demonstrate this feature by combining IASRA with the state-of-the-art IMNPDR and PMRME architectures, which results in 33% to 42% complexity reduction with less than 1% bit rate increase. Jinjia Zhou, Dajiang Zhou, Satoshi Goto |
ICIP | 2 |
| 2012 | A Low-Complexity HEVC Intra Prediction Algorithm Based on Level and Mode FilteringabstractHEVC achieves a better coding efficiency relative to prior standards, but also involves increased complexity. For intra prediction, complexity is especially intensive due to a highly flexible coding unit structure and a large number of prediction modes. This paper presents a low-complexity intra prediction algorithm for HEVC. A fast preprocessing stage based on a simplified cost model is proposed. Based on its results, a level filtering scheme reduces the number of prediction unit levels that requires fine processing from 5 to 2. To supply level filtering decision with appropriate thresholds, a fast training method is also designed. A mode filtering scheme further reduces the maximum number of angular modes to be evaluated from 34 to 9. Complexity reduction from HM 3.0 is over 50% and stable for various sequences, which makes the proposed algorithm suitable for real-time applications. The corresponding bit rate increase is lower than 2.5%. Heming Sun, Dajiang Zhou, Satoshi Goto |
ICME | 2 |
| 2012 | An Advanced Hierarchical Motion Estimation Scheme With Lossless Frame Recompression and Early-Level Termination for Beyond High-Definition Video CodingabstractIn this paper, we present a hardware-efficient fast algorithm with a lossless frame recompression scheme and early-level termination strategy for large search range (SR) motion estimation (ME) utilized in beyond high-definition video encoder. To achieve high ME quality for hierarchical motion search, we propose an advanced hierarchical ME scheme which processes the multiresolution motion search with an efficient refining stage. This enables high data and hardware reuse for much lower bandwidth and memory cost, while achieving higher ME quality than previous works. In addition, a lossless frame recompression scheme based on this ME algorithm is presented to further reduce bandwidth. A hierarchical memory organization as well as a leveling two-step data fetching strategy is applied to meet constraint of random access for hierarchical motion search structure. Also, the leveling compression strategy by allowing a lower level to refer to a higher one for compression is proposed to efficiently reduce the bandwidth. Furthermore, an early-level termination method suitable for hierarchical ME structure is also applied. This method terminates high-level redundant motion searches by establishing thresholds based on current block mode and motion search level; it also applies the early refinement termination in order to avoid unnecessary refinement for high levels. Experimental results show that the total scheme has a much lower bit rate increasing compared with previous works especially for high motion sequences, while achieving a considerable saving of memory and bandwidth cost for large SR of [-128,127]. Xuena Bao, Dajiang Zhou, Satoshi Goto |
IEEE Trans. Multim. | 2 |
| 2011 | A 98 GMACs/W 32-core vector processor in 65nm CMOS
Dajiang Zhou, Xin Jin 0002, Satoshi Goto |
ISLPED | 2 |
| 2010 | An advanced hierarchical motion estimation scheme with lossless frame recompression for ultra high definition video codingabstractIn this paper, a hardware-efficient fast algorithm with a lossless frame recompression scheme for large search range (SR) motion estimation (ME) utilized in ultra high definition video encoder is presented. To ensure high ME quality for hierarchical motion search, the advanced hierarchical ME scheme which processes the multi-resolution search with a coarse-to-fine strategy in parallel is proposed. This enables high data and hardware reuse for much lower bandwidth and memory cost, which at the same time maintains ME quality. Furthermore, a lossless frame recompression scheme based on this ME algorithm to further reduce DRAM bandwidth is presented. An efficient leveling memory organization as well as a leveling two-step data fetching strategy is proposed to meet constrains of random access for hierarchical motion search structure. Experimental results show that the total scheme increases PSNR by an average of 0.03dB with a much less bit rate increasing compared with previous works especially for high motion sequences, while achieving a considerable saving of memory and bandwidth cost for large SR of [-128, 127]. Xuena Bao, Dajiang Zhou, Satoshi Goto |
ICME | 2 |
| 2010 | A lossless frame recompression scheme for reducing DRAM power in video encodingabstractOwing to the huge bandwidth requirement, the DRAM power composes a significant portion of the power consumed by video coding system. In this paper, a lossless frame recompression scheme for reducing DRAM bandwidth is presented. The basic structure of the scheme which includes the memory organization, the data fetching strategy as well as the cache organization is proposed. Furthermore, an adaptive DPCM (Differential Pulse Code Modulation) scanning order selection is used and an efficient coding method that is suitable for compressing the DPCM samples in reference frames is also discussed. Experimental results show a 50%~60% saving of bandwidth on 720p and 1080p sequences, which indicates that the proposed scheme can be useful in reducing the system power through saving DRAM bandwidth. Xuena Bao, Dajiang Zhou, Satoshi Goto |
ISCAS | 2 |
| 2010 | An early stopping criterion for decoding LDPC codes in WiMAX and WiFi standardsabstractBased on the particular structure of parity check matrix (PCM) of LDPC codes in WiMAX and WiFi wireless communication standards, a new early stopping criterion (SC) is proposed to save the unnecessary decoding iterations. By using the proposed SC, decoding process will be stopped if all the information bits in a code word are corrected even if there are still some errors in the redundant (parity) part. Simulation shows that our proposed SC outperforms previous ones on decoding speed evaluated by average iteration number (AIN) with no bit error rate (BER) performance loss. Zhixiang Chen 0002, Xiongxin Zhao, Xiao Peng 0002, Dajiang Zhou, Satoshi Goto |
ISCAS | 4 |
| 2010 | An adaptive bandwidth reduction scheme for video codingabstractIn high definition video decoders for the standard like H.264, bandwidth requirement is a critical design issue due to the overwhelming amount of memory data access. This paper proposes a new adaptive reference frame compression scheme to reduce external memory bandwidth consumption. An efficient variable length coding is proposed to compress each processing unit efficiently. Moreover, an adaptive compression mode decision unit is proposed to adaptively choose the best compression mode according to the image characteristic and bandwidth requirement. As a result, the proposed scheme achieves efficient bandwidth reduction with little image quality loss. The experimental results show that, at the same bandwidth reduction ratio the proposed algorithm achieves up-to 0.7 dB gain compared to the existing approaches. Liu Song, Dajiang Zhou, Xin Jin 0002, Satoshi Goto |
ISCAS | 2 |
| 2010 | Intra prediction architecture for H.264/AVC QFHD encoderabstractThis paper proposes a high-performance intra prediction architecture that can support H.264/AVC high profile. The proposed MB/block co-reordering can avoid data dependency and improve pipeline utilization. Therefore, the timing constraint of real-time 4k×2k encoding can be achieved with negligible quality loss. 16×16 prediction engine and 8×8 prediction engine work parallel for prediction and coefficients generating. A reordering interlaced reconstruction is also designed for fully pipelined architecture. It takes only 160 cycles to process one macroblock (MB). Hardware utilization of prediction and reconstruction modules is almost 100%. Furthermore, PE-reusable 8×8 intra predictor and hybrid SAD & SATD mode decision are proposed to save hardware cost. The design is implemented by 90nm CMOS technology with 113.2k gates and can encode 4k×2k video sequences at 60 fps with operation frequency of 310MHz. Gang He 0002, Dajiang Zhou, Jinjia Zhou, Satoshi Goto |
PCS | 2 |
| 2009 | A 136 cycles/MB, luma-chroma parallelized H.264/AVC deblocking filter for QFHD applicationsabstractIn this paper, we present a high-throughput deblocking filter architecture for H.264/AVC in QFHD applications. In order to enhance the parallelism of filtering without notably increasing the area, we propose to parallelize the processing of luminance and chrominance samples, instead of simultaneously filtering two edges of a same component. Although the edge filter and transpose cost of the proposed architecture is a little larger than that of the single-filter solution, control logic is saved by applying an identical processing schedule to both the luminance and chrominance samples. Meanwhile, total SRAM size by bit is kept unchanged when the architecture is parallelized. As a result, throughput of this work is advanced by 50% (or processing time reduced by 33%), to be 136 cycles/MB, while area cost (17.9 k gates logic and 8 k bits SRAM) is kept comparable to the state-of-the-art works. Jinjia Zhou, Dajiang Zhou, Satoshi Goto |
ICME | 2 |
| 2009 | Block-pipelining Cache for Motion Compensation in High Definition H.264/AVC Video DecoderabstractIn this paper, we present a cache scheme targeting hardware implementation to reduce the bandwidth of motion compensation, and a block-pipelining strategy to hide long latency of the external memory in high definition H.264/AVC video decoder. Hardware architecture is also implemented for the proposed algorithms. Experimental results show that the cache succeeds in reducing external memory bandwidth of motion compensation by 66%∼78% and the block-pipelining strategy can solve the latency problem better than previous solutions. Our proposed hardware architecture can averagely process one macroblock within 297 cycles, capable of real-time processing 1920×1088@30fps H.264 sequence at lower than 80MHz. Xianmin Chen, Jiayi Zhu 0001, Dajiang Zhou, Satoshi Goto |
ISCAS | 4 |
| 2009 | Prioritized Reference Decision for Efficient Motion Vector CodingabstractIn the latest video coding frameworks, efficiency of motion vector (MV) coding is becoming increasingly important because of the growing bit rate portion of motion information. In this paper, we present a prioritized reference decision scheme for MV coding based on the H.264/AVC framework. This scheme makes use of a boolean indicator coded into the modified code words of mb_type and sub_mb_type, so as to designate whether the conventional median prediction or an alternative reference decision method is to be used. If the latter is indicated, a validation function is applied to a prioritized list of neighboring MVs, so that the minimum-cost MV difference (MVD) is likely to be generated using multiple reference MVs, with no extra overhead. Experimental result shows the proposed scheme achieves considerable bit rate reduction on MVDs, when compared with the conventional median prediction method. It also achieves a better and much stabler performance than the state-of-the-art MBP-based MV coding. Meanwhile, the harmful feedback issue of MBP is prevented in the proposed scheme. Dajiang Zhou, Jinjia Zhou, Satoshi Goto |
ISCAS | 1 |
| 2008 | An SDRAM controller optimized for high definition video coding applicationabstractThe huge SDRAM bandwidth requirement is an architectural bottleneck of video decoders. Besides the large amounts of data transmission cycles, many extra overhead (up to over 50%) are incurred by the ACTIVE or PRECHARGE (A\P) operations in conventional SDRAM controllers. In this paper, we propose an optimized SDRAM controller, which can improve the bandwidth efficiency by eliminating most of the extra overhead. An access management scheme that enables consecutive data transmission is employed to reduce the need of A\P operations. In addition, a scheduler is designed to hide the latency for A\P operations. Experimental result shows that the SDRAM controller is able to meet the bandwidth requirement of real-time decoding 1080P H.264 video streams, at less than 100Mhz with 32-bit DDR SDRAM. Jiayi Zhu 0001, Dajiang Zhou |
ISCAS | 3 |
| 2007 | A Hardware-Efficient Dual-Standard VLSI Architecture for MC Interpolation in AVS and H.264abstractH.264 and AVS are the two latest video coding standards. Since the similarity between their structures, it is feasible to develop a dual-mode VLSI decoder for supporting both standards, with substantially less cost than the solution with two individual decoders. In this paper, we propose a dual-standard VLSI architecture for MC interpolation, which is the most calculation intensive module of the dual-mode decoder. By applying reconfigurable FIR filters and an adaptive pipeline strategy, an implementation of the architecture can process realtime video streams in 1280times720, 30fps at low cost (11.5k gates, no RAM). This design also provides scalability to meet higher performance requirements. Dajiang Zhou |
ISCAS | 1 |