Tian-Sheuan Chang

dblp:40/6848 · also Tian Sheuan Chang · DBLP profile ↗
← Back
95ranked-venue papers
2as first author
25since 2021 · last 2026
0000-0002-0561-8745ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 57 · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RCW-CIM: A Digital CIM-based LLM Accelerator with Read-Compute/Write
Yan-Cheng Guo, Tian-Sheuan Chang, Jian-Wei Su
ISCAS2
2026 VitaLLM: A Versatile and Tiny Accelerator for Mixed-Precision LLM Inference on Edge Devices
Zi-Wei Lin, Tian-Sheuan Chang
ISCAS2
2026 Low-Power Vision Transformer Accelerator With Hardware-Aware Pruning and Optimized Dataflow
abstract
Current transformer accelerators primarily focus on optimizing self-attention due to its quadratic complexity. However, this focus is less relevant for vision transformers with short token lengths, where the Feed-Forward Network (FFN) tends to be the dominant computational bottleneck. This paper presents a low power Vision Transformer accelerator, optimized through algorithm-hardware co-design. The model complexity is reduced using hardware-friendly dynamic token pruning without introducing complex mechanisms. Sparsity is further improved by replacing GELU with ReLU activations and employing dynamic FFN2 pruning, achieving a 61.5% reduction in operations and a 59.3% reduction in FFN2 weights, with an accuracy loss of less than 2%. The hardware adopts a row-wise dataflow with output-oriented data access to eliminate data transposition, and supports dynamic operations with minimal area overhead. Implemented in TSMC’s 28nm CMOS technology, our design occupies 496.4K gates and includes a 232KB SRAM buffer, achieving a peak throughput of 1024 GOPS at 1GHz, with an energy efficiency of 2.31 TOPS/W and an area efficiency of 858.61 GOPS/mm2.
Ching-Lin Hsiung, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. I Regul. Pap.2
2026 A PVT-Resilient Subthreshold SRAM-Based In-Memory Computing Accelerator With In-Situ Regulation for Energy-Efficient Spiking Neural Networks
abstract
This paper presents a PVT-resilient, subthreshold SRAM-based computing-in-memory (CIM) macro tailored for energy-efficient spiking neural networks (SNNs). The macro integrates in-situ current sensors and distributed voltage regulators to enable robust large-scale (1024 wordlines, 1304 bitlines and 128 shared neuron cells) subthreshold current-mode CIM, mitigating energy overheads and process-voltage-temperature (PVT) sensitivity. The neuron cells adopt a programmable, memory cell-based firing threshold to enhance neuron robustness against PVT variations. The architecture uses a stride-tick batching schedule to significantly reduce buffer overhead with enhanced input data reuse. Exploiting the high sparsity of SNNs, the proposed system demonstrates significant improvements in energy efficiency and variation tolerance. Fabricated in 28-nm CMOS, the prototype attains 93.64\% accuracy on keyword spotting, delivers up to 1181.42 TOPS/W, and achieves 7.24 TOPS/mm^2, demonstrating a viable and efficient solution for high-performance edge SNN processing.
Shih-Hang Kao, Yang-Chan Hung, I-Wen Wang, Bing-Han Liu, Yu-Chia Chen, Tian-Sheuan Chang, Shyh-Jye Jou, Chien-Nan Jimmy Liu, Hung-Ming Chen, Wei-Zen Chen
IEEE Trans. Circuits Syst. I Regul. Pap.6
2026 A 129 FPS Full HD Real-Time Accelerator for 3D Gaussian Splatting
abstract
Rendering large-scale, unbounded scenes on AR/VR-class devices is constrained by the computation, bandwidth, and storage cost of 3D Gaussian Splatting (3DGS). We propose a low-power, low-cost 3DGS hardware accelerator that renders full-HD images in real time, together with a hardware-friendly compression pipeline that combines iterative Gaussian pruning and fine-tuning, progressive spherical harmonics (SH) degree reduction, and vector quantization of all SH coefficients and colors. The scheme achieves a $51.6\times$51.6× model-size reduction with a 0.743 dB PSNR loss. The accelerator uses a frame-level pipeline that integrates point-based culling and projection with tile-based sorting and rasterization, skips zero-Jacobian matrix multiplications (reducing processing elements by 63% and computation by 53%), and adopts comparison-free tile-based sorting with deterministic latency. Implemented in a TSMC 28-nm process at 800MHz, the design occupies $\text{0.66}\;\text{mm}^{2}$0.66mm2 with 1.1438 M gates and 120 kB SRAM, consumes 0.219 W, and delivers 1219 Mpixels/J at 267.5 Mpixels/s, enabling 1080p at 129 FPS. Overall, it is $5.98\times$5.98× smaller in area, $5.94\times$5.94× higher throughput, and delivers $7.5\times$7.5× higher energy efficiency than prior 3DGS accelerators.
Fang-Chi Chang, Tian-Sheuan Chang
IEEE Trans. Vis. Comput. Graph.2
2025 Hardware Efficient Accelerator for Spiking Transformer With Reconfigurable Parallel Time Step Computing
abstract
This paper introduces the first low-power hardware accelerator for Spiking Transformers, an emerging alternative to traditional artificial neural networks. By modifying the base Spikformer model to use IAND instead of residual addition, the model exclusively utilizes spike computation. The hardware employs a fully parallel tick-batching dataflow and a time-step reconfigurable neuron architecture, addressing the delay and power challenges of multi-timestep processing in spiking neural networks. This approach processes outputs from all time steps in parallel, reducing computation delay and eliminating membrane memory, thereby lowering energy consumption. The accelerator supports 3x3 and 1x1 convolutions and matrix operations through vectorized processing, meeting model requirements. Implemented in TSMC’s 28nm process, it achieves 3.456 TSOPS (tera spike operations per second) with a power efficiency of 38.334 TSOPS/W at 500MHz, using 198.46K logic gates and 139.25KB of SRAM.
Bo-Yu Chen, Tian-Sheuan Chang
ISCAS2
2025 A Low-Power Sparse Deep Learning Accelerator with Optimized Data Reuse
abstract
Sparse deep learning has significantly reduced computational costs; however, its irregular distribution of non-zero data complicates data flow, limits data reuse, and increases on-chip SRAM access, thereby elevating chip power consumption. To address these challenges, this work maximizes data reuse to minimize SRAM access through two approaches. First, we propose Effective Index Matching (EIM), which efficiently searches and arranges non-zero operations from compressed data. Second, we propose Shared Index Data Reuse (SIDR), which coordinates operations across Processing Elements (PEs) to regularize their SRAM data access, thereby enabling all data to be reused efficiently. Our approach reduces the access of the SRAM buffer by 86% when compared to the previous design, SparTen. As a result, our design achieves a 2.5× improvement in power efficiency compared to state-of-the-art methods while maintaining a simpler dataflow.
Kai-Chieh Hsu, Tian-Sheuan Chang
ISCAS2
2025 An Efficient Data Reuse with Tile-Based Adaptive Stationary for Transformer Accelerators
abstract
Transformer-based models have become the de facto backbone across many fields, such as computer vision and natural language processing. However, as these models scale in size, external memory access (EMA) for weight and activations becomes a critical bottleneck due to its significantly higher energy consumption compared to internal computations. While most prior work has focused on optimizing the self-attention mechanism, little attention has been given to optimizing data transfer during linear projections, where EMA costs are equally important. In this paper, we propose the Tile-based Adaptive Stationary (TAS) scheme that selects the input or weight stationary in a tile granularity, based on the input sequence length. Our experimental results demonstrate that TAS can significantly reduce EMA by more than 97% compared to traditional stationary schemes, while being compatible with various attention optimization techniques and hardware accelerators.
Tseng-Jen Li, Tian-Sheuan Chang
ISCAS2
2024 A Multi-Bit Near-RRAM based Computing Macro with Highly Computing Parallelism for CNN Application
abstract
Resistive random-access memory (RRAM) based compute-in-memory (CIM) is an emerging approach to address the demand for practical implementation of artificial intelligence (AI) on resource constrained edge devices by reducing the power-hungry data transfer between memory and processing unit. However, the state-of-the-art RRAM CIM designs fail to strike a balance between precision, energy efficiency, throughput, and latency. This work merges the techniques of CIM and compute-near-memory (CNM) to deliver high precision, high energy efficiency, high throughput, and low latency. In this paper, a 256Kb RRAM based CNM macro fabricated in TSMC 40 nm process is presented featuring: 1) opposite weight mapping with variation-robust SA to mitigate the impact of RRAM device variations on MAC (Multiply-Accumulate) results; 2) switched-capacitor-based analog multiplication circuit to achieve highly parallel computing of 128 4-bit by 4-bit MAC result with low power consumption and high operation speed; and 3) joint optimization of hardware and software to compensate for the accuracy loss after considering the non-idealities of circuits. The macro achieves a low latency of 17ns and high energy efficiency of 71 TOPS/W for MAC operations with 4-bit input, 4-bit weight and 4-bit output precision. It is used to accelerate the convolution process in the Light-CSPDenseN et AI model, resulting in a high accuracy of 86.33% on Visual Wake Words dataset.
Kuan-Chih Lin, Hao Zuo, Hsiang-Yu Wang, Yuan-Ping Huang, Ci-Hao Wu, Yan-Cheng Guo, Shyh-Jye Jou, Tuo-Hung Hou, Tian-Sheuan Chang
DATE9
2024 CIMR-V: An End-to-End SRAM-based CIM Accelerator with RISC-V for AI Edge Device
abstract
Computing-in-memory (CIM) is renowned in deep learning due to its high energy efficiency resulting from highly parallel computing with minimal data movement. However, current SRAM-based CIM designs suffer from long latency for loading weight or feature maps from DRAM for large AI models. Moreover, previous SRAM-based CIM architectures lack end-to-end model inference. To address these issues, this paper proposes CIMR-V, an end-to-end CIM accelerator with RISC-V that incorporates CIM layer fusion, convolution/max pooling pipeline, and weight fusion, resulting in an 85.14% reduction in latency for the keyword spotting model. Furthermore, the proposed CIM-type instructions facilitate end-to-end AI model inference and full stack flow, effectively synergizing the high energy efficiency of CIM and the high programmability of RISC-V. Implemented using TSMC 28nm technology, the proposed design achieves an energy efficiency of 3707.84 TOPS/W and 26.21 TOPS at 50 MHz.
Yan-Cheng Guo, Tian-Sheuan Chang, Chih-Sheng Lin, Bo-Cheng Chiou, Chih-Ming Lai, Shyh-Shyuan Sheu, Wei-Chung Lo, Shih-Chieh Chang 0001
ISCAS2
2024 Dynamic Gradient Sparse Update for Edge Training
abstract
Training on edge devices enables personalized model fine-tuning to enhance real-world performance and maintain data privacy. However, the gradient computation for backpropagation in the training requires significant memory buffers to store intermediate features and compute losses. This is unacceptable for memory-constrained edge devices such as microcontrollers. To tackle this issue, we propose a training acceleration method using dynamic gradient sparse updates. This method updates the important channels and layers only and skips gradient computation for the less important channels and layers to reduce memory usage for each update iteration. In addition, the channel selection is dynamic for different iterations to traverse most of the parameters in the update layers along the time dimension for better performance. The experimental result shows that the proposed method enables an ImageNet pre-trained MobileNetV2 trained on CIFAR-10 to achieve an accuracy of 85.77% while updating only 2% of convolution weights within 256KB on-chip memory. This results in a remarkable 98% reduction in feature memory usage compared to dense model training.
I-Hsuan Li, Tian-Sheuan Chang
ISCAS2
2024 ESSR: An 8K@30FPS Super-Resolution Accelerator With Edge Selective Network
abstract
Deep learning-based super-resolution (SR) is challenging to implement in resource-constrained edge devices for resolutions beyond full HD due to its high computational complexity and memory bandwidth requirements. This paper introduces an 8K@30FPS SR accelerator with edge-selective dynamic input processing. Dynamic processing chooses the appropriate subnets for different patches based on simple input edge criteria, achieving a 50% MAC reduction with only a 0.1dB PSNR decrease. The quality of reconstruction images is guaranteed and maximized its potential withresource adaptive model switchingeven under resource constraints. In conjunction with hardware-specific refinements, the model size is reduced by 84% to 51K, but with a decrease of less than 0.6dB PSNR. Additionally, to support dynamic processing with high utilization, this design incorporates aconfigurable group of layer mappingthat synergizes with thestructure-friendly fusion block, resulting in 77% hardware utilization and up to 79% reduction in feature SRAM access. The implementation, using the TSMC 28nm process, can achieve 8K@30FPS throughput at 800MHz with a gate count of 2749K, 0.2075W power consumption, and 4797Mpixels/J energy efficiency, exceeding previous work.
Chih-Chia Hsu, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 ACNPU: A 4.75TOPS/W 1080P@30FPS Super Resolution Accelerator With Decoupled Asymmetric Convolution
abstract
Deep learning-driven superresolution (SR) outperforms traditional techniques but also faces the challenge of high complexity and memory bandwidth. This challenge leads many accelerators to opt for simpler and shallow models like FSRCNN, compromising performance for real-time needs, especially for resource-limited edge devices. This paper proposes an energy-efficient SR accelerator, ACNPU, to tackle this challenge. The ACNPU enhances image quality by 0.34dB with a 27-layer model, but needs 36% less complexity than FSRCNN, while maintaining a similar model size, with the decoupled asymmetric convolution and split-bypass structure. The hardware-friendly 17K-parameter model enables holistic model fusion instead of localized layer fusion to remove external DRAM access of intermediate feature maps. The on-chip memory bandwidth is further reduced with the input stationary flow and parallel-layer execution to reduce power consumption. Hardware is regular and easy to control to support different layers by processing elements (PEs) clusters with reconfigurable input and uniform data flow. The implementation in the 40 nm CMOS process consumes 2333 K gate counts and 198Â KB SRAMs. The ACNPU achieves 31.7 FPS and 124.4 FPS for$\times 2$and$\times 4$scales Full-HD generation, respectively, which attains 4.75 TOPS/W energy efficiency.
Tun-Hao Yang, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 A 71.2-μW Speech Recognition Accelerator With Recurrent Spiking Neural Network
abstract
This paper introduces a 71.2-$\mu$W speech recognition accelerator designed for edge devices’ real-time applications, emphasizing an ultra low power design. Achieved through algorithm and hardware co-optimizations, we propose a compact recurrent spiking neural network with two recurrent layers, one fully connected layer, and a low time step (1 or 2). The 2.79-MB model undergoes pruning and 4-bit fixed-point quantization, shrinking it by 96.42% to 0.1 MB. On the hardware front, we take advantage ofmixed-level pruning,zero-skippingandmerged spiketechniques, reducing complexity by 90.49% to 13.86 MMAC/S. Theparallel time-step executionaddresses inter-time-step data dependencies and enables weight buffer power savings through weight sharing. Capitalizing on the sparse spike activity, an input broadcasting scheme eliminates zero computations, further saving power. Implemented on the TSMC 28-nm process, the design operates in real time at 100 kHz, consuming 71.2$\mu$W, surpassing state-of-the-art designs. At 500 MHz, it has 28.41 TOPS/W and 1903.11 GOPS/mm$^2$in energy and area efficiency, respectively.
Chih-Chyau Yang, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 ASC: Adaptive Scale Feature Map Compression for Deep Neural Network
abstract
Deep-learning accelerators are increasingly in demand; however, their performance is constrained by the size of the feature map, leading to high bandwidth requirements and large buffer sizes. We propose an adaptive scale feature map compression technique leveraging the unique properties of the feature map. This technique adopts independent channel indexing given the weak channel correlation and utilizes a cubical-like block shape to benefit from strong local correlations. The method further optimizes compression using a switchable endpoint mode and adaptive scale interpolation to handle unimodal data distributions, both with and without outliers. This results in$4\times $and up to$7.69\times $compression rates for 16-bit data in constant and variable bitrates, respectively. Our hardware design minimizes area cost by adjusting interpolation scales, which facilitates hardware sharing among interpolation points. Additionally, we introduce a threshold concept for straightforward interpolation, preventing the need for intricate hardware. The TSMC 28nm implementation showcases an equivalent gate count of 6135 for the 8-bit version. Furthermore, the hardware architecture scales effectively, with only a sublinear increase in area cost. Achieving a$32\times $throughput increase meets the theoretical bandwidth of DDR5-6400 at just$7.65\times $the hardware cost.
Tian-Sheuan Chang
IEEE Trans. Circuits Syst. I Regul. Pap.2
2023 FPCIM: A Fully-Parallel Robust ReRAM CIM Processor for Edge AI Devices
abstract
Computing-in-memory (CIM) is popular for deep learning due to its high energy efficiency owing to massive parallelism and low data movement. However, current ReRAM based CIM designs only use partial parallelism since fully parallel CIM could suffer lower model accuracy due to severe nonideal effects. This paper proposes a robust fully-parallel ReRAM-based CIM processor for deep learning. The proposed design exploits the fully-parallel computation of a$1024\mathrm{x}1024$array to achieve 110.59 TOPS and reduces nonideal effects with in-ReRAM computing (IRC) training and hybrid digital/IRC design to minimize the accuracy loss with only 1.55%. This design is programmable with a compact CIM-oriented instruction set to support various 2-D convolution neural networks (NN) as well as hybrid digital/IRC designs. The final implementation achieves a 2740.41 TOPS/W energy efficiency at 125MHz with TSMC 40nm technology, which is superior to previous designs.
Yan-Cheng Guo, Wei-Tien Lin, Tuo-Hung Hou, Tian-Sheuan Chang
ISCAS4
2023 Memory Bandwidth Efficient Design for Super-Resolution Accelerators With Structure Adaptive Fusion and Channel-Aware Addressing
abstract
State-of-the-art (SOTA) super-resolution (SR) models can generate high-quality images. However, they require a large external memory bandwidth, making it impossible to implement these models on hardware. Although some work has presented different kinds of layer fusion to reduce memory traffic, they can only work on simple model architectures and only consider the feature extraction part. To solve the above issues, this article proposes structure adaptive fusion (SAF) for the feature extraction part to avoid intermediate feature map I/O. This method selects the repetitive structure as the fusion unit and fuses multiple ones to meet buffer size and memory bandwidth constraints, which can deal with different SR models. In addition, we also propose channel-aware addressing for the upscale part to avoid off-chip data transfers. The proposed methods achieve over 90% of memory traffic reduction in all tested SOTA models. Compared to the SOTA fusion method, our approach requires a 52% smaller buffer size and up to 61% lower memory bandwidth for the same number of fusion layers.
An-Jung Huang, Jo-Hsuan Hung, Tian-Sheuan Chang
IEEE Trans. Very Large Scale Integr. Syst.3
2023 A 1.6-mW Sparse Deep Learning Accelerator for Speech Separation
abstract
Low-power deep learning accelerators (DLAs) on the speech processing enable real-time applications on edge devices. However, most of the existing accelerators suffer from high-power consumption and focus on image applications only. This article presents a low-power accelerator for speech separation through algorithm and hardware optimizations. At the algorithm level, the model is compressed with structured sensitivity as well as unstructured pruning, and further quantized to the shifted 8-bit floating-point format instead of the 32-bit floating-point format. The computations with the zero kernel and zero activation values are skipped by decomposition of the dilated and transposed convolutions. At the hardware level, the compressed model is then supported by an architecture with eight independent multipliers and accumulators (MACs) with a simple zero-skipping hardware to take advantage of the activation sparsity and low-power processing. The proposed approach reduces the model size by 95.44% and computation complexity by 93.88%. The final implementation with the TSMC 40-nm process can achieve real-time speech separation and consumes 1.6-mW power when operated at 150 MHz. The normalized energy efficiency and area efficiency are 2.344 TOPS/W and 14.42 GOPS/mm2, respectively.
Chih-Chyau Yang, Tian-Sheuan Chang
IEEE Trans. Very Large Scale Integr. Syst.2
2022 A Real Time Super Resolution Accelerator with Tilted Layer Fusion
abstract
Deep learning based superresolution achieves high-quality results, but its heavy computational workload, large buffer, and high external memory bandwidth inhibit its usage in mobile devices. To solve the above issues, this paper proposes a real-time hardware accelerator with the tilted layer fusion method that reduces the external DRAM bandwidth by 92% and just needs 102KB on-chip memory. The design implemented with a 40nm CMOS process achieves 1920xl080@60fps throughput with 544. 3K gate count when running at 600MHz; it has higher throughput and lower area cost than previous designs.
An-Jung Huang, Kai-Chieh Hsu, Tian-Sheuan Chang
ISCAS3
2022 PSCNN: A 885.86 TOPS/W Programmable SRAM-based Computing-In-Memory Processor for Keyword Spotting
abstract
Computing-in-memory (CIM) has attracted significant attentions in recent years due to its massive parallelism and low power consumption. However, current CIM designs suffer from large area overhead of small CIM macros and bad programmablity for model execution. This paper proposes a programmable CIM processor with a single large sized CIM macro instead of multiple smaller ones for power efficient computation and a flexible instruction set to support various binary 1-D convolution Neural Network (CNN) models in an easy way. Furthermore, the proposed architecture adopts the pooling write-back method to support fused or independent convolution/pooling operations to reduce 35.9% of latency, and the flexible ping-pong feature SRAM to fit different feature map sizes during layer-by-layer execution. The design fabricated in TSMC 28nm technology achieves 150.8 GOPS throughput and 885.86 TOPS/W power efficiency at 10 MHz when executing our binary keyword spotting model, which has higher power efficiency and flexibility than previous designs.
Shu-Hung Kuo, Tian-Sheuan Chang
ISCAS2
2022 BSRA: Block-based Super Resolution Accelerator with Hardware Efficient Pixel Attention
abstract
Increasingly, convolution neural network (CNN) based super resolution models have been proposed for better reconstruction results, but their large model size and complicated structure inhibit their real-time hardware implementation. Current hardware designs are limited to a plain network and suffer from lower quality and high memory bandwidth requirements. This paper proposes a super resolution hardware accelerator with hardware efficient pixel attention that just needs 25. 9K parameters and simple structure but achieves 0. 38dB better reconstruction images than the widely used FSRCNN. The accelerator adopts full model block wise convolution for full model layer fusion to reduce external memory access to model input and output only. In addition, CNN and pixel attention are well supported by PE arrays with distributed weights. The final implementation can support full HD image reconstruction at 30 frames per second with TSMC 40nm CMOS process.
Dun-Hao Yang, Tian-Sheuan Chang
ISCAS2
2022 Sparse Compressed Spiking Neural Network Accelerator for Object Detection
abstract
Spiking neural networks (SNNs), which are inspired by the human brain, have recently gained popularity due to their relatively simple and low-power hardware for transmitting binary spikes and highly sparse activation maps. However, because SNNs contain extra time dimension information, the SNN accelerator will require more buffers and take longer to infer, especially for the more difficult high-resolution object detection task. As a result, this paper proposes a sparse compressed spiking neural network accelerator that takes advantage of the high sparsity of activation maps and weights by utilizing the proposed gated one-to-all product for low power and highly parallel model execution. The experimental result of the neural network shows 71.5% mAP with mixed (1,3) time steps on the IVS 3cls dataset. The accelerator with the TSMC 28nm CMOS process can achieve$1024\times 576.29$frames per second processing when running at 500MHz with 35.88TOPS/W energy efficiency and 1.05mJ energy consumption per frame.
Hong-Han Lien, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 A Real-Time 1280 × 720 Object Detection Chip With 585 MB/s Memory Traffic
abstract
Memory bandwidth has become the real-time bottleneck of current deep learning accelerators (DLAs), particularly for high definition (HD) object detection. Under resource constraints, this article proposes a low memory traffic DLA chip with joint hardware and software optimization. To maximize hardware utilization under memory bandwidth, we morph and fuse the object detection model into a group fusion-ready model to reduce intermediate data access. This reduces the YOLOv2’s feature memory traffic from 2.9 to 0.15 GB/s. To support group fusion, our previous DLA-based hardware employees a unified buffer with write-masking for simple layer-by-layer processing in a fusion group. When compared to our previous DLA with the same processing element (PE) numbers, the chip implemented in a 40-nm process supports$1280\times 720$at 30 frames per second (FPS) object detection and consumes$7.9\times $less external dynamic random access memory (DRAM) access energy, from 2607 to 327.6 mJ.
Kuo-Wei Chang, Hsu-Tung Shih, Tian-Sheuan Chang, Shang-Hong Tsai, Chih-Chyau Yang, Chien-Ming Wu, Chun-Ming Huang
IEEE Trans. Very Large Scale Integr. Syst.3
2022 A 14 μJ/Decision Keyword-Spotting Accelerator With In-SRAMComputing and On-Chip Learning for Customization
abstract
Keyword spotting (KWS) has gained popularity as a natural way to interact with consumer devices in recent years. However, because of its always on nature and the variety of speech, it necessitates a low-power design as well as user customization. This article describes a low-power, energy-efficient KWS accelerator with static random access memory (SRAM)-based in-memory computing (IMC) and on-chip learning for user customization. However, IMC is constrained by macro size, limited precision, and nonideal effects. To address the issues mentioned above, this article proposes bias compensation and fine-tuning using an IMC-aware model design. Furthermore, because learning with low-precision edge devices results in zero error and gradient values due to quantization, this article proposes error scaling and small gradient accumulation to achieve the same accuracy as ideal model training. The simulation results show that with user customization, we can recover the accuracy loss from 51.08% to 89.76% with compensation and fine-tuning and further improve to 96.71% with customization. The chip implementation can successfully run the model with only 14$\mu \text{J}$per decision. When compared to the state-of-the-art works, the presented design has higher energy efficiency with additional on-chip model customization capabilities for higher accuracy.
Yu-Hsiang Chiang, Tian-Sheuan Chang, Shyh-Jye Jou
IEEE Trans. Very Large Scale Integr. Syst.2
2021 VSA: Reconfigurable Vectorwise Spiking Neural Network Accelerator
abstract
Spiking neural networks (SNNs) that enable low- power design on edge devices have recently attracted significant research. However, the temporal characteristic of SNNs causes high latency, high bandwidth and high energy consumption for the hardware. In this work, we propose a binary weight spiking model with IF-based Batch Normalization for small time steps and low hardware cost when direct training with input encoding layer and spatio-temporal back propagation (STBP). In addition, we propose a vectorwise hardware accelerator that is reconfigurable for different models, inference time steps and even supports the encoding layer to receive multi-bit input. The required memory bandwidth is further reduced by two-layer fusion mechanism. The implementation result shows competitive accuracy on the MNIST and CIFAR-10 datasets with only 8 time steps, and achieves power efficiency of 25.9 TOPS/W.
Hong-Han Lien, Chung-Wei Hsu, Tian-Sheuan Chang
ISCAS3
2020 Efficient Accelerator for Dilated and Transposed Convolution with Decomposition
abstract
Hardware acceleration for dilated and transposed convolution enables real time execution of related tasks like segmentation, but current designs are specific for these convolutional types or suffer from complex control for reconfigurable designs. This paper presents a design that decomposes input or weight for dilated and transposed convolutions respectively to skip redundant computations and thus executes efficiently on existing dense CNN hardware as well. The proposed architecture can cut down 87.8% of the cycle counts to achieve 8.2X speedup over a naive execution for the ENet case.
Kuo-Wei Chang, Tian-Sheuan Chang
ISCAS2
2020 Zebra: Memory Bandwidth Reduction for CNN Accelerators with Zero Block Regularization of Activation Maps
abstract
The large amount of memory bandwidth between local buffer and external DRAM has become the speedup bottleneck of CNN hardware accelerators, especially for activation maps. To reduce memory bandwidth, we propose to learn pruning unimportant blocks dynamically with zero block regularization of activation maps (Zebra). This strategy has low computational overhead and could easily integrate with other pruning methods for better performance. The results show that the proposed method for Resnet-20 on Tiny-Imagenet can reduce 70% of memory bandwidth and further improve to 76% with the combination of Network Slimming, all within 1% accuracy drops.
Hsu-Tung Shih, Tian-Sheuan Chang
ISCAS2
2020 Real-Time Wearable Gait Phase Segmentation for Running And Walking
abstract
Previous gait phase detection as convolutional neural network (CNN) based classification task requires cumbersome manual setting of time delay or heavy overlapped sliding windows to accurately classify each phase under different test cases, which is not suitable for streaming Inertial-Measurement-Unit (IMU) sensor data and fails to adapt to different scenarios. This paper presents a segmentation based gait phase detection with only a single six-axis IMU sensor, which can easily adapt to both walking and running at various speeds. The proposed segmentation uses CNN with gait phase aware receptive field setting and IMU oriented processing order, which can fit to high sampling rate of IMU up to 1000Hz for high accuracy and low sampling rate down to 20Hz for real time calculation. The proposed model on the 20Hz sampling rate data can achieve average error of 8.86 ms in swing time, 9.12 ms in stance time and 96.44% accuracy of gait phase detection and 99.97% accuracy of stride detection. Its real-time implementation on mobile phone only takes 36 ms for 1 second length of sensor data.
Jien-De Sui, Tzyy-Yuang Shiang, Tian-Sheuan Chang
ISCAS4
2019 NV-BNN: An Accurate Deep Convolutional Neural Network Based on Binary STT-MRAM for Adaptive AI Edge
abstract
Binary STT-MRAM is a highly anticipated embedded nonvolatile memory technology in advanced logic nodes < 28 nm. How to enable its in-memory computing (IMC) capability is critical for enhancing AI Edge. Based on the soon-available STT-MRAM, we report the first binary deep convolutional neural network (NV-BNN) capable of both local and remote learning. Exploiting intrinsic cumulative switching probability, accurate online training of CIFAR-10 color images (~ 90%) is realized using a relaxed endurance spec (switching ≤ 20 times) and hybrid digital/IMC design. For offline training, the accuracy loss due to imprecise weight placement can be mitigated using a rapid non-iterative training-with-noise and fine-tuning scheme.
Chih-Cheng Chang, Ming-Hung Wu, Jia-Wei Lin, Chun-Hsien Li, Vivek Parmar, Heng-Yuan Lee, Jeng-Hua Wei, Shyh-Shyuan Sheu, Manan Suri, Tian-Sheuan Chang, Tuo-Hung Hou
DAC10
2019 VSCNN: Convolution Neural Network Accelerator with Vector Sparsity
abstract
Hardware accelerator for convolution neural network (CNNs) enables real time applications of artificial intelligence technology. However, most of the accelerators only support dense CNN computations or suffers complex control to support fine grained sparse networks. To solve above problem, this paper presents an efficient CNN accelerator with 1-D vector broadcasted input to support both dense network as well as vector sparse network with the same hardware and low overhead. The presented design achieves 1.93X speedup over the dense CNN computations.
Kuo-Wei Chang, Tian-Sheuan Chang
ISCAS2
2019 Run Time Adaptive Network Slimming for Mobile Environments
abstract
Modern convolutional neural network (CNN) models offer significant performance improvement over previous methods, but suffer from high computational complexity and are not able to adapt to different run-time needs. To solve above problem, this paper proposes an inference-stage pruning method that offers multiple operation points in a single model, which can provide computational power-accuracy modulation during run time. This method can perform on shallow CNN models as well as very deep networks such as Resnet101. Experimental results show that up to 50% savings in the FLOP are available by trading away less than 10% of the top-1 accuracy.
Hong Ming Chiu, Kuan-Chih Lin, Tian-Sheuan Chang
ISCAS3
2017 Fast rate distortion optimization with adaptive context group modeling for HEVC
abstract
Rate distortion optimization helps decide the best coding mode and partition to improve coding efficiency, but suffers from serious data dependency and complexity that hinders an efficient hardware encoder implementation. Thus, this paper presents a hardware-friendly fast rate-distortion method and its design. For rate estimation, we propose a context group adaptive entropy based method for more precise estimation and parallel computation that is applied to both intra and inter predictions instead of intra prediction only as in previous approaches. For distortion estimation, we use the transform domain instead of spatial domain computation to save inverse discrete transform computation and image reconstruction, and reduce cost further by adopting fixed zero sub-blocks in the high frequency part for 32×32 and 16×16 blocks. The simulation results shows 1.77% BD-rate loss in average. The proposed hardware design adopts the interleaved Luma/Chroma coding schedule to improve hardware utilization. The final implementation with TSMC 40nm CMOS process can achieve real time 4K×2K@30fps encoding with 57.95K gate count while operating under 400MHz clock frequency.
Hung-Cheng Chen, Tian-Sheuan Chang
ISCAS2
2016 Fast intra prediction algorithm and design for high efficiency video coding
abstract
To meet the real time demand of HEVC intra encoding, this paper proposed a fast intra prediction algorithm and its design with a gradient weight controlled block size selection to reduce number of PU (prediction unit) sizes to two. These two PU sizes will be further selectively reduced to one based on its SATD distribution. The simulation results show that the proposed algorithm can save 79% encoding time for all-intra main case compared to HM-9.0rc1, with 3.4% BD-rate increase. The hardware design costs 224K gate count and 1.7KB SRAM for 4Kx2k@30fps processing with TSMC 90 nm CMOS technology when operated at 270 MHz operating frequency.
Han-Chiou Fang, Hung-Cheng Chen, Tian-Sheuan Chang
ISCAS3
2016 A QFHD 30-frames/s HEVC Decoder Design
abstract
The High Efficiency Video Coding (HEVC) standard provides superior compression with large and variablesize coding units and advanced prediction modes, which leads to high buffer costs, memory bandwidth, and irregular computation for ultra high-definition video decoding hardware. Thus, this paper presents an HEVC decoder with a four-stage mixed block size pipeline to reduce the pipeline stage buffer size by approximately 91% compared with the 64×64 block-based pipeline. The high memory bandwidth due to motion compensation problem was solved by 16 × 16 block-based data access, precision-based data access, and a smart buffer to reduce the data bandwidth by 88%. In addition, for irregular computation, a reconfigurable architecture was adopted to unify the variable-size transform. A common intra-prediction module was also designed with a 4 × 4 block-based bottom-up computation for variable-size intra prediction and modes in a regular manner. Furthermore, the corner position computation for the motion vector predictor was applied to handle variable-size motion compensation. Finally, the implementation with the TSMC 90-nm CMOS process used 467k logic gates and 15.778 kB of on-chip memory and supported 4096×2160 at 30-frames/s video decoding at a 270-MHz operation frequency.
Pai-Tse Chiang, Yi-Ching Ting, Hsuan-ku Chen, Shiau-Yu Jou, I-Wen Chen, Hang-Chiu Fang, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. Video Technol.7
2015 A hardware-efficient deblocking filter design for HEVC
abstract
This paper presents a hardware-efficient deblocking filter architecture for High Efficiency Video Coding (HEVC) to reduce visual artifacts at block boundaries. This design proposes an interleaved scheduling to reduce the intermediate data storage to be 1536 bits instead of whole 8192 bits. The implementation with 90 nm CMOS technology can support real-time deblocking operation of 7682×4320@30 fps under 141.5 MHz with only 31K gate count.
Chih-Chung Fang, I-Wen Chen, Tian-Sheuan Chang
ISCAS3
2015 Fast Motion Estimation Algorithm and Design for Real Time QFHD High Efficiency Video Coding
abstract
Motion estimation (ME) in the latest High Efficiency Video Coding standard adopts the quadtree coding structure and up to a 64 × 64 prediction unit (PU) size to improve the coding gain. However, these techniques also have serious design problems regarding the complexity, data dependency, external memory bandwidth, and on-chip buffer size compared with previous standards, especially for real-time ultrahigh-definition video coding. To solve these problems, this paper proposes an efficient ME design with a joint algorithm and architecture optimization. To reduce complexity, we propose a predictive integer ME (IME) algorithm that selects the most probable search directions and steps through a statistical analysis to reduce the number of search points by 90.5%. We also employ a PU size-dependent fractional ME (FME) algorithm to reduce the interpolation filtering by 62.4% compared with the reference software. To resolve the corresponding dependency, we cascade the IME and FME computations via interlaced scheduling and propose an early motion vector prediction candidate approach. We use this scheduling with a 16 × 16 processing unit to compute the partial matching cost of all PUs with the same 16 × 16 current block in an interlaced order and share their common reference block to reduce the on-chip buffer size and off-chip memory bandwidth. The bandwidth is further reduced by a cache with double Z scan indexed addressing to simplify the cache controller. Implementation with a Taiwan Semiconductor Manufacturing Company 90-nm CMOS process supports the real-time encoding of 4 K × 2 K at 60 frames/s operated at 270 MHz with 778.7k logic gates and 17.4 KB of on-chip memory.
Shiaw-Yu Jou, Shan-Jung Chang, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. Video Technol.3
2014 Gradient-based PU size selection for HEVC intra prediction
abstract
This paper proposes a fast hardware friendly intra block size selection algorithm with simple gradient calculations and the bottom-up structure to solve the high computational demand of intra prediction in the latest HEVC (High Efficiency Video Coding) standard. The simulation results show that the proposed algorithm can save 57.3% encoding time in average for all-intra main case compared to the default encoding scheme in HM-9.0rc1, with slightly 2.2% BD-rate increases.
Yi-Ching Ting, Tian-Sheuan Chang
ISCAS2
2014 Low Complexity Formant Estimation Adaptive Feedback Cancellation for Hearing Aids Using Pitch Based Processing
abstract
This paper proposes a novel algorithm and architecture for the adaptive feedback cancellation (AFC) based on the pitch and the formant information for hearing aid (HA) applications. The proposed method, named as Pitch based Formant Estimation (PFE-AFC), has significantly low complexity compared to Prediction Error Method AFC (PEM-AFC). The proposed PFE-AFC consists of a forward and a backward path processing. The forward path processing includes a low complexity pitch based formant estimator for decorrelation filter coefficients update and a pitch based voice activity detector for speech detection, which facilitates the feedback cancellation filter in the backward path to reduce feedback component and maintain speech quality. From system point of view, the PFE-AFC has low complexity overhead since it is easy to share computation resource with other components in the HA system, such as noise reduction and auditory compensation. In addition, the PFE-AFC is suitable for hardware implementation owing to its regular structure. Complexity evaluations show that the PFE-AFC has four orders lower complexity than the PEM-AFC. Simulation results show that the PFE-AFC and the PEM-AFC can achieve similar PESQ (perceptual evaluation speech quality) and ASG (added stable gain). Moreover, the proposed PFE-AFC can outperform the conventional AFC.
Yi FanChiang, Cheng-Wen Wei, Yi-Le Meng, Yu-Wen Lin, Shyh-Jye Jou, Tian-Sheuan Chang
IEEE ACM Trans. Audio Speech Lang. Process.6
2014 Correction to "Low complexity formant estimation adaptive feedback cancellation for hearing aids using pitch based processing"
abstract
In the above paper (ibid., vol. 22, no. 8, pp.1248-1259, Aug. 2014), Table IV is incorrect. The correct table is presented here.
Yi FanChiang, Cheng-Wen Wei, Yi-Le Meng, Yu-Wen Lin, Shyh-Jye Jou, Tian-Sheuan Chang
IEEE ACM Trans. Audio Speech Lang. Process.6
2013 A reconfigurable inverse transform architecture design for HEVC decoder
abstract
In this paper, we present a reconfigurable hardware design which can support the inverse transform size from 4×4 to 32×32 in HEVC (High Efficiency Video Coding). We explore the coefficient properties of various inverse transforms such that a base inverse transform unit can be reconfigured or refined to generate other size of inverse transform. The implementation in 90nm technology can support 3840×2160@30fps processing and only needs about 133.8K gate count, which can save 53% of gate count when compared with previous work.
Pai-Tse Chiang, Tian-Sheuan Chang
ISCAS2
2013 Fast zero block detection and early CU termination for HEVC Video Coding
abstract
This paper proposed a fast zero block detection for various transform size from 32×32 to 4×4 in the new generation of the High Efficiency Video Coding (HEVC) standard. The derivation is based on sum-of-absolute-difference (SAD) value available in the inter prediction computation. The proposed method achieves detection accuracy to 90% in average, and saves transform unit computation by 44% (QP at 22) and 65% (QP at 32) with negligible coding performance loss, when compared with that of HM4.0rc1. Additionally, this pre-skip detection could further help decide the CU inter mode efficiently with about 50% time saving.
Pai-Tse Chiang, Tian-Sheuan Chang
ISCAS2
2013 An Efficient Mode Preselection Algorithm for Fractional Motion Estimation in H.264/AVC Scalable Video Extension
abstract
The video coding standard, H.264/AVC scalable video extension (SVC), adopts various advanced interlayer prediction modes to explore the data redundancies between layers for better coding efficiency but at the expense of significantly increased computational complexity and data access bandwidth, especially for hardware realization of fractional motion estimation mode decision. To deal with this problem, this paper proposes a mode preselection algorithm for fractional motion estimation in scalable video coding. We first analyze the rate distortion cost relationship between different prediction modes. With the statistical results, several mode preselection rules are proposed to filter out the potentially skippable prediction modes. Simulation results show that our proposed algorithm reduces up to 65.97% prediction modes and 79.79% coding time on average with only 0.036 dB and 0.496% BD-peak signal-to-noise ratio (PSNR) degradation and BD-rate increase, respectively. Furthermore, the proposed mode preselection algorithm has been implemented in hardware and it costs only 9 k gate counts, when synthesized by 90 nm CMOS technology.
Gwo-Long Li, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. Video Technol.2
2013 Fast SIFT Design for Real-Time Visual Feature Extraction
abstract
Visual feature extraction with scale invariant feature transform (SIFT) is widely used for object recognition. However, its real-time implementation suffers from long latency, heavy computation, and high memory storage because of its frame level computation with iterated Gaussian blur operations. Thus, this paper proposes a layer parallel SIFT (LPSIFT) with integral image, and its parallel hardware design with an on-the-fly feature extraction flow for real-time application needs. Compared with the original SIFT algorithm, the proposed approach reduces the computational amount by 90% and memory usage by 95%. The final implementation uses 580-K gate count with 90-nm CMOS technology, and offers 6000 feature points/frame for VGA images at 30 frames/s and ∼ 2000 feature points/frame for 1920 × 1080 images at 30 frames/s at the clock rate of 100 MHz.
Liang-Chi Chiu, Tian-Sheuan Chang, Jiun-Yen Chen, Nelson Yen-Chung Chang
IEEE Trans. Image Process.2
2013 Algorithm and Architecture Design of Bandwidth-Oriented Motion Estimation for Real-Time Mobile Video Applications
abstract
This paper proposes a data bandwidth-oriented motion estimation design for resource-limited mobile video applications using an integrated bandwidth rate distortion optimization framework. This framework predicts and allocates the appropriate data bandwidth for motion estimation under a limited bandwidth supply to fit a dynamically changing bandwidth supply. The simulation results show that our proposed algorithm can achieve 66% and 41% memory bandwidth savings while maintaining an equivalent rate-distortion performance and meeting real-time targets, when compared with conventional approaches for low-motion and high-motion D1 (704 × 576)-size video, respectively. The final implementation costs 122 K gate counts with TSMC 0.13-μ m CMOS technology and consumes 74 mW of power for D1 resolution at 30 frames/s which is 40% of that achieved in previous designs.
Jui-Hung Hsieh, Tian-Sheuan Chang
IEEE Trans. Very Large Scale Integr. Syst.2
2013 135-MHz 258-K Gates VLSI Design for All-Intra H.264/AVC Scalable Video Encoder
abstract
To satisfy the video application diversities, an extension of H.264/advanced video coding (AVC), called scalable video coding (SVC), is designed to provide multiple demanded video data via a single video encoder. However, constructed on the fundamental of H.264/AVC, the complexity of SVC is much higher than that of H.264/AVC. In this paper, a VLSI design for all-intra scalable video encoder is proposed to aim at efficient scalable video encoding. First, the memory bandwidth requirements for several encoding methods are analyzed to find out the best encoding method which can achieve best tradeoff between internal memory usage and external memory access. Afterward, an all-intra SVC encoder combined with several advanced techniques, including fast intra prediction algorithm, efficient syntax element encoding approach in context-adaptive variable-length coding, and hardware-efficient techniques, are implemented in a macroblock (MB)-level pipeline to increase data throughput. Implementation results demonstrate that our proposed SVC encoder can process more than 594-k MBs per second, which is equivalent to the summation of 60 high-definition, 1080-p, SD 480-p, and common intermediate format frames under 135-MHz working frequency. The proposed design consumes 258-K gate counts when synthesized by 90-nm CMOS technology.
Gwo-Long Li, Tzu-Yu Chen, Meng-Wei Shen, Meng-Hsun Wen, Tian-Sheuan Chang
IEEE Trans. Very Large Scale Integr. Syst.5
2012 A high throughput CAVLC design for HEVC
abstract
This paper proposes a high throughput context adaptive variable length coding (CAVLC) hardware design for high bit rate HEVC standard. The proposed design adopts a multi-coefficient encoding architecture with the input-parallel information-cascade method to solve the data dependency while attain high throughput. The final implementation with 90nm CMOS technology can process at least 3.2 coefficients per cycle with 12193 gate count when operate at 270MHz. This processing rate can support real video coding with 4K×2K@60fps at the high bit rate case.
Hsuan-ku Chen, Tian-Sheuan Chang
ISCAS2
2012 A low complexity speech coder for binaural communication in hearing aids
abstract
This paper presents a low complexity speech coder with the adaptive Minimum-Maximum Scalar Quantization (MMSQ) method suitable for binaural communication in hearing aids. Compared to the fixed rate MMSQ, the adaptive MMSQ can achieve similar or better compression ratio with 13dB SNR improvement. The final codec implementation with 0.18um CMOS process only costs 6K gate count with 118uW power consumption.
Shuo-Wen Hsu, Tian-Sheuan Chang
ISCAS2
2012 Fast disparity estimation for 3DTV applications
abstract
The depth estimation reference software (DERS) algorithm developed by the MPEG 3-D Video Coding could produce high-quality disparity maps for 3DTV applications but suffers from high computational complexity due to its complicated graph-cut optimization. Therefore, this paper proposed a new fast disparity estimation algorithm that could significantly reduce the computational complexity by the downsampled matching cost method. To address the temporal consistency in videos, this paper proposed the no-motion registration for the foreground copy artifact and the still-edge preservation for the flicker artifact. In addition, the occlusion problem is also solved in the proposed algorithm. The experimental results show that our algorithm could generate comparable disparity maps to the DERS algorithm, and only takes 10.8% of its execution time.
Yu-Cheng Tseng, Tian-Sheuan Chang
VCIP2
2012 A 135 MHz 542 k Gates High Throughput H.264/AVC Scalable High Profile Decoder
abstract
To satisfy the requirement of application heterogeneities, the latest H.264/AVC based video coding standard called scalable video coding additional includes temporal, SNR, and spatial scalabilities for frame rate, quality, and frame resolution adaptation. However, these inclusions significantly increase chip design difficulties such as decoding time, memory bandwidth, and area cost. This paper presents an H.264/AVC scalable high profile decoder realization with several optimization techniques to provide high throughput video decoding. For decoding flow, this paper proposes an one-pass macroblock-based quality layer decoding flow for SNR scalability and 71% of external memory bandwidth and 66% of macroblock processing cycles can be saved. For texture padding in interlayer intra prediction, the modified padding flow can save 26% of decoding time. For interlayer predictor design, this paper proposes a centralized concept for accumulation-based calculation of corresponding spatial position, simplified poly-phase interpolator, and efficient motion vector generator to save area cost and decoding time. Furthermore, the residual reconstruction path with the parallel-pipeline architecture is also proposed to cope with the additional decoding complexity and thus leads to 54% of gate count savings compared to the traditional serial-pipeline architecture. Finally, the proposed H.264/AVC scalable high profile decoder design is implemented with 90 nm CMOS technology and it costs 542 k gate count and 39.66 Kbytes on-chip memory while is capable to decode 60 frames/s for resolution with three quality layers at 135 MHz operating frequency.
Gwo-Long Li, Yuan-Hsin Liao, Po-Yuan Hsu, Meng-Hsun Wen, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. Video Technol.6
2012 A Highly Efficient VLSI Architecture for H.264/AVC Level 5.1 CABAC Decoder
abstract
In this paper, a high throughput context-based adaptive binary arithmetic coding decoder design is proposed. This decoder employs a syntax element prediction method to solve pipeline hazard problems. It also uses a new hybrid memory two-symbol parallel decoding in order to enhance performance as well as to reduce costs. The critical path delay of the two-symbol binary arithmetic decoding engine is improved by 28% with an efficient mathematical transform. Experimental results show that the throughput of our proposed design can reach 485.76 Mbins/s in the high bit-rate coding and 446.2 Mbins/s on average at 264MHz operating frequency, which is sufficient to support H.264/AVC level 5.1 real-time decoding.
Yuan-Hsin Liao, Gwo-Long Li, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. Video Technol.3
2012 A 385 MHz 13.54 K Gates Delay Balanced Two-Level CAVLC Decoder for Ultra HD H.264/AVC Video
abstract
To satisfy the heavy performance requirement in real-time high-resolution H.264/AVC, very large-scale integrated implementation of the entropy decoder is necessary since it dominates the overall decoder throughput. In this paper, we propose a high-throughput delay balanced two-level context-based adaptive variable length coding (CAVLC) decoder with 21% shorter critical path delay in comparison to the traditional two-level decoder design. Furthermore, redundant decoding processes are removed by a skipping mechanism. The proposed CAVLC decoder only needs 127.13 cycles per macroblock on average to support level 5.1 decoding with 13.54k gate counts under 90-nm CMOS technology.
Yuan-Hsin Liao, Gwo-Long Li, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. Video Technol.3
2012 Sub µW Noise Reduction for CIC Hearing Aids
abstract
This paper presents a sub noise reduction design to enhance speech for completely-in-the-canal (CIC) type hearing aids by optimizing its algorithm and associated architecture. In algorithm optimization, a low-complexity mixed perceptual-discrete wavelet packet transform (P-DWPT) and fast Hartley transform (FHT) are adopted for spectral decomposition and reconstruction. A simple yet efficient denoise method with 4-zone-voice activity detection (VAD) supports a consonant protection to improve speech quality and a skip scheme to reduce power consumption. In the designed architecture, mixed P-DWPT and FHT are folded into one 8-by-8 configurable butterfly computation unit with on-time scheduling for low power operation. The circuit is implemented with 0.18 μm CMOS process and consumes only 0.65 μW power at 1.0 V with a speech quality that is comparable to that achieved using other high-complexity algorithms.
Cheng-Wen Wei, Sheng-Jie Su, Tian-Sheuan Chang, Shyh-Jye Jou
IEEE Trans. Very Large Scale Integr. Syst.3
2011 Real-time high-definition stereo matching on FPGA
abstract
Although many fast stereo matching designs have been proposed in the past decades, it is still very challenging to achieve real-time speed at high definition resolution while maintaining high matching accuracy. In this paper, we propose a real-time high definition stereo matching design on FPGA. By using the Mini-Census transform and the Cross-based cost aggregation, the proposed algorithm is robust to radiometric differences and produces accurate disparity maps. The algorithm modules have been optimized for efficient hardware implementations and instantiated in an SoC environment. Implemented on a single EP3SL150 FPGA, our design achieves 60 frames per second for 1024 × 768 stereo images. Evaluated with the Middlebury stereo benchmark, the proposed design also delivers leading stereo matching accuracy among prior related work.
Lu Zhang 0018, Ke Zhang 0012, Tian-Sheuan Chang, Gauthier Lafruit, Georgi Kuzmanov, Diederik Verkest
FPGA3
2011 A 94fps view synthesis engine for HD1080p video
abstract
This paper presents a low-cost and high-throughput view synthesis engine based the view synthesis reference software (VSRS) algorithm. With the horizontal shift mode, we propose the row-based pipelined architecture to save the memory cost for original camera rotation issue. Owing to row- based method, internal Z-buffers for storing depth data can be reduced, and also the external bandwidth can be reduced. With the 90nm technology process, our view synthesis engine can achieve the throughput of 94.5 frame/sec for the HD1080p input with the gate count of 142.9k and the low memory cost of 54.72Kbytes.
Fu-Jen Chang, Yu-Cheng Tseng, Tian-Sheuan Chang
VCIP3
2011 VLSI Architecture for Real-Time HD1080p View Synthesis Engine
abstract
This paper presents a real-time HD1080p view synthesis engine based on the reference algorithm from 3-D video coding team by solving high computational complexity and high memory cost problems. For the computational complexity, we propose the bilinear interpolation to simplify the hole filling process, and the Z scaling method with floating-point format to reduce the cost of homography calculation. For the memory cost, we propose the frame-level pipelining to reduce the requirement of warped depth maps, and the column-order warping method to remove the Z-buffer in occlusion handling. With the 90 nm complementary metal-oxide-semiconductor technology, our view synthesis engine can archive the throughput of 32.4 f/s for HD1080p videos with the gate count of 268.5 K and the internal memory of 69.4 kbytes. The experimental result shows our implementation has the similar synthesis quality as the original reference algorithm.
Ying-Rung Horng, Yu-Cheng Tseng, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. Video Technol.3
2011 A 124 Mpixels/s VLSI Design for Histogram-Based Joint Bilateral Filtering
abstract
This paper presents an efficient and scalable design for histogram-based bilateral filtering (BF) and joint BF (JBF) by memory reduction methods and architecture design techniques to solve the problems of high memory cost, high computational complexity, high bandwidth, and large range table. The presented memory reduction methods exploit the progressive computing characteristics to reduce the memory cost to 0.003%-0.020%, as compared with the original approach. Furthermore, the architecture design techniques adopt range domain parallelism and take advantage of the computing order and the numerical properties to solve the complexity, bandwidth, and range-table problems. The example design with a 90-nm complementary metal-oxide-semiconductor process can deliver the throughput to 124 Mpixels/s with 356-K gate counts and 23-KB on-chip memory.
Yu-Cheng Tseng, Po-Hsiung Hsu, Tian-Sheuan Chang
IEEE Trans. Image Process.3
2010 Efficient inter-layer prediction hardware design with extended spatial scalability for H.264/AVC scalable extension
abstract
To support inter-layer prediction with arbitrary frame resolution ratio between successive spatial layers, the scalable video coding (SVC) adopts the mechanism of extended spatial scalability (ESS) to achieve it but with noticeable hardware implementation complexity due to the numerous multiplication operations. Therefore, this paper proposes a hardware efficient inter-layer prediction architecture design with ESS by means of accumulator approach. In addition, an area efficient inter-layer interpolator architecture and simplified transform block identification scheme are also proposed to further reduce hardware costs. Simulation results demonstrate that our proposed architecture can significantly save gate count when compared to direct implementation approach.
Gwo-Long Li, Tian-Sheuan Chang
ISCAS3
2010 Stereoscopic images generation with directional Gaussian filter
abstract
The depth-based image rendering (DIBR) uses monoscopic image with its corresponding depth map to generate stereoscopic images for 3DTV. The main difficulty in the DIBR is that holes appear in the stereoscopic images because of the occlusion between objects. However, the previous depth smoothing methods to reduce holes usually destroy the perceived depth quality or the perceived naturalness. This paper proposed a new depth smoothing method to solve these problems by applying a new directional Gaussian filter guided by edge direction in the hole-flag area iteratively. The subjective evaluation results show that our proposed method can generate better stereoscopic images with good perceived depth quality as well as perceived naturalness.
Ying-Rung Horng, Yu-Cheng Tseng, Tian-Sheuan Chang
ISCAS3
2010 Low memory cost bilateral filtering using stripe-based sliding integral histogram
abstract
Bilateral filter can well smooth images while preserving edges, and thus has been widely used in various applications. However, it needs significant memory cost due to a whole image storage. This paper develops three methods to significantly reduce the memory cost. The runtime updating method (RUM) discards unnecessary data in runtime. The stripe-based integral histogram method (SBM) divides image into vertical image-stripes. The sliding origin method (SOM) sliding moves the origin of integration region to achieve the most memory reduction. The final result shows that the proposed approach only needs 24Kbits memory, which saves 99.993% of memory cost for smoothing a VGA-sized image with 30-fps speed, when compared to the previous optimized constant time integral histogram approach.
Po-Hsiung Hsu, Yu-Cheng Tseng, Tian-Sheuan Chang
ISCAS3
2010 A high throughput VLSI design with hybrid memory architecture for H.264/AVC CABAC decoder
abstract
A high throughput context-based adaptive binary arithmetic coding (CABAC) decoding design with hybrid memory architecture for H.264/AVC is presented in this paper. To accelerate the decoding speed with hardware cost consideration, a new hybrid memory two-symbol parallel decoding technique is proposed. In addition, an efficient mathematical transform method is also proposed to further decrease the critical path of two-symbol binary arithmetic decoding procedure. The proposed architecture is implemented by UMC 90nm technology and experimental results show that our proposal can operate at 264 MHz with 42.37k gate count, and the throughput is 483.1 Mbins/sec, which surpasses previous design with 48.6% hardware cost saving.
Yuan-Hsin Liao, Gwo-Long Li, Tian-Sheuan Chang
ISCAS3
2010 Fast stereo matching with predictive search range
abstract
Local stereo matching could deliver accurate disparity maps by the associated method, like adaptive support-weight, but suffers from the high computational complexity, O(NL), where N is pixel count in spatial domain, and L is search range in disparity domain. This paper proposes a fast algorithm that groups similar pixels into super-pixels for spatial reduction, and predicts their search range by simple matching for disparity reduction. The proposed algorithm could be directly applied to other local stereo matching, and reduce its computational complexity to only 8.2%-17.4% with slight 1.5%-3.2% of accuracy degradation.
Yu-Cheng Tseng, Po-Hsiung Hsu, Tian-Sheuan Chang
PCS3
2010 Algorithm and Architecture of Disparity Estimation With Mini-Census Adaptive Support Weight
abstract
High-performance real-time stereo vision system is crucial to various stereo vision applications, such as robotics, autonomous vehicles, multiview video coding, freeview TV, and 3-D video conferencing. In this paper, we proposed a high-performance hardware-friendly disparity estimation algorithm called mini-census adaptive support weight (MCADSW) and also proposed its corresponding real-time very large scale integration (VLSI) architecture. To make the proposed MCADSW algorithm hardware-friendly, we proposed simplification techniques such as using mini-census, removing proximity weight, using YUV color representation, using Manhattan color distance, and using scaled-and-truncate weight approximation. After applied these simplifications, the MCADSW algorithm was not only hardware-friendly, but was also 1.63 times faster. In the corresponding real-time VLSI architecture, we proposed partial column reuse and access reduction with expanded window to significantly reduce the bandwidth requirement. The proposed architecture was implemented using United Microelectronics Corporation (UMC) 90 nm complementary metal-oxide-semiconductor technology and can achieve a disparity estimation frame rate of 42 frames/s for common intermediate format size images when clocked at 95 MHz. The synthesized gate-count and memory size is 563 k and 21.3 kB, respectively.
Nelson Yen-Chung Chang, Tsung-Hsien Tsai, Po-Hsiung Hsu, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. Video Technol.5
2010 RD Optimized Bandwidth Efficient Motion Estimation and Its Hardware Design With On-Demand Data Access
abstract
Data bandwidth dominates the performance and power consumption in the video encoder design. In which a low bandwidth and bandwidth aware motion estimation design enables smooth and better video quality as well as lower power consumption on data accesses. This paper proposes a bandwidth efficient motion estimation and its hardware implementation to deal with the bandwidth issues. First, an on-demand data access mechanism is proposed to acquire the reference data according to the video content for motion estimation process and thus can avoid unnecessary reference data loading. Furthermore, the available bandwidth constraint is properly modeled into our proposed rate distortion optimization framework to efficiently use the data bandwidth. Simulation results show that our proposed algorithm not only allocates proper data bandwidth for motion estimation according to video content but also saves 79.15% data bandwidth demand with 0.03 dB PSNR drop and 2.50% bitrate increase in maximum for 4 CIF resolution sequences, when compared to the fully data reuse full search motion estimation which reuses the overlapped reference data to avoid unnecessary data reloading. In addition, under the available data bandwidth constraint, our proposed algorithm can achieve 2.43%, 0.08%, and 0.20% BD-bitrate saving with 0.17 dB, 0.01 dB, and 0.01 dB BD-PSNR increase on average for high, median, and low motion sequences when compared to the full search motion estimation algorithm. The resulted design only needs 75.27 K gate counts when running at 23 MHz operating frequency for 4 CIF at 30 frames/s with 90 nm CMOS process due to its search range independent buffer design.
Gwo-Long Li, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. Video Technol.2
2010 Architecture Design of Belief Propagation for Real-Time Disparity Estimation
abstract
Belief propagation based algorithms perform best in disparity estimation but suffer from high computational complexity and storage, especially in message passing. This paper proposes an efficient architecture design with three techniques to solve the problems. For the memory storage, we propose the spinning-message and the sliding-bipartite node plane that can reduce memory cost to 1.2% for image-scale algorithms and 23.4% for block-scale algorithms, when compared to the traditional approach. For the logic complexity, we propose a buffer-free processing element architecture that has 3.6 times hardware efficiency of the previous work. The three proposed techniques could be applied to various belief propagation based algorithms to save significant hardware cost as well as approach real-time speed.
Yu-Cheng Tseng, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. Video Technol.2
2009 Bandwidth-rate-distortion optimized motion estimation
abstract
Motion estimation process is the most computational and memory intensive component in video encoders. However, traditional motion estimation algorithms focus on rate and distortion performance and do not take memory bandwidth into consideration. As a result, its performance will be degraded under bandwidth constraint environment. In this paper, a bandwidth-rate-distortion optimized motion estimation algorithm is introduced to solve the above issue. Simulation results show that the proposed method could achieve similar performance as the bandwidth unconstrained method with up to 70% bandwidth saving.
Wei-Cheng Tai, Gwo-Long Li, Tian-Sheuan Chang
ICME3
2009 Memory Analysis for H.264/AVC Scalable Extension Encoder
abstract
In this paper, the first overall memory analysis for H.264/AVC scalable extension (SVC) is presented. Since the dependency among layers is a major discrepancy between SVC and H.264/AVC, the hardware design about them is quite different. Therefore, we analyze the memory bandwidth and memory access requirement of different coding flows for SVC. Based on the analysis results, we deduce that the frame-level method enjoys the best performance. With some improvements from previous work and proposal by our observation, the external memory access could be further reduced by at least 53% than initial design. Furthermore, through our analysis, a solution is provided for the trade-off between internal memory storage and external memory bandwidth in hardware design.
Tzu-Yu Chen, Gwo-Long Li, Tian-Sheuan Chang
ISCAS3
2009 A Memory Efficient Fine Grain Scalability Coefficient Encoding Method for H.264/AVC Scalable Video Extension
abstract
In this paper, a memory efficient Fine Grain Scalability (FGS) coefficient encoding method is proposed to reduce the external memory access requirement. In the H.264/AVC Scalable Video Extension, the FGS coefficients encoding is frame based. However, the frame based mechanism results in the difficulty of hardware implementation due to large internal memory requirements and external memory accesses. Therefore, a non-uniform memory size design which can achieve low external memory access is proposed to realize the macroblock based FGS coefficients encoding. Compared to previous work, our proposed method can save at least 38 KB external memory accesses per frame in average.
Meng-Wei Shen, Gwo-Long Li, Tian-Sheuan Chang
ISCAS3
2009 Low-memory Cost Belief Propagation Architecture for Disparity Estimation
abstract
In disparity estimation, belief propagation can deliver better disparity quality than other algorithms but suffer from large storage cost, especially at the message update processing. To reduce the storage cost, this paper proposes low-memory cost architectures for the message update PE to satisfy the real-time application. We propose four architectures which are post-normalization, shadow buffer, no memory, and no memory+double PE architectures. Compared to the previous design, the proposed no memory+double PE architecture can save 28% of the hardware cost at most for 320times240@30 fps and 64 disparity levels.
Yu-Cheng Tseng, Nelson Yen-Chung Chang, Tian-Sheuan Chang
ISCAS3
2009 A 140-MHz 94 K Gates HD1080p 30-Frames/s Intra-Only Profile H.264 Encoder
abstract
This paper presents a HD1080p 30-frames/s H.264 intra encoder operated at 140 MHz with just 94 K gate count and 0.72-mm2core area for digital video recorder or digital still camera applications. To achieve high throughput and low area cost for high-definition video, we apply the modified three-step fast intra prediction technique to reduce the cycle count while keeping the quality as close as full search. Then, in architecture scheduling, we further adopt the variable pixel parallelism instead of constant four-pixel parallelism to speed up performance on the critical intra prediction part while keeping other parts unchanged for low area cost. The achieved design only needs half of the working frequency and reduces the gate count cost by 23.5% compared with the previous design with the same HD720p 30-frames/s requirement. Besides, our design at 140 MHz can support HD1080p 30 frames/s for digital video encoder or 4096 times2304 images with 6.78 frames/s for digital still camera application.
Yu-Kun Lin, Chun-Wei Ku, De-Wei Li, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. Video Technol.4
2008 A 242mW, 10mm21080p H.264/AVC high profile encoder chip
abstract
A 1080p high profile H.264 encoder is designed by the robust reusable silicon IP methodology and fabricated in a 0.13μm CMOS technology with an area of 10 mm2 and 242mW at 145MHz. Compared to the state-of-the-art design targeted at 720p baseline, this design reduces 53.4% power and 46.7% area through parallelism enhanced throughput and cross stage sharing pipeline.
Yu-Kun Lin, De-Wei Li, Chia-Chun Lin, Tzu-Yun Kuo, Sian-Jin Wu, Wei-Cheng Tai, Wei-Cheng Chang, Tian-Sheuan Chang
DAC8
2008 ISID : In-order scan and indexed diffusion segmentation algorithm for stereo vision
abstract
Existing segmentation algorithms have irregular computing order, expensive sorting, or inefficient backtracking procedure which would reduce their processing speed. In this paper, an in-order scan and indexed diffusion (ISID) segmentation algorithm for stereo vision which is more regular and does not need sorting nor backtracking is proposed. The inorder scan plateau detection is the first step in ISID which detects whether pixels in a 3×3 sliding window belongs to the same region or not. Then the indexed upward diffusion assigns a label to an undetermined pixel using a label diffusion method. Simulation results show that with the introduced regularity and lower complexity, the proposed ISID algorithm reduces 54% and 36% of the execution time when compared with the immersion-based and toboggan-based watershed algorithm in average.
Jing-Chu Chan, Nelson Yen-Chung Chang, Tian-Sheuan Chang
ISCAS3
2008 Data reuse analysis of local stereo matching
abstract
External memory bandwidth and internal memory size have been major bottlenecks in designing VLSI architecture for real-time stereo matching hardware because of large amount of pixel data and disparity range. To address these bottlenecks, this work explores the impact of data reuse on disparity-order and pixel-order along with the partial column reuse (PCR) and vertically expanded row reuse (VERR) techniques we proposed. The analysis suggest that a disparity-order reuse with both PCR and VERR techniques is suitable for low memory cost and low external bandwidth design, whereas the pixel-order reuse with both techniques is more suitable for low computation resource requirement.
Tsung-Hsien Tsai, Nelson Yen-Chung Chang, Tian-Sheuan Chang
ISCAS3
2008 Architecture Design of Shape-Adaptive Discrete Cosine Transform and Its Inverse for MPEG-4 Video Coding
abstract
This paper presents efficient VLSI architectures of the shape-adaptive discrete cosine transform (SA-DCT) and its inverse transform (SA-IDCT) for MPEG-4. Two of the challenges encountered during the exploitation of more efficient architectures for the SA-DCT and SA-IDCT are addressed. One challenge is to handle the architectural irregularity due to the shape-adaptive nature. The other one is to provide acceptable throughput using minimal hardware. In the algorithm-level optimization, this work exploits the numerical properties found in the transform matrices of various lengths, and derives a fine-grained zero-skipping scheme for the IDCT which can perform 22.6% more zero-skipping than the common vector-based coarse-grained zero-skipping scheme does. In the architecture-level design, the 1-D variable-length DCT/IDCT architectures designed on the basis of the numerical properties are proposed. An auto-aligned transpose memory that aligns the data of different lengths is also incorporated. In addition, a zero-index table is also included in the transpose memory to support the fine-grained zero-skipping in the SA-IDCT. The synthesized designs of the SA-DCT and SA-IDCT are implemented using UMC 0.18-mum technology. The SA-DCT architecture has 26 635 gates, and its average cycle-throughput is 0.66 pixels/cycle, which is comparable to other proposed architectures. On the other hand, the SA-IDCT architecture has 29 960 gates, and its cycle-throughput is 6.42 pixels/cycle. While decoding for CIF@30FPS, the SA-IDCT is clocked at 0.7 MHz, and the power consumption is 0.14 mW. Both the throughput and power consumption of the proposed SA-IDCT architecture are an order better than those of the existing SA-IDCT architectures.
Hui-Cheng Hsu, Kun-Bin Lee, Nelson Yen-Chung Chang, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. Video Technol.4
2008 Adaptive De-Interlacing With Robust Overlapped Block Motion Compensation
abstract
The paper proposes a new method for de-interlacing video. It combines simple techniques such as field insertion and edge-adaptive line averaging as well as motion-compensated (MC) methods by a robust motion detector. The simpler techniques are preferred, and the MC method is limited to places where it is really advantageous to limit processing power without sacrificing too much quality. The final experimental results show that the proposed method needs lower complexity but achieves 2 dB PSNR gain when compared to the previous method.
Sin-Bor Wang, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. Video Technol.2
2007 A Low Cost Context Adaptive Arithmetic Coder for H.264/MPEG-4 AVC Video Coding
abstract
This paper presents a fast and low cost context adaptive binary arithmetic encoder for H.264/MPEG-4 AVC video coding standard through both algorithm level and architecture level optimizations. First in the algorithm level, we process the binarization and context generation in parallel to reduce the encoding iteration cycles to three or four cycles from five cycles in the previous design. Second, in the architecture level, we reduce the cycles of renormalization loops by employing one-skipping and bit-parallelism, and save hardware cost of arithmetic coder by merging three different modes. The implemented design shows that it can achieve the 333 MHz frequency with only 13.3K gate count.
Jian-Long Chen, Yu-Kun Lin, Tian-Sheuan Chang
ICASSP (2)3
2007 SIFME: A Single Iteration Fractional-Pel Motion Estimation Algorithm and Architecture for HDTV Sized H.264 Video Coding
abstract
This paper presents a set of fast algorithm and VLSI architecture for HDTV-sized H.264 fractional motion estimation. To solve the long computational latency in HD-sized application, we propose to use the single iteration algorithm with only six search points. This single iteration method halves the cycle count of two iteration methods in previous approaches. Moreover, we propose to use 4×4 Hadamard instead of 8×8 Hadamard as cost function for H.264 high profiles without significant video quality loss. By these techniques, the resulted architecture can save 20% of area and provide over 40% of throughput improvement than the previous work, and is able to support HDTV applications.
Tzu-Yun Kuo, Yu-Kun Lin, Tian-Sheuan Chang
ICASSP (1)3
2007 A 61MHz 72K Gates 1280x720 30FPS H.264 Intra Encoder
abstract
This paper presents an HD720p 30 frames per sec H.264 intra encoder operated at 61 MHz with just 72 K gate count. We achieve the low cost and low operating frequency with the highly utilized variable pixel scheduling, and a modified three-step fast algorithm. Thus, the resulted design only needs half of operating frequency and reduces 30% of area cost compared to the previous HD720p intra encoder design.
De-Wei Li, Chun-Wei Ku, Chao-Chung Cheng, Yu-Kun Lin, Tian-Sheuan Chang
ICASSP (2)5
2007 PMRME: A Parallel Multi-Resolution Motion Estimation Algorithm and Architecture for HDTV Sized H.264 Video Coding
abstract
The paper presents a hardware-efficient fast algorithm and its architecture for large search range motion estimation (ME) used in HDTV sized H.264 video coding. To solve the high cost and latency in large search range cases, the proposed algorithm processes ME in parallel multi-resolution levels instead of serial processing in the previous approach. This enables high data reuse for lower bandwidth and low memory cost. Further combining with our previous proposed mode filtering and bit truncation, the algorithm only increases the bit rate within -0.58% and 3.06% and at most 0.04 dB and 0.07 dB PSNR degradation for 720p and 1080p sequences respectively. The hardware implementation can save up to 49.5% of area cost and 65% of memory cost compared to the previous approach for large search range to [-128, 127].
Chia-Chun Lin, Yu-Kun Lin, Tian-Sheuan Chang
ICASSP (2)3
2007 Real-Time DSP Implementation on Local Stereo Matching
abstract
Real-time DSP stereo matching solution has been important to various applications relying on stereo vision. We proposed a 4times5 jigsaw matching template and the dual-block parallel processing technique to enhance VLIW DSP stereo matcher's performance. The 4times5 jigsaw template improves the matching quality by 1% compared with regular 4times5 block template while consuming the same amount of memory access bandwidth. Along with the benefit of the jigsaw template, the dual-block parallel processing technique, which doubles the throughput, is possible to be implemented for DSP. Together with instruction scheduling and operation pipelining, our DSP stereo matcher can achieve 50 FPS of 16 disparity levels for a 384times288 stereo image pair. Both quantitative and qualitative stereo matching results are provided at the end of this work.
Nelson Yen-Chung Chang, Ting-Min Lin, Tsung-Hsien Tsai, Yu-Cheng Tseng, Tian-Sheuan Chang
ICME5
2007 Low Memory Cost Block-Based Belief Propagation for Stereo Correspondence
abstract
The typical belief propagation has good accuracy for stereo correspondence but suffers from large run-time memory cost. In this paper, we propose a block-based belief propagation algorithm for stereo correspondence that partitions an image into regular blocks for optimization. With independently partitioned blocks, the required memory size could be reduced significantly by 99% with slightly degraded performance with a 32times32 block size when compared to original one. Besides, such blocks are also suitable for parallel hardware implementation. Experimental results using Middlebury stereo test bed demonstrate the performance of the proposed method.
Yu-Cheng Tseng, Nelson Yen-Chung Chang, Tian-Sheuan Chang
ICME3
2007 A Fast Algorithm and Its VLSI Architecture for Fractional Motion Estimation for H.264/MPEG-4 AVC Video Coding
abstract
This paper presents a fast algorithm and its VLSI architecture for H.264 fractional motion estimation. Motivated by the high correlation of cost between neighboring fractional pel position, the proposed algorithm efficiently explores the neighborhood position around the minimum one and thus skips other unlikely ones. Thus, the proposed search pattern and early termination under constant quantization parameter can reduce about 50% of computation complexity compared to that in reference software but only with 0.1-0.2 dB peak signal-to-noise ratio degradation and less than 2% of bit rate increase. The VLSI architecture of the proposed algorithm thus can save 40% of area cost due to only half of the processing elements and save 14% of searching time when compared with the previous design
Yu-Jen Wang, Chao-Chung Cheng, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. Video Technol.3
2006 Algorithms and DSP implementation of H.264/AVC
abstract
This survey paper intends to provide a comprehensive coverage of the techniques that are pertinent to the processor-based implementation of H.264/AVC video codec, particularly on DSP. Most of this paper is devoted to the computationally efficient algorithms, or the fast algorithms. Fast algorithms for motion estimation, intra-prediction and mode decision are described to reduce the computational complexity. In addition, in order to port the H.264/AVC codec to DSP, we also outline the basic principles of DSP code optimization
Hung-Chih Lin, Yu-Jen Wang, Kai-Ting Cheng, Shang-Yu Yeh, Wei-Nien Chen, Chia-Yang Tsai, Tian-Sheuan Chang, Hsueh-Ming Hang
ASP-DAC7
2006 A 1280×720 pixels 30 frames/s H.264/MPEG-4 AVC intra encoder
abstract
This paper presents an HDTV size H.264/MPEG-4 AVC intra encoder suitable for DSC and digital video camera applications. The chip reduces the gate count by saving the costly plane mode and enhances the video quality with the improved cost function. With careful scheduling and high performance function unit, the developed chip can easily support 29.46M pixels/s still image encoding and real-time moving picture intra coding of HDTV 720p@30fps video application when clocked at 117.28MHz under 0.18mum CMOS process
Chao-Chung Cheng, Chun-Wei Ku, Tian-Sheuan Chang
ISCAS3
2006 A fast fractional pel motion estimation algorithm for H.264/MPEG-4 AVC
abstract
This paper presents a fast algorithm for H.264 fractional ME. Motivated by the highly correlation of cost between neighboring fractional pel position, the proposed algorithm efficiently explores the neighborhood position around the minimum one and thus skips other unlikely ones. Thus, the proposed algorithm can complete the search by only examining 8 or 9 search points instead of 17 search points in the search algorithm of reference software. Early termination technique is applied in every single search point. The simulation result shows that the proposed algorithm can reduce about 50% of computation complexity compared to that in reference software but only with 0.1-0.2 dB PSNR degradation and less than 2% of bit rate increase.
Yu-Jen Wang, Chao-Chung Cheng, Tian-Sheuan Chang
ISCAS3
2006 A zero-skipping multi-symbol CAVLC decoder for MPEG-4 AVC/H.264
abstract
This paper presents a high-performance CAVLC decoding VLSI architecture for MPEG-4 AVC/H.264. Instead of just skipping zero block, the proposed design explores the features of CAVLC decoding process to efficient skip possible processes if none needed to be decoded, and can decode multiple symbols in sign and run before stage. The proposed design just needs average 90 cycles for one MB decoding, which can meet real time HDTV requirement and saves 64% of cycle count in average when compared with previous design. The hardware cost is about 13192 gates when synthesized at 125 MHz.
Guo-Shiuan Yu, Tian-Sheuan Chang
ISCAS2
2006 Combined Frame Memory Motion Compensation for Video Coding
abstract
The frame memory has long been the dominant component in a video decoder in terms of energy, area, and latency. We proposed a non-combined frame memory motion compensation (CFMMC) for video decoding which facilitates the characteristic of the perfect-matched macroblock (MB) to avoid unnecessary memory access and to save energy. The statistic result confirms that some sequences have more than 70% of MBs being perfect-matched MB. The CFMMC hardware architecture is further evaluated for latency, area, and energy. The hardware architecture shows that with SRAM-base frame memory, the equivalent gate count can be reduced by 37.7%, and the energy consumption and the latency may also be improved for sequences with enough percentage of perfect-matched MBs. Since the benefit of the CFMMC is highly dependent on the percentage of perfect-matched MBs, it is best suited for applications with large portion of static background, such as video surveillance, video telephony, and video conferencing
Nelson Yen-Chung Chang, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. Video Technol.2
2006 A High-Definition H.264/AVC Intra-Frame Codec IP for Digital Video and Still Camera Applications
abstract
This paper presents a real-time high-definition 720p@30fps H.264/MPEG-4 AVC intra-frame codec IP suitable for digital video and digital still camera applications. The whole design is optimized in both the algorithm and architecture levels. In the algorithm level, we propose to remove the area-costly plane mode, and enhance the cost function to reduce hardware cost and to increase the processing speed while provide nearly the same quality. In the architecture design, in additional to the fast module implementation the process is arranged in the macroblock-level pipelining style together with three careful scheduling techniques to avoid the idle cycles and improve the data throughput. The whole codec design only needs 103 K gate count for a core size of 1.28times1.28mm2and achieves real-time encoding and decoding at 117 and 25.5 MHz, respectively, when implemented by 0.18-mum CMOS technology
C.-W. Ku, C.-C. Cheng, G.-S. Yu, M.-C. Tsai, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. Video Technol.5
2006 Fast Variable Block Size Motion Estimation by Adaptive Early Termination
abstract
This paper presents a fast motion estimation algorithm by adaptively changing the early termination threshold for the current accumulated partial sum of absolute difference (SAD) value. The simulation results show that the proposed algorithm can provide the similar quality while saving 77.9% and 50.6% of SAD computation when comparing with early termination methods in MPEG-4 VM18.0 and H.264 JM9.0, respectively.
Esam A. Al Qaralleh, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. Video Technol.2
2006 An Efficient Binary Motion Estimation Algorithm and its Architecture for MPEG-4 Shape Encoding
abstract
This paper presents a fast binary motion estimation (BME) algorithm and its architecture for MPEG-4 shape encoding. The proposed algorithm explores the property of the binary-value in BME to quickly skip the unnecessary sum of absolute differences (SAD) computation. When comparing with the full search algorithm, simulation results show that it can efficiently save in the search positions to an average$-hbox99.58$% of that in the full search algorithm with the same PSNR quality. Due to the algorithm's simplicity and regularity, the resulting hardware implementation also exhibits simple and regular control and data flow. It can achieve real-time encoding with only 11582 gate count.
Esam A. Al Qaralleh, Tian-Sheuan Chang, Kun-Bin Lee
IEEE Trans. Circuits Syst. Video Technol.2
2005 A bandwidth efficient subsampling-based block matching architecture for motion estimation
abstract
We have developed a new pel subsampling-based search hardware for motion estimation called quartet-pel motion estimation (QME). The memory access of search range memory can be reduced to 25%. The computational complexity can also be reduced to 25% with respect to full-search block matching algorithm (FBMA). On the other hand, flexible and efficient hardware architecture is also implemented. The flexibility is based on the configuration of processing unit, and adjustable candidate number. In addition, complete verification and testing methods are also considered.
Hao-Yun Chin, Chao-Chung Cheng, Yu-Kun Lin, Tian-Sheuan Chang
ASP-DAC4
2005 Fast block type decision algorithm for intra prediction in H.264 FRext
abstract
This paper presents a fast block type decision algorithm for 4/spl times/4, 8/spl times/8 and 16/spl times/16 intra prediction in H.264/AVC FRExt (fidelity range extension). With additional 8/spl times/8 intra prediction size, the complexity of intra prediction is increased by almost 50%. Thus, the proposed algorithm uses the smoothness of the macroblock to select the intra prediction block size and save computational cost. When further combining with our previously proposed three-step fast intra prediction algorithm, the proposed algorithm can save up to 32.1% and 63.4% of computation time of all I frames and intra prediction coding respectively, and achieves average 1.05% bit rate decrease and 0.046 dB quality degradation when comparing with the reference software.
Yu-Kun Lin, Tian-Sheuan Chang
ICIP (1)2
2005 A memory-efficient realization of cyclic convolution and its application to discrete cosine transform
abstract
This paper presents a memory-efficient approach to realize the cyclic convolution and its application to the discrete cosine transform (DCT). We adopt the way of distributed arithmetic (DA) computation, exploit the symmetry property of DCT coefficients to merge the elements in the matrix of DCT kernel, separate the kernel to be two perfect cyclic forms, and partition the content of ROM into groups to facilitate an efficient realization of a one-dimensional (1-D) N-point DCT kernel using (N-1)/2 adders or subtractors, one small ROM module, a barrel shifter, and ((N-1)/2)+1 accumulators. The proposed memory-efficient design technique is characterized by rearranging the content of the ROM using the conventional DA approach into several groups such that all the elements in a group can be accessed simultaneously in accumulating all the DCT outputs for increasing the ROM utilization. Considering an example using 16-bit coefficients, the proposed design can save more than 57% of the delay-area product, as compare with the existing DA-based designs in the case of the 1-D seven-point DCT. Finally, a 1-D DCT chip was implemented to illustrate the efficiency associated with the proposed approach.
Hun-Chen Chen, Jiun-In Guo, Tian-Sheuan Chang, Chein-Wei Jen
IEEE Trans. Circuits Syst. Video Technol.3
2002 On the data reuse and memory bandwidth analysis for full-search block-matching VLSI architecture
abstract
This work explores the data reuse properties of full-search block-matching (FSBM) for motion estimation (ME) and associated architecture designs, as well as memory bandwidth requirements. Memory bandwidth in high-quality video is a major bottleneck to designing an implementable architecture because of large frame size and search range. First, the memory bandwidth in ME is analyzed and the problem is solved by exploring data reuse. Four levels are defined according to the degree of data reuse for previous frame access. With the highest level of data reuse, one-access for frame pixels is achieved. A scheduling strategy is also applied to data reuse of the ME architecture designs and a seven-type classification system is developed that can accommodate most published ME architectures. This classification can simplify the work of designers in designing more cost-effective ME architectures, while simultaneously minimizing memory bandwidth. Finally, a FSBM architecture suitable for high quality HDTV video with a minimum memory bandwidth feature is proposed. Our architecture is able to achieve 100% hardware efficiency while preserving minimum I/O pin count, low local memory size, and bandwidth.
Jen-Chieh Tuan, Tian-Sheuan Chang, Chein-Wei Jen
IEEE Trans. Circuits Syst. Video Technol.2
2000 A simple processor core design for DCT/IDCT
abstract
This paper presents a cost-effective processor core design that features the simplest hardware and is suitable for discrete cosine transform/indiscrete cosine transform (DCT/IDCT) operations in H.263 and digital camera. This design combines the techniques of fast direct two-dimensional DCT algorithm, the bit level adder-based distributed arithmetic, and common subexpression sharing to reduce the hardware cost and enhance the computing speed. The resulting architecture is very simple and regular such that it can be easily scaled for higher throughput rate requirements. The DCT design has been implemented by 0.6 /spl mu/m SPDM CMOS technology and only costs 1493 gate count, or 0.78 mm/sup 2/. The proposed design can meet real-time DCT/IDCT requirements of the H.263 codec system for QCIF image frame size at 10 frames/s with 4:2:0 color format. Moreover, the proposed design still possesses additional computing power for other operations when operating at 33 MHz.
Tian-Sheuan Chang, Chin-Sheng Kung, Chein-Wei Jen
IEEE Trans. Circuits Syst. Video Technol.1
1998 Low power FIR filter realization with differential coefficients and input
abstract
Most FIR filter realizations use the inputs and coefficients directly to compute the convolution. We present a low power and high speed FIR filter designs by using first order difference between inputs and various orders of differences between coefficients. This design first reformulates the FIR operations with the differences in the algorithm level. Then, in the architecture level, we adopt the distributed arithmetic (DA) architecture to exploit the probability distribution such that the power consumption can be reduced further. The design is applied to an example FIR filter to quantify the energy savings and speedup. It shows lower power consumption than the previous design with the comparable performance.
Tian-Sheuan Chang, Chein-Wei Jen
ICASSP1