Shen-Fu Hsiao

dblp:90/1566 · DBLP profile ↗
← Back
34ranked-venue papers
30as first author
6since 2021 · last 2024
0000-0002-4627-570XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 22 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 first-authorTheory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2024 Hardware Accelerator for MobileViT Vision Transformer with Reconfigurable Computation
abstract
With the great success of the Transformer model in Natural Language Processing (NLP), Vision Transformer (ViT) was proposed achieving comparable performance to traditional Convolutional Neural Network (CNN) models in tasks such as image classification and object detection. This paper focuses on the acceleration of a new lightweight hybrid model, named MobileViT, which has less computation complexity and higher accuracy compared with ViT and other CNN-based lightweight models such as MobileNets. We introduce an adaptive systolic array (SA) design with a flexible shape size, called LEGO SA, that enhances the efficiency of hardware utilization and memory accesses during standard convolution, Depth-wise Separable Convolution (DWC), and self-attention operations. Furthermore, matrix transpose in self-attention is implemented efficiently with significantly reduced wastage of execution time, memory buffers, and power consumption. The proposed MobileViT hardware accelerator with 112KB on-chip buffers occupies an area of just 1.64mm^2 on the TSMC 40nm process, and achieves a performance of 1.2 TOPS at 600 MHz with energy efficiency of 5.34 TOPS/W.
Shen-Fu Hsiao, Tzu-Hsien Chao, Yen-Che Yuan, Kun-Chih Chen
ISCAS1
2024 Neural Network Acceleration Using Digit-Plane Computation with Early Termination
abstract
Deep neural network (DNN) hardware accelerator designs can be divided into two categories, bit-parallel and bit-serial, depending on whether the input of the multiplication is in bit-parallel or bit-serial style. Bit-serial DNN designs with only shift-add operations are more flexible for computation when the per-layer bit-widths are different in aggressive model quantization. Rectified Linear Unit (ReLU) is a common non-linear activation function after each layer of DNN computation where negative values are replaced by zeros. Thus, early-termination could be employed to reduce ineffective computation, resulting in improved speed performance and reduced energy consumption. In this paper, we propose a new bit-serial computation style by re-organizing data in bit-plane order. We observe that bit-plane computation allows more efficient early-termination. The proposed bit-plane DNN design is also extended to digit-plane through radix-4 Booth recoding, leading to a smaller area cost. Experimental results show that the speed performance of the proposed designs is 2.25x and 1.48x times that of the bit-parallel and bit-serial designs respectively when executing the VGG-16 model.
Shen-Fu Hsiao, Hou-Chun Kuo, Yu Kuo, Kun-Chih Chen
ISCAS1
2022 Dynamically Swappable Digit-Serial Multi-Precision Deep Neural Network Accelerator with Early Termination
abstract
We propose a digit-serial Deep Neural Network (DNN) hardware accelerator design using flexible digit-serial multiplication and early termination mechanism which supports multi-precision computation with different bit-widths in different DNN layers. The dynamically swappable digit-serial multiplier design employs radix-4 Booth recoding for either weights or activations whichever having the smaller quantized bit-width and performs digit-serial sequential multiplication with one input in the Booth-recoded digit-serial form and the other in bit-parallel form. Furthermore, we add a so-called early-termination technique that stops DNN computation whenever the partially accumulated sums are smaller than pre-scribed thresholds, leading to additional speedup when the Rectified Linear Unit (ReLU) is used as the non-linear activation function. Implementation results show that the proposed design can achieve more than 50% speedup compared with the baseline 16-bit fixed-point design.
Shen-Fu Hsiao, Hung-Ching Li, Yu-Che Yen, Po-Chang Li
ISCAS1
2022 A 40.96-GOPS 196.8-mW Digital Logic Accelerator Used in DNN for Underwater Object Recognition
abstract
This investigation presents a digital logic accelerator (DLA) design of a neural network hardware that utilizes output reuse. The DLA is used in the detection mechanism of underwater objects that was deployed in an underwater vehicle. A modified YoloV3-tiny network was also implemented to detect more than 20 underwater objects. The proposed DLA uses processing units that have parallel architectures of output windows, and output channels. Moreover, a new Inter-Controller is designed to control the direct memory access (DMA) together with a new Reshape module to improve the performance and power efficiency. A detailed description of the design as well as the measurements on silicon are presented. The chip is realized using a typical 180-nm CMOS process. It showed a performance result of 40.96 GOPS and the power consumption is 196.8 mW. The DLA was tested to demonstrate 19.88 frames per second and 40.96 GOPS.
Chua-Chin Wang, Ralph Gerard B. Sangalang, Chien-Ping Kuo, Hsin-Che Wu, Yi Hsu, Shen-Fu Hsiao, Chia-Hung Yeh
IEEE Trans. Circuits Syst. I Regul. Pap.6
2021 Comparison of Digit-Serial and Bit-Level Designs for Acceleration of Convolutional Neural Network Computation
abstract
Since the bit-width requirements of weights in different neural network (NN) layers might be different, we can design more efficient NN accelerators considering the dynamic adjustment of bit-width in the involved hardware components. In this paper, two different bit-width adjustable design approaches are designed, digital-serial (DS) and bit-level (BL). In the DS design, the multiplication-accumulation is performed digit-by-digit after radix-4 Booth recoding, leading to fewer execution cycles compared to the conventional bit-serial design. In the BL design, the arithmetic components of higher bit-width is constructed-from sub-units of smaller bit-width, leading to more execution parallelism in case of smaller bit-width. We compare both DS and BL designs for the quantized VGG NN models with previous alternative designs and find that the DS design has significant speed-up against the BL design and other previous NN hardware accelerators which adopt uniform bit-width for all layer of computations.
Shen-Fu Hsiao, Jian-Ming Chen, Yu-Hong Chen, Hung-Ching Li, Yi Hsu
ISCAS1
2021 Efficient Quantization and Multi-Precision Design of Arithmetic Components for Deep Learning
abstract
We present a quantization algorithm to find the per-layer bit-width in deep neural networks (DNN) models and then propose multi-precision designs for two fundamental DNN arithmetic operations: multiplication and non-linear activation function computation. The multi-precision multiplier design with truncated partial product bits supports four precision modes (4-bit, 8-bit, 12-bit, and 16-bit) with shared hardware resource so that the power consumption in low-precision modes can be reduced by turning off unnecessary hardware circuits. For the evaluation of non-linear activation functions, we compare three different approaches in various precision requirements and observe that the best design method depends on the bit-accuracy.
Shen-Fu Hsiao, Yu-Chang Chen, Yu-Che Yen
ISCAS1
2020 Sparsity-Aware Deep Learning Accelerator Design Supporting CNN and LSTM Operations
abstract
Sparsity of data and weights appears in many convolution neural networks (CNN) and recurrent neural networks such as long-short term memory (LSTM). In this paper, we design a sparsity-aware deep learning hardware accelerator exploiting both data and weight sparsity in CNN and LSTM models. The proposed hardware accelerator significantly reduces memory accesses and computations, leading to much lower power consumption.
Shen-Fu Hsiao, Hsuan-Jui Chang
ISCAS1
2020 Hardware Efficient Function Computation Based on Optimized Piecewise Polynomial Approximation
abstract
This paper presents an optimization method for function computation based on piecewise polynomial approximation with truncated multipliers. The proposed design jointly considers all the error sources, including the approximation errors, quantization errors, truncation errors, and rounding errors. Thus, the total error budget can be utilized more efficiently and the bit widths of the hardware components can be optimized, leading to significant area improvement. Due to the trade-off between the size of lookup tables (LUT) and the area of arithmetic components, we consider two minimization goals: LUT size and total area. Experimental results show that the proposed optimization method has chance of finding better hardware design variables compared with state-of-the-art designs.
Shen-Fu Hsiao, Chia-Yang Wong, Yu-Chang Chen
ISCAS1
2019 Dual-Precision Acceleration of Convolutional Neural Network Computation with Mixed Input and Output Data Reuse
abstract
Memory access dominates power consumption in hardware acceleration of deep neural networks (DNN) computation due to the movement of huge data and weights. This paper design a DNN accelerator using mixed input and output data reuse scheme to achieve balance between internal memory size and memory access amount, two contradictory design goals in resource limited embedded systems. First, analytical forms for memory size and accesses are derived for different data reuse methods in DNN convolution. After comparing the analysis results across different convolutional layers of the VGG-16 model with different levels of hardware parallelism, we implement a low-cost DNN hardware accelerator using mixed input and output data reuse scheme with 32 processing elements (PEs) operating in parallel. Furthermore, the design supports two precision modes (8-bit and 16-bit) allowing variable precision requirements across DNN layers, resulting in more efficient computation compared with single-precision designs through sharing of hardware resource.
Shen-Fu Hsiao, Pei-Hsuan Wu, Jien-Min Chen, Kun-Chih Chen
ISCAS1
2018 Architectural Exploration of Function Computation Based on Cubic Polynomial Interpolation with Application in Deep Neural Networks
abstract
Computation of function values is widely used in applications of computer vision, communication, and artificial neural networks. Piecewise polynomial approximation (PPA) with lookup tables (LUT) assisted by arithmetic units is usually adopted in hardware implementation. In high-precision requirements, the degree of per-segment polynomial is usually increased in order to reduce the size of LUT, but the cost of arithmetic units usually dominates the entire design. This paper analyzes different architectural designs of function computation units based on degree-three polynomial interpolation where the arithmetic units include truncated multipliers, squaring unit (squarer), and cubic unit (cuber). Significant savings in area and power are achieved by selecting architectures without cuber. Furthermore, we design reconfigurable functional computation units supporting four different precision modes with shared hardware resources. The designs are applied to the computation of several arithmetic functions including some popular activation functions for deep neural networks where the bit-accuracy requirement across network layer might vary.
Shen-Fu Hsiao, Yu-Chang Chen, Hsiang-Hao Liang
DSD1
2018 Design and Implementation of Low-Cost LK Optical Flow Computation for Images of Single and Multiple Levels
abstract
Lucas-Kanade (LK) optical flow algorithm is widely used for moving object detection and tracking by computing the motion vectors of pixels in image sequences. Due to the high computation complexity, optical flow computation is one of the crucial operations in many computer vision applications. This paper presents a low-cost hardware implementation of the LK optical flow algorithm. In particular, we design a low-cost divider used in the matrix inversion, leading to significant reduction in delay and area for calculation of small optical flow vectors. Furthermore, we also design a pyramidal LK optical flow computation unit that can process images of different sizes in order to increase the magnitude range of optical flow vectors.
Shen-Fu Hsiao, Chen-Yen Tsai
DSD1
2018 Optimization of Lookup Table Size in Table-Bound Design of Function Computation
abstract
Computation of function values is critical for designing datapath units in digital signal processors (DSP) and graphics processing units (GPU). Most function computation methods requires lookup tables (LUT) and simple arithmetic components. In table-bound methods, LUT size takes a significant portion of total hardware area, in particular for high-precision applications. This paper presents a new multi-level lossless table decomposition to further reduce total table size in a recently proposed hierarchical multipartite (HMP) table method which is a generation of the prior bipartite/multipartite table-addition methods. Furthermore, hardware design parameters are optimized by jointly considering all the error sources. Experimental results shows that the proposed design has the chance of further reducing the total table size of HMP, which already has significant table-size saving over previous similar designs.
Shen-Fu Hsiao, Kun-Chih Chen, Yi-Hau Chen
ISCAS1
2017 Hierarchical Multipartite Function Evaluation
abstract
Function evaluation is an important arithmetic computation in many signal processing applications, such as special function units in modern graphics processing units (GPUs). Hardware implementations of function evaluation usually consists of lookup tables (LUT) and some simple arithmetic units of multipliers and/or adders. LUT usually takes a significant portion of total area cost, especially when function evaluators are allowed to compute several different arithmetic functions with shared arithmetic units where evaluation of each function needs separate LUT. In this paper, we focus on the category of table-addition (TA) function evaluators that are composed of two types of LUT, table of initial values (TI) and table of offset values (TO), followed by a multi-operand adder. It has been shown that multipartite table method (MP) has significant improvement over prior similar designs such as symmetric bipartite table methods (SBTM) and symmetric table addition methods (STAM) for applications with low-to-medium precision requirements. This paper presents an extension of MP, called hierarchical multipartite (HMP), which further reduces total table size by applying several levels of table decompositions. Furthermore, we perform the bit-width optimization by jointly considering the impacts of all error sources during the search of best table decompositions, leading to more efficient hardware design. Besides, a new lossless decomposition of TI is presented, resulting in additional saving of table size without incurring any extra errors. Experimental results show that the proposed design can efficiently reduce the total area cost in ASIC and FPGA implementations.
Shen-Fu Hsiao, Chia-Sheng Wen, Yi-Hau Chen, Kuei-Chun Huang
IEEE Trans. Computers1
2014 Compression of Lookup Table for Piecewise Polynomial Function Evaluation
abstract
Function evaluation is an important arithmetic computation in many signal processing applications, such as the special function unit in modern graphics processing units (GPUs). Lookup table (LUT) usually takes a significant portion of total area in function evaluation using piecewise polynomial approximation. Many papers have proposed various approaches to reduce table size without sacrificing precision requirement. This paper presents new LUT compression methods that do not introduce extra errors but can effectively further reduce the table size in the piecewise polynomial approximation with uniform or non-uniform segmentations.
Shen-Fu Hsiao, Chia-Sheng Wen, Po-Han Wu
DSD1
2014 Design and Implementation of Multiple-Vehicle Detection and Tracking Systems with Machine Learning
abstract
A vehicle detection system is realized in two stages: hypothesis generation (HG) and hypothesis verification (HV). HG adopts frame division and shadow detection to find possible candidates of vehicles within a plausible region of the image frame. Then, during HV, object ratio constraint is first used to eliminate unreasonable hypotheses. Afterward, based on the training results of the support vector machine (SVM) with the proposed vehicle feature extraction, the kernel function is found and employed in the classifier to find the real vehicles. Both pure software and combined software/hardware (SW/HW) implementations of the proposed vehicle detection system are presented where the combined SW/HW implementation can achieve correct detection rate of more than 90% in most test conditions.
Shen-Fu Hsiao, Guan-Fu Yeh, Je-Chi Chen
DSD1
2014 Design of low-leakage multi-port SRAM for register file in graphics processing unit
abstract
This paper combines several low-leakage and low-cost techniques to design multi-port static random access memory (SRAM) for register file in a vertex shader processor for OpenGL ES 2.0 graphics applications. First, precharge control is employed to eliminate unnecessary precharge operations. Then, dynamic forward body-bias control for leakage reduction is proposed to adjust the threshold voltage of transistors depending on whether memory cell is accessed or idle. Different power-gating methods are presented for SRAM cells and for sense amplifiers. The thin-cell-like lithography-friendly layout style is adopted for the multi-port SRAM cell to achieve more compact and flexible layout with easy addition of read ports. The cross-coupled inverters in the SRAM cell have asymmetric design to reduce area and also increase the writability. A 32×128-bit register file supporting 5-read and 1-write operations is implemented and verified, which has smaller area and lower leakage compared with that from ARM register file generator.
Shen-Fu Hsiao, Pu-Cheng Wu
ISCAS1
2013 Design of Hardware Function Evaluators Using Low-Overhead Nonuniform Segmentation With Address Remapping
abstract
In the piecewise function evaluation with polynomial approximation, nonuniform segmentation can effectively reduce the size of lookup tables for some arithmetic functions compared to uniform segmentation approaches, at the cost of the extra segment address (index) encoder that results in area and delay overhead. Also, it is observed that the nonuniform segmentation reflects a design tradeoff between the ROM size and the area cost of the subsequent arithmetic computation hardware. In this paper, we propose a new nonuniform segmentation method that searches for the optimal segmentation scheme with the goal of minimized ROM, total area, or delay. For some high-variation arithmetic functions, the proposed segmentation method achieves significant area reduction compared to the uniform segmentation method. We also demonstrate the design tradeoff among uniform and nonuniform segmentation, and degree-one and degree-two polynomial approximations, with respect to precision ranging from 12 to 32 bits for the elementary function of reciprocal.
Shen-Fu Hsiao, Hou-Jen Ko, Yu-Ling Tseng, Wen-Liang Huang, Shin-Hung Lin, Chia-Sheng Wen
IEEE Trans. Very Large Scale Integr. Syst.1
2012 Low latency design of Depth-Image-Based Rendering using hybrid warping and hole-filling
abstract
A low-latency design of Depth-Image-Based Rendering (DIBR) stereoscopic image generation hardware is presented. We propose a new algorithm that performs the operations of warping and hole-filling in parallel so that the overall computation latency is significantly reduced. Furthermore, two new approaches, “raised disparity around edge” and “horizontal mirroring” are employed to reduce visual artifacts of synthesized virtual images. Experimental results show that the new design can effectively reduce the computation time and improve the synthesized image quality compared to previous DIBR designs.
Shen-Fu Hsiao, Jin-Wen Cheng, Wen-Ling Wang, Guan-Fu Yeh
ISCAS1
2010 A new non-uniform segmentation and addressing remapping strategy for hardware-oriented function evaluators based on polynomial approximation
abstract
This paper presents a new non-uniform segmentation method for arithmetic function evaluation based on polynomial approximations. The merging of several uniform segments can reduce the required ROM size compared with normal uniform segmentation. The previously proposed hierarchical segmentation turns out to be special cases of this new approach. In general, non-uniform segmentation leads to irregular address indexing that needs extra computation hardware. Here, an address rearrangement and mapping method is proposed that does not need any additional address computation, and thus the critical path delay can be reduced. Experimental results show that this new segmentation and address remapping method can efficiently reduce the table size for some elementary arithmetic functions, such as log2x and 1/x.
Hou-Jen Ko, Shen-Fu Hsiao, Wen-Liang Huang
ISCAS2
2009 An 8.69 Mvertices/s 278 Mpixels/s tile-based 3D graphics SoC HW/SW development for consumer electronics
abstract
This paper presents an 8.69 Mvertices/s, 278 Mpixels/s, 15.7 mm2tiled-based 3D graphics SoC HW/SW supporting OpenGL ES 1.0 running at 139 MHz. The SoC also includes embedded circuitry to monitor run time characteristics, detect bus protocol error/inefficiency, and capture bus traces at various abstraction levels with compression ratio up to 98%.
Liang-Bi Chen, Ruei-Ting Gu, Wei-Sheng Huang, Chien-Chou Wang, Wen-Chi Shiue, Tsung-Yu Ho, Yun-Nan Chang, Shen-Fu Hsiao, Chung-Nan Lee, Ing-Jer Huang
ASP-DAC8
2008 Area oriented pass-transistor logic synthesis using buffer elimination and layout compaction
abstract
This paper presents a cell-based AISC design flow where the traditional CMOS cell library is replaced by pass-transistor logic (PTL) cell library. In particular, we develop an automatic PTL logic synthesizer to perform area-oriented synthesis by exploiting the characteristics of the PTL cell circuits. Two methods are used to reduce the area cost. The first method, called buffer elimination, for the pre-layout area minimization is to reduce the redundant inverters in the gate-level netlist during the logic mapping stage and results in an area saving of more than 50%. The second method, called layout compaction, is to reduce the layout area in the physical level by considering the design rules imposed on the PTL cell circuits, and lead to an additional 30% area saving.
Shen-Fu Hsiao, Ming-Yu Tsai, Chia-Sheng Wen
ISCAS1
2008 An automatic hardware generator for special arithmetic functions using various ROM-based approximation approaches
abstract
In this paper we develop an automatic generator to produce hardware units that compute various single-value functions using table-method approaches and compare their performance and area costs with other alternative implementations. Experimental results show that making choices between these approaches depends on the tradeoffs between speed, area and especially the required output precision. In particular, the proposed hardware function units produced have better delay and/or area cost compared with those available from Synopsys design Ware library, a popular commercial library for generating various arithmetic units.
Shen-Fu Hsiao, Ping-Chung Wei, Ching-Pin Lin
ISCAS1
2004 A memory-efficient and high-speed sine/cosine generator based on parallel CORDIC rotations
abstract
The sine/cosine function generator is based on parallelization of the original CORDIC algorithm by predicting all the rotation directions directly from the binary bits of the initial input angle. Unlike previous approaches that require complicated circuits or exponentially increased ROM, our proposed architecture has a relatively simple prediction scheme through an efficient angle recoding. The critical path delay is also reduced by utilizing the predicted rotation directions to design an efficient multioperand carry-save addition structure.
Shen-Fu Hsiao, Yu Hen Hu, Tso-Bing Juang
IEEE Signal Process. Lett.1
2003 Design and implementation of a video-oriented network-interface-card system
abstract
We design a specific Ethernet network interface card (NIC) for accelerating the video delivery by offloading the overheads of protocol headers identification/appending and CRC/checksums calculation, and speeding video bit streams with a dedicated video interface. Compared with the same operations of a 50MHz ARM micro-controller, the NIC system saves 47,000 ns per frame. This NIC card also supports the coexistence of the IPv4 and IPv6 standard for the future extension. Both FPGA prototyping and 0.35um cell-based design of the specific NIC system are given.
Mingchih Chen, Shen-Fu Hsiao, Cheng-Hsien Yang
ASP-DAC2
2001 A new hardware-efficient algorithm and architecture for computation of 2-D DCTs on a linear array
abstract
A new recursive algorithm with hardware complexity of O(log/sub 2/N) is derived for fast computation of N/spl times/N 2-D discrete cosine transforms (2-D DCTs). It first converts the original 2-D data matrices into 1-D vectors and then employs different partition methods for the input and output indices in the 1-D vector space. Afterward, the algorithm computes the corresponding 2-D complex DCT (2-D CCT) and then uses a post-addition step to produce simultaneously two 2-D DCT outputs. The decomposed form of the 2-D recursive algorithm looks like a radix-4 fast Fourier transform algorithm. The common entries in each row of the butterfly-like matrix are factored out in order to reduce the number of multipliers needed during implementation. A new linear architecture for the derived algorithm is presented which leads to a hardware-efficient architectural design requiring only log/sub 2/N complex multipliers plus 3log/sub 2/N complex adders/subtractors for the computation of a 2-D N/spl times/N CCT.
Shen-Fu Hsiao, Wei-Ren Shiue
IEEE Trans. Circuits Syst. Video Technol.1
2000 Low-cost unified architectures for the computation of discrete trigonometric transforms
abstract
New hardware-oriented matrix formulations and their corresponding VLSI architectures are proposed for both radix-2 and radix-4 fast algorithms of a variety of discrete trigonometric transforms (DXTs), including the discrete Fourier transform (DFT), discrete cosine transform (DCT), discrete sine transform (DST), and discrete Hartley transform (DHT). All the DXTs have the same complex kernel operation consisting of products of diagonal matrices and band matrices with equally spaced nonzero diagonals. New linear arrays based on the systolic mapping of the common kernel operation are designed which lead to hardware-efficient architectures requiring much fewer processing elements compared to other previously proposed unified DXT architectures of the same throughput rate.
Shen-Fu Hsiao, Wei-Ren Shiue
ICASSP1
2000 High-performance multiplexer-based logic synthesis using pass-transistor logic
abstract
An automatic logic/circuit synthesizer is developed which takes as input several Boolean functions and generates netlist output with basic composing cells from the pass-transistor cell library containing only two types of cells: 2-to-1 multiplexers and inverters. The synthesis procedure first constructs efficient binary decision diagrams (BDDs) for these Boolean functions considering both multi-function sharing and minimum width. Each node in the BDD trees can be realized by a 2-to-1 multiplexer (MUX) designed with pass-transistor logic. Then inverters are inserted along all the MUX paths in order to improve the speed performance and to alleviate the voltage-drop problem. Compared to the recently proposed pass-transistor based top-down design, our synthesizer has better speed and area performance due to the reduced number of cascaded inverters.
Shen-Fu Hsiao, Jia-Siang Yeh, Da-Yen Chen
ISCAS1
1999 A high-throughput, low power architecture and its VLSI implementation for DFT/IDFT computation
abstract
A recursive algorithm for computation of both forward and backward DFT has been proposed where the common entries in the decomposed matrices are factored out in order to reduce the number of multipliers needed during implementation. The derived algorithm is essentially the band-matrix-vector multiplication with matrix bandwidth of 3. By exploiting the heterogeneous dependency graphs for the matrix-vector multiplication and using an efficient mapping technique, only log/sub 2/N adders and log/sub 2/N-1 multipliers are needed to compute the DFT of size N, a great saving from a previously proposed systolic architecture which calls for 3log/sub 2/N adders and 3log/sub 2/N multipliers. Furthermore, due to the simplicity and regularity of the architectures, it is possible to design a low power processor by turning off the hardware components of no operation at proper time steps. VLSI implementation of the DFT/IDFT processor with distributed finite state machine (FSM) for timing control is also presented.
Shen-Fu Hsiao, Wei-Ren Shiue
ICASSP1
1999 New hardware-efficient algorithm and architecture for the computation of 2-D DCT on a linear systolic array
abstract
A new recursive algorithm for fast computation of two-dimensional discrete cosine transforms (2-D DCT) is derived by converting the 2-D data matrices into 1-D vectors and then using different partition methods for the time and frequency indices. The algorithm first computes the 2-D complex DCT (2-D CCT) and then produces two 2-D DCT outputs simultaneously through a post-addition step. The decomposed form of the 2-D recursive algorithm looks very like a radix-4 FFT algorithm and is in particular suitable for VLSI implementation since the common entries in each row of the butterfly-like matrix are factored out in order to reduce the number of multipliers. A new linear systolic architecture is presented which leads to a hardware-efficient architectural design requiring only logN multipliers plus 3logN adders/subtractors for the computation of two N/spl times/N DCTs.
Shen-Fu Hsiao, Wei-Ren Shiue
ICASSP1
1997 New unified VLSI architectures for computing DFT and other transforms
abstract
Fast computation of DFT and other popular transforms is essential in high-speed DSP applications. This paper proposes new architectures with low hardware cost and high throughput rate. The new architectures are very suitable for VLSI implementation since they are very regular and require much fewer complex multipliers compared to the recently proposed approaches. Furthermore, the same architectures may be exploited to compute a variety of frequently-used transforms.
Shen-Fu Hsiao, Chung-Yi Yen
ICASSP1
1995 Adaptive Jacobi method for parallel singular value decompositions
abstract
The Jacobi method has been used on special-purpose multiprocessor VLSI systems for parallel singular value decomposition (SVD) of dense matrices, and CORDIC processors are often used as the basic processing elements to implement the two-sided rotations, the fundamental operations in the Jacobi method. Generalizations of the original CORDIC algorithm to multi-dimensional spaces have been used in the SVD of complex matrices to achieve faster computation speed. A further speed-up of more than 2 can be gained by gradually refining the resolution of the CORDIC algorithms used in the Jacobi method.
Shen-Fu Hsiao
ICASSP1
1995 Householder CORDIC Algorithms
abstract
Matrix computations are often expressed in terms of plane rotations, which may be implemented using COordinate Rotation Digital Computer (CORDIC) arithmetic. As matrix sizes increase multiprocessor systems employing traditional CORDIC arithmetic, which operates on two-dimensional (2D) vectors, become unable to achieve sufficient speed. Speed may be increased by expressing the matrix computations in terms of higher dimensional rotations and implementing these rotations using novel CORDIC algorithms-called Householder CORDIC-that extend CORDIC arithmetic to arbitrary dimensions. The method employed to prove the convergence of these multidimensional algorithms differs from the one used in the 2D case. After a discussion of scaling factor decomposition, range extension and numerical errors, VLSI implementations of Householder CORDIC processors are presented and their speed and area are estimated. Finally, some applications of the Householder CORDIC algorithms are listed.>
Shen-Fu Hsiao, Jean-Marc Delosme
IEEE Trans. Computers1
1994 Parallel processing of complex data using quaternion and pseudo-quaternion CORDIC algorithms
abstract
Many statistical signal processing algorithms operating on real data are based on rotations on two-dimensional (2-D) real vectors. The original, 2-D CORDIC algorithms are well suited to parallel implementations of these algorithms. However they are not as well suited to the operations on 2-D complex vectors that arise when dealing with complex data. We briefly describe extensions of the 2-D CORDIC algorithms to 4-D space and show how to use them to speed up parallel computations when solving overdetermined complex linear systems and Hermitian systems.>
Shen-Fu Hsiao, Jean-Marc Delosme
ASAP1
1991 The CORDIC Householder algorithm
abstract
A novel n-dimensional (n-D) CORDIC algorithm for Euclidean and pseudo-Euclidean rotations is proposed. This algorithm is closely related to Householder transformations. It is shown to converge faster than CORDIC algorithms developed earlier for n=3 and 4. Processor architectures for the algorithm are presented. The area and time performance of n-D CORDIC processors are evaluated. For a comparable time performance, the processors require significantly less area than parallel Householder processors. Furthermore, arrays of n-D Euclidean CORDIC processors are shown to speed up the QR decomposition of rectangular matrices by a factor of n-1 in comparison with a 2-D CORDIC processor array.>
Shen-Fu Hsiao, Jean-Marc Delosme
IEEE Symposium on Computer Arithmetic1