Shih-Chang Hsia

dblp:75/473 · DBLP profile ↗
← Back
31ranked-venue papers
27as first author
4since 2021 · last 2026
0000-0001-9828-0773ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 21 first-authorSystems, architecture and hardware · 6 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 High resolution thermopile -Chip temperature sensor combined with dual-slope converter
Shih-Chang Hsia, Shang-Ze Yang
Integr.1
2026 A Novel Chip Design and System Integration for High-Speed Bridge Silicon-Carbide Driver
abstract
Silicon carbide (SiC) devices are widely used in electric vehicles due to their low on-resistance and excellent thermal conductivity. In this study, we design a high-speed SiC gate driver chip using the TSMC$0.18\mu $m HV process for a half-bridge driver. Both high-side and low-side driving circuits incorporate overdriving techniques to reduce turn-on time. A new signal isolator is implemented to ensure isolation between the high-side and low-side circuits. Two driver chips, one for the high-side and one for the low-side, were successfully fabricated, occupying only 1.69 mm2of silicon area. For system integration, a microprocessor controls the gate driver chip, voltage converter, and SiC devices to drive a motor. A high-performance crosstalk suppression circuit is designed to enhance system efficiency. Experimental results show that the proposed circuit reduces crosstalk levels by 87% without requiring a filtering capacitor, and a negative voltage source.
Shih-Chang Hsia, Yuan-Heng Wang, Shin-Chi Lai, Tsung-Heng Tsai
IEEE Trans. Circuits Syst. I Regul. Pap.1
2025 VLSI Architecture Design for Compact Shortcut Denoising Autoencoder Neural Network of ECG Signal
abstract
The Electrocardiogram (ECG) test detects and records cardiac-related electrical activity of the heart. The ECG test identifies and documents cardiac-related electrical activity in the heart. The use of ECG signals for cardiovascular disease nursing as a crucial component of preoperative evaluation is increasing. ECG signals need to denoise and display in a clear waveform due to the numerous noises. We have introduced Compact Shortcut Denoising Auto-encoder (CS-DAE) neural network, which reduces the noise from ECG signals. The Compact Shortcut approach compresses the features passed through the shortcut layers, which lowers the operation’s memory needs and improves the noise reduction impact. In addition, the encoder and decoder process the Pixel-Unshuffled and Pixel-Shuffled, which effectively mitigates the feature loss caused by down-sampling and up-sampling operations. As a result, the CS-DAE algorithm decreases the computation and required memory size while maintaining higher accuracy. We have used MITDB and NSTDB datasets for training and testing the proposed CS-DAE model, resulting in the average Percentage of Root Mean Square Difference (PRD) being 46.30% and the improvement of Signal-to-Noise Ratio (SNRimp) being 10.50. In addition, we have designed VLSI architect ure for the proposed CS-DAE neural network to accelerate low hardware cost and less computation. The TUL PYNQTM-Z2 development platform runs the Verilog code, which is used for VLSI architecture and has the lowest power consumption of 1.65W.
Shin-Chi Lai, Szu-Ting Wang, S. M. Salahuddin Morsalin, Jia-He Lin, Shih-Chang Hsia, Chuan-Yu Chang, Ming-Hwa Sheu
IEEE Trans. Circuits Syst. I Regul. Pap.5
2022 Fast computation of deep neural network and its real-time implementation for image recognition
abstract
Abstract The convolution is widely used for deep neural networks to extract the key features, which requires many additions and multiplications. In this study, the fast computational algorithm is presented to reduce the number of arithmetic when the accuracy is kept. The order of deep convolution is alternative to save the computational operators. To verify the performance, the proposed algorithm is embedded to the typical deep neural network VggNet. The structure of VggNet is further modified using the proposed summation and concatenation techniques to improve the computational accuracy and to reduce the processing time. Compared with the original VggNet, the simulations show that the operational FLOPs can be greatly reduced at least 50% with various datasets testing. Besides, the training time with epoch per batch can save about 10%–20%. The proposed fast algorithm can lessen the parameters and the mode size over 90%. The recognition accuracy can be improved with 1%–4% from various datasets testing. Based on the fast network, real‐time FPGA had been realized, which the hardware performance can achieve 371 GOPs with 642 DSP cores. The processing speed can achieve near to 1 k frames per second, and the real‐time recognition rate can achieve over 90%.
Shih-Chang Hsia, Szu-Hong Wang, Feng-Yang Kuo
Comput. Intell.1
2020 Fast search real-time face recognition based on DCT coefficients distribution
abstract
The authors propose an adaptive face recognition algorithm based on the discrete cosine transform (DCT) coefficients approach. For the database's establishment, the face images are pre‐processed with colour transform, hair cutting, and background removing to eliminate non‐face information. The recognised kernel applied the weights of DCT coefficient distribution with the entire image transformation, to avoid position mismatch and reduce the light effect. The key coefficients of DCT are chosen from the training database by maximum variance. The fast search mode can reject 90% weak candidates with few coefficients to fasten the processing speed. The significant coefficients weighting methods are used to enhance face features. Only using 50 coefficients per picture, the recognition rate can achieve 95% for ORL face database testing. For real‐time recognition, camera imaging is processed with algorithms using C‐programming based on Windows system. The recognition rate can achieve 95% and the speed is about nine frames per second for real‐time recognition in practice.
Shih-Chang Hsia, Szu-Hong Wang, Chia-Jung Chen
IET Image Process.1
2018 Fast-efficient algorithm of high-profile intra prediction for H.264 encoding system
abstract
This study presents a fast intra‐prediction algorithm for a high‐profile H.264 encoder. First, a pre‐decision algorithm is proposed to reject the impossible block size based on image variance. Then a fast 4 × 4 block prediction algorithm is proposed to select four possible modes from nine predicted modes based on the edge‐filtering detection. With a hierarchical approach, the 8 × 8 block prediction is based on the result of the chosen 4 × 4 block mode. This approach selects only one to five modes of H.264 coding from nine prediction modes. Following the prediction of the 16 × 16 luminance block and 8 × 8 chrominance block is mapped by the result of the 8 × 8 luminance block by checking only one to two modes. The pre‐decision algorithm, the fast 4 × 4, 8 × 8 and 16 × 16 block prediction algorithms can be combined to improve the coding speed. Simulations demonstrate that the proposed algorithm can save about 70% coding time at most for intra‐frame coding in the H.264 system while increasing only about 1% bit‐rate with a negligible peak signal‐to‐noise ratio drop.
Shih-Chang Hsia, Wing-Kwong Wong, Yen-Hung Shih
IET Image Process.1
2016 Fast-efficient shape error concealment technique based on block classification
abstract
This study presents a fast‐efficient error concealment method for recovering information related to shape. The proposed technique comprises block classification, edge direction interpolation and filtering interpolation. Missing blocks are classified into four categories: transparent, opaque, edge and isolated blocks. Most of the computation is spent on edge blocks and isolated blocks to maximise the cost and performance tradeoff. For the recovery of edge blocks, the edge slope is computed by referring to the nearest available block, from which the missing shape is interpolated parallel to the edge. Isolated blocks are dealt with using a cascade filter to approximate the actual shape. Experimental results show that the proposed method provides better cost performance in the restoration of shapes than that afforded by comparable algorithms, both in numerical parameters and the resulting shapes. The processing speed is approximately two to three times faster than previous methods and low computational load makes the proposed technique applicable to real‐time MPEG‐4 systems.
Shih-Chang Hsia, Cheng Hung Hsiao
IET Image Process.1
2015 High-performance high dynamic range image generation by inverted local patterns
abstract
This study presents an image processing algorithm capable of generating high dynamic range (HDR) images from a single frame. On the basis of liquid crystal display (LCD) backlight theory, the proposed algorithm uses an inverted local pattern approach to adjust the histogram of lighting distribution. To improve processing efficiency, sampled images are classified into three types according to exposure characteristics, which are used to obtain suitable processing factor. The inverting factor is mixed with the original sample to darken the bright areas and brighten the dark areas in the images. Brightness enhancement and auto‐gain control are then added to expand the range of the grey levels. Authors’ results demonstrate the efficacy of the proposed HDR algorithm in improving shadow details. In addition, single‐pass processing is used to reduce computational complexity, making the algorithm applicable for low‐power portable cameras and video recorders.
Shih-Chang Hsia, Ting-Tseng Kuo
IET Image Process.1
2014 Efficient scrolling videotext detection with adaptive temporal differential approach
abstract
This study presents an efficient algorithm for the detection of scrolling text on continuous frames. The proposed scheme adopts both spatial and temporal computation as well as pre‐processing to differentiate among noise, text and background information. Scrolling text is detected using temporal differentiation among inter‐frames and then the region in which the scrolling text appears is identified as a rectangle. Simulation results demonstrate that the proposed algorithm is capable of precisely differentiating scrolling text in any direction, along any of the frame boundaries without false detection or missed text. In a comparison using quantised measurements, the proposed method outperforms all competing algorithms.
Shih-Chang Hsia, Nan-Tsai Chang-Jian
IET Image Process.1
2012 Low-complexity high-quality adaptive deblocking filter for H.264/AVC system
Shih-Chang Hsia, Wei-Chih Hsu, Sheng-Chieh Lee
Signal Process. Image Commun.1
2009 Adaptive video coding control for real-time H.264/AVC encoder
Shih-Chang Hsia, Szu-Hong Wang
J. Vis. Commun. Image Represent.1
2007 VLSI Implementation of High-Performance Error Concealment Processor for TV Broadcasting
abstract
This paper presents an error concealment processor to increase the performance of the TV receiver while the decoding bit stream over error-prone channel suffers from damage. An efficient error-concealment algorithm is advised with the adaptation of the spatial interpolation and the temporal prediction to reduce the nonmatched error for high motion regions and achieve fine resolution for still or low motion regions. Based on the adaptive algorithm, we proposed a parallel VLSI architecture with pipeline scheduling for real time implementation. The error concealment processor consists of the computational core, RAM block and interface, and then is integrated to the video decoder by using a complex processing schedule. The computational sources are commonly used for various frame types processing to reduce the hardware cost. The chip occupies about 27 k gates and includes one on-chip line-buffer. The silicon area is about 9 mm2and the throughput rate can achieve 50 Mpixels/s, when implemented by 0.35-mum CMOS technology.
Shih-Chang Hsia, Shih Wen Chou
IEEE Trans. Circuits Syst. Video Technol.1
2007 Shift-Register-Based Data Transposition for Cost-Effective Discrete Cosine Transform
abstract
This paper presents a cost-effective 2-D-discrete cosine transform (DCT) architecture based on the fast row/column decomposition algorithm. We propose a new schedule for 2-D-DCT computing to reduce the hardware cost. With this approach, the transposed memory can be simplified using shift-registers for the data transposition between two 1-D-DCT units. A special shift cell with MOS circuit is designed by using the energy transferring methodology. The memory size can be greatly reduced, and the address generator and its READ/WRITE control all can be saved. For an 8 times 8-block transformation, the number of transistors is only 4 k for the shift-register array. The maximum frequency of shift-operation can achieve about 120 MHz, when implemented by 0.35-mum technology.
Shih-Chang Hsia, Szu-Hong Wang
IEEE Trans. Very Large Scale Integr. Syst.1
2006 A cost-effective line-based light-balancing technique using adaptive processing
abstract
The camera imaging system has been widely used; however, the displaying image appears to have an unequal light distribution. This paper presents novel light-balancing techniques to compensate uneven illumination based on adaptive signal processing. For text image processing, first, we estimate the background level and then process each pixel with nonuniform gain. This algorithm can balance the light distribution while keeping a high contrast in the image. For graph image processing, the adaptive section control using piecewise nonlinear gain is proposed to equalize the histogram. Simulations show that the performance of light balance is better than the other methods. Moreover, we employ line-based processing to efficiently reduce the memory requirement and the computational cost to make it applicable in real-time systems.
Shih-Chang Hsia, Ming-Huei Chen, Yu-Min Chen
IEEE Trans. Image Process.1
2006 VLSI implementation of low-power high-quality color interpolation processor for CCD camera
abstract
This paper presents a color interpolation technique for a single-chip charge-coupled device with color-filter-array format. We propose edge-direction weighting and the local gain approach to reconstruct missing color components. Simulations show that the proposed method can achieve better quality-complexity tradeoff than other algorithms. For real-time implementation, a cost-effective architecture consisting of a pipeline schedule is designed based on our new algorithm. With the time-sharing method, the VLSI architecture can interpolate various colors using a common computational kernel, reducing the circuit complexity. The prototype of the color interpolation processor has been successfully verified with a field-programmable gate array device. The chip only uses about 10K gates and two line buffers.
Shih-Chang Hsia, Ming-Huei Chen, Po-Shien Tsai
IEEE Trans. Very Large Scale Integr. Syst.1
2005 A fast efficient restoration algorithm for high-noise image filtering with adaptive approach
Shih-Chang Hsia
J. Vis. Commun. Image Represent.1
2005 Efficient light balancing techniques for text images in video presentation systems
abstract
This paper presents a novel light balancing technique to compensate uneven illumination based on adaptive gain control, which is specially designed for text images. A segmentation technique is used to remove contents from a sampling image with block-based processing. To achieve high performance, the block size is adaptively selected for different character size according to image features. Then the linear interpolation is employed to generate new values instead of text pixels to obtain a background image. Referring to this image, an adaptive technique is presented to process the sampling image to achieve light balancing. Simulations demonstrate that our light balance processing can achieve better quality than the conventional methods for camera systems.
Shih-Chang Hsia, Po-Shien Tsai
IEEE Trans. Circuits Syst. Video Technol.1
2005 Efficient adaptive error concealment technique for video decoding system
abstract
This paper presents a novel error concealment method for video decoding system. The proposed algorithm adaptively combines the spatial interpolation and the temporal prediction technique based on block variance and interframe correlation, to recover the lost data. The adaptive function depends on the scene change detection, motion distance and spatial information from the nearby blocks of the previous and current frames to determine the weighting of the spatial interpolation and the temporal compensation. Simulations demonstrate that the proposed technique can achieve well subjective and quantitative results, and outperforms all the others against which are compared. Even if the scene changes in the videos, this algorithm also can efficiently recover the damaged blocks for Intra(I), Predictive(P), and Bidirectional (B) frames.
Shih-Chang Hsia, Shyi-Chyi Cheng, Shih Wen Chou
IEEE Trans. Multim.1
2004 A High Robust Watermarking Technique Using Sub-Band Filtering
abstract
Although many watermarking algorithms have been developed in recent years, most of them concentrated on using binary data such as logos or texts. We propose an efficient algorithm of gray-level watermarks using the sub-band filtering approach. With signal processing in the frequency domain, the feature of progressive transformation is utilized to disperse the watermark into the entire image and to enhance the robustness. Simulation results show that the quality of hiding is superb with a 45 dB PSNR on the average. Moreover, our algorithm is capable to provide a reliable watermark protection even the watermarked image suffered from damages or JPEG compressions.
Shih-Chang Hsia, I-Chang Jou
MMM1
2004 An edge-oriented spatial interpolation for consecutive block error concealment
abstract
The coding scheme is currently the most commonly used process of compressing image data. However, this process often results in very serious distortion in the decoded image if the bit-stream suffers from damages. In order to address this drawback, this letter aims to develop a new error concealment technique based on edge-oriented interpolation for still image or intra-frame correction. The first step involves finding the edge direction of a lost block by using one-dimensional matching techniques from two boundaries of neighboring blocks. Then, the error pixels are recovered by weighting linear interpolation along the estimated edge direction. Afterwards, the median filter is used to recover residual damaged-pixels. Simulations demonstrate that the important edge information can be significantly recovered, and also prove that this method performs better compared to other methods that use both subjective and objective measures.
Shih-Chang Hsia
IEEE Signal Process. Lett.1
2004 A real-time chip implementation for adaptive video coding control
abstract
The paper presents an adaptive coding control for real-time video coding systems. Based on temporal correlations, the group-of-pictures (GOP) of a video sequence is split into one basic GOP (BGOP) and many adaptive GOPs (AGOPs) and then processed accordingly. The advantage of this method is to improve the coding efficiency, particularly solving the scene change problem. Even if the coding bit rate is over the budget, the coding scheme does not require re-encoding, hence it is especially suitable for real-time coding. Based on the adaptive algorithm, the chip is realized to process the functions of picture type decision, coding rate estimation, quantization scale decision, scene-change detection and coding mode decision in the operation. The gate count of this chip is only about 2000 and its silicon area is 2.85 mm/sup 2/ using a 0.35 /spl mu/m CMOS process.
Shih-Chang Hsia, Shyi-Chyi Cheng, Chung-Long Chen
IEEE Trans. Circuits Syst. Video Technol.1
2003 Fast algorithms for color image processing by principal component analysis
Shyi-Chyi Cheng, Shih-Chang Hsia
J. Vis. Commun. Image Represent.2
2003 Efficient memory IP design for HDTV coding applications
abstract
The memory intellectual property (IP) is a key component for video coding systems as using system on one chip design methodology. In this paper, cost-effective memory design and complex address generation are presented for high-definition television coding applications. The addressing method uses a bit-allocation approach to simplify the computational circuit and significantly improves the memory access speed. For the bit-allocation requirement, a new memory structure is designed using pseudoaddress decoding concept to reduce the I/O complexity and to shorten the access time. The memory IP integrated to practical video coding systems is also presented. The experiments show that the proposed memory IP can provide better performance than the conventional one.
Shih-Chang Hsia
IEEE Trans. Circuits Syst. Video Technol.1
2003 Parallel VLSI design for a real-time video-impulse noise-reduction processor
abstract
High-quality televisions (TVs) such as improved digital TV, enhanced TV, and high-definition TV have become popular in recent years. However, impulse noise affects TV broadcasts. This paper proposes an efficient noise-removal algorithm using an adaptive digital signal-processing approach. Simulations have demonstrated that the new adaptive algorithm could efficiently reduce impulse noise even in highly corrupted images. In order to achieve real-time implementation, a cost-effective architecture is proposed using a parallel structure and pipelined processing. The proposed processor can achieve the throughput rate of 45M pixels/s using only 4k gates and two line buffers. Unlike median-filtering chips, this processor provides better filtering quality and its circuit is much less complex.
Shih-Chang Hsia
IEEE Trans. Very Large Scale Integr. Syst.1
2002 VLSI implementation for low-complexity full-search motion estimation
abstract
Although many ASICs for motion estimation have been developed, either the chip complexity is too high or the optimal accuracy was not achieved. In this study, an adaptive full-search algorithm is presented to reduce the searching complexity with a temporal correlation approach. The efficiency of the proposed full search can be promoted about 5-10 times in comparison with the conventional full search while the searching accuracy remains intact. Based on the adaptive full-search algorithm, a real-time VLSI chip is regularly designed by using the module base. For MPEG-2 applications, the computational kernel only uses eight processing elements to meet the speed requirement. The processing rate of the proposed chip can achieve 53 K blocks/s to search from -127 to +127 vectors, using only 8 K gates.
Shih-Chang Hsia
IEEE Trans. Circuits Syst. Video Technol.1
2001 A recursive full search algorithm based on temporal correlation
Shih-Chang Hsia, Chien-Cheng Tseng
Signal Process.1
2000 Computation of fractional derivatives using Fourier transform and digital FIR differentiator
Chien-Cheng Tseng, Soo-Chang Pei, Shih-Chang Hsia
Signal Process.3
1997 Efficient postprocessor for blocky effect removal based on transform characteristics
abstract
In this paper, we propose transform-domain algorithms to effectively classify the characteristics of blocks and estimate the strength of the blocky effect. The transform-domain algorithms require much lower computational complexity and much less memory than the spatial ones. Along with the estimated blocky strength, we also propose an adaptive finite impulse response (FIR) filter to effectively remove the blocky effect. Simulation results show that the proposed algorithms sharply preserve the edge information and greatly improve the decoded image quality.
Shih-Chang Hsia, Jar-Ferr Yang, Bin-Da Liu
IEEE Trans. Circuits Syst. Video Technol.1
1996 A parallel video converter for displaying 4: 3 images on 16: 9 HDTV receivers
abstract
The 16/9 widescreen TV is becoming popular in the marketplace; however, its aspect ratio is not compatible with the conventional 4/3 SDTV program. We propose a nonuniform conversion to efficiently display 4/3 SDTV programs on a 16/9 high-definition television (HDTV) monitor, for real time HDTV receivers, a parallel video converter is also proposed using nonuniform interpolation. To trade off complexity and performance, we only use one field-memory and one line-buffer for motion adaptation to interpolate the pixels, which can be finished by the parallel computation kernel in one clock cycle. In order to combine parallel computation, the memory-slice method is employed, which is efficiently controlled by a state machine for nonuniform interpolation. With parallel conversion, a high-speed video converter can be achieved for advanced HDTV receivers.
Shih-Chang Hsia, Bin-Da Liu, Jar-Ferr Yang, Chien-Hsiu Huang
IEEE Trans. Circuits Syst. Video Technol.1
1995 VLSI implementation of parallel coefficient-by-coefficient two-dimensional IDCT processor
abstract
We propose the pipelined VLSI architecture and modular design to realize the coefficient-by-coefficient two-dimensional inverse discrete cosine transform (2-D IDCT) suggested by Yang, Bel, and Hisa (see ibid., vol.5, no.1, p.25-30, 1995). Based on parallel processing, the architecture of this chip is designed with a five-stage pipeline to meet the speed requirement for real-time applications. The key building modules of this chip include a generator of cosine angle index, a pipelined multiplier, and a matrix accumulator core. Satisfying the IEEE standard of 2-D IDCT in computational accuracy, this IDCT chip, which can work at a clock rate of higher than 50 MHz, is implemented by the CMOS technology in a reasonable die size. With modular and regular structures, the IDCT VLSI chip can be operated in a progressive transform mode. In a real video decoding system, the average pixel-rate of the proposed 2-D IDCT chip achieves over 150 MHz for decoding intraframes and up to 400 MHz for decoding interframes.>
Shih-Chang Hsia, Bin-Da Liu, Jar-Ferr Yang, Bor-Long Bai
IEEE Trans. Circuits Syst. Video Technol.1
1995 An efficient two-dimensional inverse discrete cosine transform algorithm for HDTV receivers
abstract
In this paper, we propose an efficient two-dimensional inverse discrete cosine transform (2-D IDCT) algorithm based upon coefficient-by-coefficient implementations. The algorithm, which is developed in the use of the symmetrical properties of computational kernel matrices, can effectively combine with the run length code (RLC) decoder in progressive manners. The proposed method requires only N/2 multipliers in an 'N/spl times/N' 2-D IDCT processor. By evaluating the efficiency performance, the processor can complete 2N/M pixels per multiplier per system cycle where M is the number of non-zero quantized DCT coefficients. Due to its inherent parallel and coefficient-by-coefficient characteristics, the proposed efficient 2-D IDCT algorithm is suitable for the future digital HDTV receivers and recording systems.>
Jar-Ferr Yang, Bor-Long Bai, Shih-Chang Hsia
IEEE Trans. Circuits Syst. Video Technol.3