EDBT 2026 Demo / reviewers in the wild / expert
Shin-Chi Lai
dblp:52/10136
· DBLP profile ↗
7ranked-venue papers
5as first author
3since 2021 · last 2026
0000-0003-0011-3649ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 5 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Novel Chip Design and System Integration for High-Speed Bridge Silicon-Carbide DriverabstractSilicon carbide (SiC) devices are widely used in electric vehicles due to their low on-resistance and excellent thermal conductivity. In this study, we design a high-speed SiC gate driver chip using the TSMC$0.18\mu $m HV process for a half-bridge driver. Both high-side and low-side driving circuits incorporate overdriving techniques to reduce turn-on time. A new signal isolator is implemented to ensure isolation between the high-side and low-side circuits. Two driver chips, one for the high-side and one for the low-side, were successfully fabricated, occupying only 1.69 mm2of silicon area. For system integration, a microprocessor controls the gate driver chip, voltage converter, and SiC devices to drive a motor. A high-performance crosstalk suppression circuit is designed to enhance system efficiency. Experimental results show that the proposed circuit reduces crosstalk levels by 87% without requiring a filtering capacitor, and a negative voltage source. Shih-Chang Hsia, Yuan-Heng Wang, Shin-Chi Lai, Tsung-Heng Tsai |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | VLSI Architecture Design for Compact Shortcut Denoising Autoencoder Neural Network of ECG SignalabstractThe Electrocardiogram (ECG) test detects and records cardiac-related electrical activity of the heart. The ECG test identifies and documents cardiac-related electrical activity in the heart. The use of ECG signals for cardiovascular disease nursing as a crucial component of preoperative evaluation is increasing. ECG signals need to denoise and display in a clear waveform due to the numerous noises. We have introduced Compact Shortcut Denoising Auto-encoder (CS-DAE) neural network, which reduces the noise from ECG signals. The Compact Shortcut approach compresses the features passed through the shortcut layers, which lowers the operation’s memory needs and improves the noise reduction impact. In addition, the encoder and decoder process the Pixel-Unshuffled and Pixel-Shuffled, which effectively mitigates the feature loss caused by down-sampling and up-sampling operations. As a result, the CS-DAE algorithm decreases the computation and required memory size while maintaining higher accuracy. We have used MITDB and NSTDB datasets for training and testing the proposed CS-DAE model, resulting in the average Percentage of Root Mean Square Difference (PRD) being 46.30% and the improvement of Signal-to-Noise Ratio (SNRimp) being 10.50. In addition, we have designed VLSI architect ure for the proposed CS-DAE neural network to accelerate low hardware cost and less computation. The TUL PYNQTM-Z2 development platform runs the Verilog code, which is used for VLSI architecture and has the lowest power consumption of 1.65W. Shin-Chi Lai, Szu-Ting Wang, S. M. Salahuddin Morsalin, Jia-He Lin, Shih-Chang Hsia, Chuan-Yu Chang, Ming-Hwa Sheu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | High-Accuracy and Low-Multiplication Recursive Discrete Cosine Transform Algorithm Design and Its Realization in Mel-Scale Frequency Cepstral CoefficientsabstractThis brief introduces an innovative recursive discrete cosine transform (DCT) algorithm characterized by its exceptional precision and minimal multiplication requirements. Through the strategic implementation of data reordering and “q” value adjustment schemes, the proposed algorithm entails only a single constant-multiplication operation featuring a fixed cosine coefficient within the iterative phase. By judiciously selecting an appropriate “q” value (q =41), it achieves outstanding results, reaching peak signal-to-noise ratios (PSNRs) of 94.9 and 100.9 dB under 18-bit and 20-bit word length (WL) conditions, respectively, in terms of decimal places. Notably, the proposed algorithm substantially diminishes the number of multiplications by 86.08%, offset by an increase of 2688 additions. The proposed design has a simpler structure and utilizes fewer hardware resources. In field programmable gate array (FPGA) implementation, the device is composed of 43 combinational adaptive look-up tables (ALUTs) specifically allocated for constant multiplication (CM). Overall, the proposed accelerator totally takes 158 combinational ALUTs, 65 registers, a 960-bit read-only memory (ROM), and a 1024-bit random access memory (RAM) in hardware realization and can be operated at a maximum frequency of 156.62 MHz. Therefore, it is particularly well-suited for VLSI implementation in a parallel calculation of Mel-scale frequency cepstral coefficients (MFCCs). Shin-Chi Lai, Szu-Ting Wang, Yi-Chang Zhu, Ying-Hsiu Hung, Jeng-Dao Lee, Wei-Da Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | Fast algorithm and common structure design of recursive analysis and synthesis quadrature mirror filterbanks for digital radio mondialeabstractThis paper proposed a novel fast algorithm and common structure design of analysis and synthesis quadrature mirror filterbanks (AQMF, SQMF) on the spectral band replication (SBR) in digital radio mondiale (DRM). Based on recent Lai et al.'s concept, an extended issue is addressed from the view point of recursively computing the AQMF and SQMF coefficients. The proposed method also combines with the lifting scheme algorithm and canonical signed digit (CSD) multiplication. The results show that the proposed AQMF algorithm has a great improvement on computational complexity. For the recursive kernel computation (N=64), the proposed method has, respectively, 46.38% of multiplication reductions and 20.46% of addition reductions which can cover the shortcoming of the proposed SQMF. The overall complexity of the proposed algorithm (N=64) requires 1984 real multiplication and 128 CSD multiplication, 4704 real addition and 192 CSD addition, and 113 coefficients. It would be more efficient and more suitable than previous works for DRM applications. An-Kai Li, Sheau-Fang Lei, Wen-Kai Tsai, Shin-Chi Lai |
ISCAS | 4 |
| 2013 | An efficient DCT-IV-based ECG compression algorithm and its hardware accelerator designabstractThis paper presents a new efficient DCT-IV-based ECG compression algorithm with a higher Quality Score (QS) and a better Compressing Ratio (CR). The ECG signals sourced from MIT-BIT arrhythmia database with a sampling rate of 360 Hz are employed to be the test patterns for the evaluating the proposed compression algorithm. The simulation results show that the averages of CR, Percent RMS Difference (PRD), and QS are, respectively, 5.267, 0.187, and 28.223 for all 48 lead-V1 patterns of MIT-BIH database. Compared with Lee et al.'s algorithm, the QS value of the proposed method has a great improvement by 25.1%. Additionally, we use DCT-IV to be a unified transform kernel for ECG signal encoding and decoding because the formula of forward DCT-IV is same to its inverse. Also, we realize it to be a compact hardware accelerator with a fewer hardware resources. Therefore, it would be a better choice for realizing the ECG compressor in the future. Shin-Chi Lai, Wei-Che Chien, Chien-Sheng Lan, Meng-Kun Lee, Ching-Hsing Luo, Sheau-Fang Lei |
ISCAS | 1 |
| 2013 | High-performance RDFT design for applications of digital radio mondialeabstractThis paper presents a novel recursive discrete Fourier transform (RDFT) integrated with prime factor and common factor decomposition algorithms. For audio decoder and orthogonal frequency-division multiplexing (OFDM) in a DRM receiver, the proposed RDFT processor is applied to compute both DFT and IMDCT coefficients. Hence, it can greatly reduce the hardware costs in implementation of a portable DRM receiver. For computations of 256-points DFT, the proposed method can greatly reduce 82.75% of multiplications, 85.55% of additions, 79.66% of computational cycles compared with the latest Lai et al.'s RDFT. We further realize it by using TMSC 0.18um 1P6M CMOS technology. The core size is 840×840 um2and the power consumption is 14.6mW @ 25MHz. The results show that it would be more efficient and more suitable than previous works for DRM applications. Shin-Chi Lai, Wen-Ho Juang, Yueh-Shu Lee, Sheau-Fang Lei |
ISCAS | 1 |
| 2012 | Hardware-efficient filterbank design for fast recursive MDST and IMDST algorithmsabstractThis paper proposes a recursive-structure and hardware-efficient filterbank design for modified discrete sine transform (MDST) and inverse MDST (IMDST) computations. The proposed algorithm can be derived to obtain a common computation, the type-IV of Discrete Sine Transform (DST-IV), and then can be converted into a recursive structure of the modified type-II of inverse discrete sine transform (Modified IDST-II). The implementation results show that the proposed method not only can directly employ the same hardware accelerator to calculate the forward and inverse MDST but also greatly reduce the hardware costs. Compared with previous approaches from the simulation results, the Peak signal-to-noise ratio (PSNR) value can achieve 79.76 dB at best, and the time cost per transform of the proposed hardware only takes 343 μs. Shin-Chi Lai, Yi-Ping Yeh, Sheau-Fang Lei |
ISCAS | 1 |