Xin-Yu Shih

dblp:64/7127 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
3since 2021 · last 2023
0000-0002-9045-5847ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2023 Unified chip hardware architecture of KD-tree mean-based trainer and speeding-up classifier with repeat-point searching for various applications
Xin-Yu Shih, Chen-Yen Song, Yao-Yu Lu
Integr.1
2023 Design and Implementation of Decision-Tree (DT) Online Training Hardware Using Divider-Free GI Calculation and Speeding-Up Double-Root Classifier
abstract
This paper proposes a total solution in ASIC chip hardware implementation, fulfilling decision-tree (DT) online training and classification missions. We also develop three effective design techniques, including divider-free resister-in-parallel Gini Impurity calculation (DR-GIC), adaptive means with updated learning function (AM-ULF), and speeding-up classifier with double-root tree (SC-DRT). In the ASIC chip implementation with TSMC 40-nm multi-Vt CMOS process, the chip layout is well-verified and only has a total core area of 0.803 mm2. The maximum operating frequency is 429 MHz, dissipating average power of 73.7 mW. Regarding online training, it supports a maximum training size of 10.1KB data on chip. Accordingly, the worst-case training latency is only 9.98ms for dealing with 1024 data elements. For classification, the decision throughput is up to 4.29 GBps – 8.58 GBps. After being verified with various datasets for different applications, the ASIC chip with superior design performance offers a solid hardware implementation, realizing a total DT-based procedure from online training to classification.
Xin-Yu Shih, Yao Chiu, Hsiang-En Wu
IEEE Trans. Circuits Syst. I Regul. Pap.1
2023 Design and Implementation of Dual-Mode Support Vector Machine (SVM) Trainer and Classifier Chip Architecture for Human Disease Detection Applications
abstract
This paper proposes a low-latency dual-mode support vector machine (SVM) chip hardware architecture for human disease detection applications. It can simultaneously support the on-line SVM trainer and classifier both for linear and non-linear operating modes. In addition to the proposed hardware architecture, we also develop three effective design techniques, such as Queue-based Fast Selection (QFS), Dual-mode Linear and Non-linear Kernel (DLNK), and Low-latency Lagrange-multiplier-value Updater (LLU). For the ASIC implementation, our developed work with TSMC 40-nm multi-Vt CMOS technology only occupies a total core area of 0.249 mm2 in chip layout. After system validation with seven representative datasets, the operating frequency of our chip is up to 452 MHz. For linear and non-linear modes, the according power dissipates 54.2 mW and 71.4 mW, respectively. While manipulating 1024 training data for linear and non-linear modes, the worst-case training latency is only 1.8 ms and 99.0 ms, respectively. In addition, under the superior accuracy, true positive rate (TPR), and true negative rate (TNR) performance, the classification throughput for linear and non-linear modes is 2.42 GBps and 4.74 MBps, respectively.
Xin-Yu Shih, Hsiang-En Wu, Ming-Xian Cai
IEEE Trans. Circuits Syst. I Regul. Pap.1
2019 Flexible design and implementation of QC-Based LDPC decoder architecture for on-line user-defined matrix downloading and efficient decoding
Xin-Yu Shih, Hong-Ru Chou
Integr.1
2018 VLSI design and implementation of a reconfigurable hardware-friendly Polar encoder architecture for emerging high-speed 5G system
Xin-Yu Shih, Po-Chun Huang, Hong-Ru Chou
Integr.1
2009 A 52-mW 8.29mm2 19-mode LDPC decoder chip for mobile WiMAX applications
abstract
This paper presents a LDPC decoder chip supporting all 19 modes in Mobile WiMAX applications. An efficient IC design strategy is proposed to reduce 31.25% decoding latency, and enhance hardware utilization ratio from 50% to 75%. In addition, we propose a new early termination scheme that can dynamically adjust the iteration number. The multi-mode chip implemented in 8.29mm2die area can be maximally measured at 83.3MHz with only 52mW power consumption.
Xin-Yu Shih, Cheng-Zhou Zhan, An-Yeu Wu
ASP-DAC1
2009 A Triple-mode LDPC Decoder Design for IEEE 802.11n SYSTEM
abstract
This paper shows a triple-mode LDPC decoder design with two design techniques, the matrix reordering algorithm for multi-mode reconfiguration and the single-entry-multiple-data (SEMD) scheme for throughput enhancement. The matrix reordering algorithm can reduce the computational complexity from O(n!) to O(n3). The SEMD can enhance the throughput by m times with small area overhead. With TSMC 0.13 mum CMOS, the proposed design is synthesized in 1.99 mm2area at 172.4 MHz.
Min-An Chao, Jen-Yang Wen, Xin-Yu Shih, An-Yeu Wu
ISCAS3
2008 High-performance scheduling algorithm for partially parallel LDPC decoder
abstract
In this paper, we propose a new scheduling algorithm for the overlapped message passing decoding, which can be applied to general low-density parity check (LDPC) codes. The partially parallel LDPC architecture is commonly used for reducing the area cost of the processing units. The dependency of two kinds of processing units, check node unit (CNU) and bit node unit (BNU), should be considered to enhance the hardware utilization efficiency (HUE). Based on the properties of the parity check matrix of LDPC codes, the updating calculation of the CNU and BNU can be overlapped to reduce the decoding latency by enhancing the HUE with the matrix scheduling algorithm. By applying our proposed LDPC scheduling algorithm to a (1944, 972)-irregular LDPC code, we can get about 60% throughput gain in average without any performance degradation.
Cheng-Zhou Zhan, Xin-Yu Shih, An-Yeu Wu
ICASSP2