Yi-Chung Wu

dblp:189/9809 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0002-8774-623XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Reconfigurable Computing Challenge: A High-Throughput FPGA Accelerator for Real-Time 3D Gaussian Splatting
abstract
This paper presents the first FPGA accelerator for real-time sort-free 3D Gaussian Splatting. Algorithm–architecture co-optimization is applied by eliminating global depth sorting and adopting a weighted-sum accumulation scheme. The proposed design enables a fully pipelined and hardware-friendly rendering architecture with high parallelism. In addition, a Gaussian pruning strategy reduces the number of Gaussians by 64%–80% with quality degradation less than 1.35 dB. The Width Predictor is designed to eliminates 21.5% redundant pixel computations, while the Diagonal Rendering Mode further reduces latency and achieves 2× area efficiency. A scale-adaptive covariance computation unit ensures stable FP16 computation by mitigating numerical overflow. This work is implemented on a AMD-Xilinx Alveo U200 FPGA, achieving 637.6 FPS at at 600×340 resolution while consuming 2.7 W at a clock frequency of 60 MHz. Compared to prior FPGA-based rendering accelerators, this work achieves 7.64×–12.76× higher energy efficiency, demonstrating an efficient and scalable solution for real-time 3DGS rendering.
Guan-Wei Lai, Yi-Chen Hu, Chia-Yu Kuo, Wei-Chien Cheng, Yi-Chung Wu
FCCM5
2026 A 16-nm 24.8 Mbps/mJ Utilization-Aware Dynamic Allocation Read Mapping Accelerator for Next-Generation Sequencing
Chia-Yu Kuo, Wei-Chien Cheng, Yi-Chung Wu
ISCAS3
2025 A 40-nm 6.75 TOPS/W Quantization-Aware NN Processor for High-Energy-Efficiency Image Recognition
abstract
This paper presents the first VGG-like dedicated neural network (NN) processor designed to support the complete quantization-aware scheme for computer vision acceleration, including input normalization, reparameterization, and convolutional neural networks. The design utilizes a float-to-fixed point transformation scheme and architectural explorations, such as folding and data interleaving. To enhance the throughput, the processor features 9 CNN engines with 1,296 INT8 MACs for compute-intensive operations. A reparameter circuit merges multiple kernels into one, significantly reducing computational complexity. The chip is designed using TSMC 40 nm technology and occupies a core area of 1.14 mm2. Operating at a clock of 200 MHz and consuming 38.4 mW, the chip achieves a 59.6 FPS for real-time image recognition applications. Compared to state-of-the-art designs, it achieves 9.5-to-475.4× and 137.6-to-712.9× improvements in area and energy efficiency, respectively. This work offers substantial benefits for real-time image recognition, providing both high throughput and low power consumption.
Jun-Wei Yu, Jen-Chieh Cheng, Jia-You Yu, Shih-Yu Pan, Yi-Chung Wu
ISCAS5
2022 Achieving Accurate Automatic Sleep Apnea/Hypopnea Syndrome Assessment Using Nasal Pressure Signal
abstract
Automatic assessment of sleep apnea/ hypopnea syndrome (SAHS) based on fewer physiological signals is critical for the success of healthcare at home. However, previous studies that use such settings only achieve a lower assessment accuracy, causing fewer syndromes to be separated for effective diagnosis. This paper presents a 3-stage support vector machines (SVM)-based algorithm for SAHS assessment using a single-channel nasal pressure (NP) signal. In this work, NP signal is utilized for feature extraction. Amplitude features, as well as those extracted using discrete Fourier transform and discrete wavelet transform, are used for machine learning. A total of 58 sets of polysomnography recordings, each with approximately 7 h in duration, were analyzed. This work achieves a sensitivity of 95.7% and a positive predictive value of 90.9%, outperforming previous works using NP signal. Compared with prior studies using only SpO2 signal, this work still achieves better performance and supports more classification levels. Thanks to the low-complexity settings based only on the NP signal, the proposed approach provides a promising solution to SAHS assessment for remote healthcare.
Ying-Sheng Lin, Yi-Pao Wu, Yi-Chung Wu, Pei-Lin Lee, Chia-Hsiang Yang
IEEE J. Biomed. Health Informatics3
2017 Integration of energy-recycling logic and wireless power transfer for ultra-low-power implantables
abstract
This paper presents an integration of energy-recycling logic circuits with a wireless power transfer receiving module for ultra-low-power applications, such as transcutaneous biomedical implantables. In the prototype design, one inductive coil implanted inside the body receives wireless power and supplies the following electronics. While part of the loading is composed of conventional CMOS logics, the rest is implemented with energy-recycling logic circuits. Energy-recycling logic and the associated adiabatic operation achieve excellent energy efficiency by transferring and recycling energy between digital logic blocks along with the signal propagation. The required AC supplies further lead to a natural integration with wireless power transfer and therefore obviate the need for a rectifier that contributes to substantial power loss. As a proof of concept, a finite-impulse-response filter is designed in 90-nm CMOS process. Simulation results show a 59.3% power reduction as compared to static CMOS counterpart.
Hsin-Tzu Lin, Yi-Chung Wu, Ping-Hsuan Hsieh, Chia-Hsiang Yang
ISCAS2
2016 sBWT: memory efficient implementation of the hardware-acceleration-friendly Schindler transform for the fast biological sequence mapping
abstract
MOTIVATION: The Full-text index in Minute space (FM-index) derived from the Burrows-Wheeler transform (BWT) is broadly used for fast string matching in large genomes or a huge set of sequencing reads. Several graphic processing unit (GPU) accelerated aligners based on the FM-index have been proposed recently; however, the construction of the index is still handled by central processing unit (CPU), only parallelized in data level (e.g. by performing blockwise suffix sorting in GPU), or not scalable for large genomes. RESULTS: To fulfill the need for a more practical, hardware-parallelizable indexing and matching approach, we herein propose sBWT based on a BWT variant (i.e. Schindler transform) that can be built with highly simplified hardware-acceleration-friendly algorithms and still suffices accurate and fast string matching in repetitive references. In our tests, the implementation achieves significant speedups in indexing and searching compared with other BWT-based tools and can be applied to a variety of domains. AVAILABILITY AND IMPLEMENTATION: sBWT is implemented in C ++ with CPU-only and GPU-accelerated versions. sBWT is open-source software and is available at http://jhhung.github.io/sBWT/Supplementary information: Supplementary data are available at Bioinformatics online. CONTACT: [email protected] or [email protected] (also [email protected]).
Chia-Hua Chang, Min-Te Chou, Yi-Chung Wu, Ting-Wei Hong, Yun-Lung Li, Chia-Hsiang Yang, Jui-Hung Hung
Bioinform.3
2016 Comparative study of singing voice detection methods
Shingchern D. You, Yi-Chung Wu, Shih-Hsien Peng
Multim. Tools Appl.2