EDBT 2026 Demo / reviewers in the wild / expert
Jongseok Woo
dblp:364/2639
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2024
0000-0001-8030-2356ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Hardware Accelerated Autoencoder for RF Communication Using Short-Time-Fourier- Transform Assisted Convolutional Neural NetworkabstractThis paper presents a hardware-accelerated autoencoder (AE) for wireless communication using a Short-Time-Fourier-Transform Assisted Convolutional Neural Network (STFT-CNN-AE). The design aims to reduce the autoencoder's resource requirements and power dissipation while maintaining its performance even in low Signal-to-Noise Ratio (SNR) wireless channels. The STFT-CNN-AE was implemented and tested on a Zynq UltraScale+ FPGA platform. Prototype measurements show that the STFT-CNN-AE achieves 3.5 times higher throughput at 2.6 times faster frequency, consumes 59% less power, and requires 76% fewer hardware resources (LUT and DSP) compared to a prior Multi-Layer Perceptron-based AE (MLP-AE). These improvements were achieved while maintaining comparable performance in low SNR (<7.5dB) channels. Kuchul Jung, Jongseok Woo, Saibal Mukhopadhyay |
DATE | 2 |
| 2024 | Enhancing IoT Security with a Hardware Accelerated Machine Learning Model coupling Autoencoder and Long-Short-Term-Memory for Anomaly DetectionabstractThis paper proposes a hardware accelerator for machine learning-based anomaly detection to enhance IoT security. Our model integrates Multilayer Perceptron (MLP) with Long Short-Term Memory (LSTM), utilizing an MLP-based Autoencoder and Isolation Forest algorithm for data dimensionality reduction and computational complexity reduction. Prototyped on a Zynq UltraScale+ XCZU9EG FPGA, our AE-LSTM model surpasses baseline MLP-only and LSTM-only models in resource utilization efficiency and detection accuracy. Compared to these baselines, it reduces parameters by 79.4% and 98% and LUT usage by 61.4% and 90.8%, respectively, while minimizing other resource utilization. Furthermore, power consumption is lowered to about 40% of the MLP-based model's consumption rate and 36% of the LSTM-based model's rate, with latency reduced to less than one-third from both baselines. Kuchul Jung, Jongseok Woo, Saibal Mukhopadhyay |
ISCAS | 2 |
| 2024 | Efficient Hardware Design of DNN for RF Signal Modulation RecognitionabstractThis paper presents an efficient deep neural network (DNN) accelerator design for the application of modulation recognition of Radio Frequency (RF) signals. A low complexity DNN model utilizing the ternary weights is demonstrated with co-analysis of the classification accuracy and the hardware design. In order to maximize the benefits of the ternary weight quantization, the dedicated hardware design called merged layer architecture is proposed. Physical design analysis shows that the proposed method can improve the bandwidth of the received signal and reduce hardware costs significantly. The physical design analysis is based on the Application Specific Integrated Circuit (ASIC) to evaluate the dedicated hardware design, and the functionality of the full DNN system is verified on the FPGA platform. Jongseok Woo, Kuchul Jung, Saibal Mukhopadhyay |
ISCAS | 1 |
| 2024 | Hardware-friendly Hessian-driven Row-wise Quantization and FPGA Acceleration for Transformer-based ModelsabstractRecent advancements in using FPGAs as co-processors for language model acceleration, particularly in terms of energy efficiency and flexibility, face challenges due to limited memory capacity. This issue hinders the deployment of transformer-based language models. To address these issues, we propose a novel software-hardware co-optimization approach. Our approach incorporates a hardware-friendly hessian-based row-wise mixed-precision quantization algorithm and an intra-layer mixed-precision compute fabric. The software algorithm, based on Hessian analysis, quantizes important rows with high precision and unimportant rows with low precision to compress the parameters effectively while maintaining accuracy, enabling fine-grained mixed-precision computation on the FPGA accelerator. Moreover, the integration of row-wise mixed-precision quantization and our energy-efficient data flow optimization scheme enables the accommodation of all necessary parameters of the BERT-base model on the FPGA, eliminating the need for off-chip memory access during runtime. The experimental results demonstrate that our FPGA accelerator generally outperforms existing FPGA accelerators, exhibiting energy efficiency improvements ranging from 4.12X to 14.57X compared to existing FPGA accelerators. Woohong Byun, Jongseok Woo, Saibal Mukhopadhyay |
ISLPED | 2 |