EDBT 2026 Demo / reviewers in the wild / expert
Sangwon Shin
dblp:252/8820
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Performance Characterization and Optimization of LLM Inference on Tenstorrent AI Accelerators
Jangho Lim, Dongin Shin, Uichan Kim, Jinhyeok Choi, Sangwon Shin, Sangwoo Park 0005, Gunjae Koo, Taeweon Suh |
Euro-Par (2) | 6 |
| 2026 | Three Birds, One Stone: Fast, Accurate-aware and Cost-Efficient Accelerator for Ternary LLMabstractOn-device LLM inference is increasingly important for latency- and privacy-sensitive applications, yet it remains challenging due to the high compute and storage demands. Ternary-weight LLMs are a promising direction because they dramatically reduce model size and simplify arithmetic. In practice, deploying pretrained models on edge devices typically relies on post-training quantization (PTQ), but ternary PTQ often needs fine-grained scaling to preserve accuracy, which amplifies scale-metadata traffic and sub-byte decoding overhead that fits poorly with conventional NPU datapaths. This paper presents T-ACE, a Ternary Accuracy-aware Compute Engine that enables efficient ternary LLM inference under PTQ by jointly designing the data representation and execution pipeline. T-ACE co-packs 64 ternary weights and power-of-two scale metadata into a naturally aligned 16-byte block, eliminating separate scale fetches and preserving aligned memory access. To decode compact ternary packing efficiently, T-ACE proposes a compact two-stage 5-trit unpacker and integrates on-the-fly decoding and scaling directly into the ternary GEMM pipeline. The evaluation on an FPGA prototype shows that decoding and scaling are fully overlapped with GEMM execution, incurring no additional cycles over baseline. Moreover, the comparison against A100/H100 baselines in a normalized setting shows that T-ACE improves accuracy-adjusted compute density (ACD) by 66.8% and accuracy-adjusted energy efficiency (AEE) by 17.6% over the best GPU baseline. Wonseok Jung, Sangwon Shin, Hongjun Um, Jangho Lim, Yongjun Park 0001, Gunjae Koo, Sangwoo Park 0005, Taeweon Suh |
ICS | 3 |
| 2026 | SumcheckPIM: An Efficient HBM-Based PIM Architecture for Linear Complexity Zero Knowledge ProofsabstractZero-knowledge proofs (ZKPs) are emerging as a core technology for privacy-preserving computation. Despite steady progress in protocol and algorithm design, generating these proofs remains computationally intensive, driving growing interest in hardware acceleration for kernels such as number-theoretic transform (NTT) and multi-scalar multiplication (MSM). Among them, the sumcheck protocol offers a compelling alternative with O(n) prover complexity compared to O(nlog n) for NTT-based approaches, yet our analysis reveals its execution is fundamentally memory-bound, with severely underutilized compute resources. This characteristic demands a memory-centric acceleration strategy, in contrast to compute-centric approaches of prior work. Sunchae Kim, Taewoon Kang, Sangwon Shin, Taeweon Suh, Yibin Yang 0001, Gunjae Koo |
ICS | 3 |
| 2025 | SkipNZ: Non-zero Value Skipping for Efficient CNN Acceleration
Joonyup Kwon, Jinhyeok Choi, Ngoc-Son Pham, Sangwon Shin, Taeweon Suh |
Euro-Par (2) | 4 |
| 2025 | HBM-Aware Number Theoretic Transform Accelerator for Zero-Knowledge ProofabstractZero-Knowledge Proof (ZKP) cryptographic algorithms have garnered significant attention for their ability to enhance privacy. However, the practical deployment of these algorithms remains challenging because they demand extremely high computational effort and handle huge volumes of data, especially in the Number Theoretic Transform (NTT) step. In this work, we propose an HBM-aware dataflow that employs sub-tiling and row-shuffling techniques to overcome the nonuniform stride access problem and to maximize HBM bandwidth utilization. We also design the NTT accelerator to use minimal FPGA resources. In particular, we explore diverse design options for the 256-bit modular multiplier and adopt an efficient design that optimizes resource usage and performance. Experimental results demonstrate that the proposed accelerator achieves lower latency and enhanced resource utilization compared to state-of-the-art FPGA-based designs. Sangwon Shin, Ngoc-Son Pham, Lei Xu 0012, Larry Shi, Taeweon Suh |
ICCD | 1 |
| 2024 | Demo: Real-Time Spectrum Segmentation and Classification with Over-The-Air DataabstractSpectrum usage is increasing daily, necessitating new methods for efficient utilization. Spectrum sharing allows the coexistence of multiple wireless communication systems in the same spectrum. Effective spectrum segmentation and classification are essential for this, yet existing methods treat them as separate processes and often focus on specific communication techniques. Our application addresses these issues by jointly segmenting and classifying narrowband signals in wideband IQ samples. This demo paper presents the application, demonstrating its end-to-end approach of spectrum segmentation and classification. The application achieves an accuracy of 92.6% on the over-the-air (OTA) wireless communication spectrum. This represents an improvement of over 9% compared to the state-of-the-art solution, highlighting its effectiveness. Sangwon Shin, Prashant Subedi, Mehmet Can Vuran |
LCN | 1 |
| 2024 | Seek and Classify: End-to-end Joint Spectrum Segmentation and Classification for Multi-signal Wideband Spectrum SensingabstractThe rise in the use of wireless communication has led to the problem of spectrum scarcity in licensed bands. The popularity of the Internet of Things (IoT) requires innovative solutions that maximize the use of the available spectrum to support the increasing number of connected devices. The ability to detect and classify modulation of the signals efficiently can enable a cognitive radio to monitor the spectrum activity in real-time and utilize unused frequencies. In this work, Seek and Classify, an end-to-end framework for joint spectrum segmentation and classification for narrowband signals from wideband IQ samples, is developed. Seek and Classify includes a novel intersection of unions-based training methodology and machine learning architectures that advances this unique area. Evaluations performed through both synthetically generated radio signals and over-the-air experiments with software defined radios reveal that the proposed training strategy and models increase the classification accuracy from 41% to 99%. Moreover, the end-to-end framework reduces the sensing time for narrowband signals by 2-10 times, depending on hardware capabilities. The extensive evaluations provide guidance for the choice of training methods, machine learning architectures, and preprocessing tools for the most effective joint segmentation and classification performance. Prashant Subedi, Sangwon Shin, Mehmet Can Vuran |
LCN | 2 |