EDBT 2026 Demo / reviewers in the wild / expert
Sangwoo Park 0005
dblp:09/1315-5
· DBLP profile ↗
5ranked-venue papers
1as first author
3since 2021 · last 2026
0000-0002-5831-2176ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Performance Characterization and Optimization of LLM Inference on Tenstorrent AI Accelerators
Jangho Lim, Dongin Shin, Uichan Kim, Jinhyeok Choi, Sangwon Shin, Sangwoo Park 0005, Gunjae Koo, Taeweon Suh |
Euro-Par (2) | 7 |
| 2026 | Three Birds, One Stone: Fast, Accurate-aware and Cost-Efficient Accelerator for Ternary LLMabstractOn-device LLM inference is increasingly important for latency- and privacy-sensitive applications, yet it remains challenging due to the high compute and storage demands. Ternary-weight LLMs are a promising direction because they dramatically reduce model size and simplify arithmetic. In practice, deploying pretrained models on edge devices typically relies on post-training quantization (PTQ), but ternary PTQ often needs fine-grained scaling to preserve accuracy, which amplifies scale-metadata traffic and sub-byte decoding overhead that fits poorly with conventional NPU datapaths. This paper presents T-ACE, a Ternary Accuracy-aware Compute Engine that enables efficient ternary LLM inference under PTQ by jointly designing the data representation and execution pipeline. T-ACE co-packs 64 ternary weights and power-of-two scale metadata into a naturally aligned 16-byte block, eliminating separate scale fetches and preserving aligned memory access. To decode compact ternary packing efficiently, T-ACE proposes a compact two-stage 5-trit unpacker and integrates on-the-fly decoding and scaling directly into the ternary GEMM pipeline. The evaluation on an FPGA prototype shows that decoding and scaling are fully overlapped with GEMM execution, incurring no additional cycles over baseline. Moreover, the comparison against A100/H100 baselines in a normalized setting shows that T-ACE improves accuracy-adjusted compute density (ACD) by 66.8% and accuracy-adjusted energy efficiency (AEE) by 17.6% over the best GPU baseline. Wonseok Jung, Sangwon Shin, Hongjun Um, Jangho Lim, Yongjun Park 0001, Gunjae Koo, Sangwoo Park 0005, Taeweon Suh |
ICS | 8 |
| 2024 | I2SR: Immediate Interrupt Service Routine on RISC-V MCU to Control mmWave RF TransceiversabstractIn 5G technology, millimeter Wave (mmWave), known as Frequency Range 2 (FR2), is one of the candidates for next-generation mobile communications. However, the adoption of mmWave technology in transceiver systems faces challenges in terms of strict latency requirements, due to the significantly increased number of control and compensation tasks. In this paper, we propose a novel interrupt architecture, Immediate Interrupt Service Routine (I2SR) on a RISC-V MCU, enabling zero-latency context switching. Applied to the mmWave transceiver, our I2SR MCU efficiently handles the mmWave tasks by removing context switching overhead. In our experiment, I2SR architecture reduces the overall MCU utilization by 30.6% compared to the baseline MCU, thereby demonstrating that I2SR architecture satisfies the strict latency requirements essential for mmWave applications. Sangwoo Park 0005, Jun-Ho Huh, Sanghyo Jeong, Inhwan Kim, Jae Min Kim |
DATE | 2 |
| 2019 | Driver Drowsiness Detection Using Condition-Adaptive Representation Learning FrameworkabstractWe propose a condition-adaptive representation learning framework for driver drowsiness detection based on a 3D-deep convolutional neural network. The proposed framework consists of four models: spatio-temporal representation learning, scene condition understanding, feature fusion, and drowsiness detection. Spatio-temporal representation learning extracts features that can describe motions and appearances in video simultaneously. Scene condition understanding classifies the scene conditions related to various conditions about the drivers and driving situations, such as statuses of wearing glasses, illumination condition of driving, and motion of facial elements, such as head, eye, and mouth. Feature fusion generates a condition-adaptive representation using two features extracted from the above models. The drowsiness detection model recognizes driver drowsiness status using the condition-adaptive representation. The condition-adaptive representation learning framework can extract more discriminative features focusing on each scene condition than the general representation so that the drowsiness detection method can provide more accurate results for the various driving situations. The proposed framework is evaluated with the NTHU drowsy driver detection video dataset. The experimental results show that our framework outperforms the existing drowsiness detection methods based on visual analysis. Jongmin Yu, Sangwoo Park 0005, Moongu Jeon |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2017 | Learning Feature Representation for Face VerificationabstractPrevious models based on Deep Convolutional Neural Networks (DCNN) for face verification focused on learning face representations. The face features extracted from the models are applied to additional metric learning to improve a verification accuracy. The models extract high-dimensional face features to solve a multi-class classification. This results in a dependency of a model on specific training sets since a dimension of the feature should be equal to the number of subjects in a training set. In this paper, we propose a method for learning feature representations which directly determine whether two input images are identical using a single model based on DCNN and residual learning. It is possible to remove the dependency since the model doesn't learn face representations based on multi-class classification. We show that the proposed method achieves the competitive performance for face verification. We demonstrate the face verification performance of the proposed method using the test dataset of Labeled Face in the Wild dataset. Sangwoo Park 0005, Jongmin Yu, Moongu Jeon |
AVSS | 1 |