VLDB 2026 Research / reviewers in the wild / expert
Dongyun Kam
dblp:234/4038
· DBLP profile ↗
14ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0002-8542-1845ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Panacea: Novel DNN Accelerator using Accuracy-Preserving Asymmetric Quantization and Energy-Saving Bit-Slice SparsityabstractLow bit-precisions and their bit-slice sparsity have recently been studied to accelerate general matrix-multiplications (GEMM) during large-scale deep neural network (DNN) inferences. While the conventional symmetric quantization facilitates low-resolution processing with bit-slice sparsity for both weight and activation, its accuracy loss caused by the activation’s asymmetric distributions cannot be acceptable, especially for largescale DNNs. In efforts to mitigate this accuracy loss, recent studies have actively utilized asymmetric quantization for activations without requiring additional operations. However, the cuttingedge asymmetric quantization produces numerous nonzero slices that cannot be compressed and skipped by recent bit-slice GEMM accelerators, naturally consuming more processing energy to handle the quantized DNN models.To simultaneously achieve high accuracy and hardware efficiency for large-scale DNN inferences, this paper proposes an Asymmetrically-Quantized bit-Slice GEMM (AQS-GEMM) for the first time. In contrast to the previous bit-slice computing, which only skips operations of zero slices, the AQS-GEMM compresses frequent nonzero slices, generated by asymmetric quantization, and skips their operations. To increase the slicelevel sparsity of activations, we also introduce two algorithm-hardware co-optimization methods: a zero-point manipulation and a distribution-based bit-slicing. To support the proposed AQS-GEMM and optimizations at the hardware-level, we newly introduce a DNN accelerator, Panacea, which efficiently handles sparse/dense workloads of the tiled AQS-GEMM to increase data reuse and utilization. Panacea supports a specialized dataflow and run-length encoding to maximize data reuse and minimize external memory accesses, significantly improving its hardware efficiency. Numerous benchmark evaluations show that Panacea outperforms existing DNN accelerators, e.g., $1.97 \times$ and $3.26 \times$ higher energy efficiency, and $1.88 \times$ and $2.41 \times$ higher throughput than the recent bit-slice accelerator Sibia and the SIMD design, respectively, on OPT-2.7B, while providing better algorithm performance with asymmetric quantization. Dongyun Kam, Myeongji Yun, Sunwoo Yoo, Seungwoo Hong, Zhengya Zhang, Youngjoo Lee 0002 |
HPCA | 1 |
| 2025 | Hybrid Ordered Statistics Decoding of Short-Length BCH Codes for URLLC Systems: Theoretical Analysis and Decoder ImplementationabstractThe ordered statistics decoding (OSD) algorithm has been gaining popularity for ultra-reliable and low-latency communication (URLLC) scenarios due to its near-maximum likelihood decoding performance, especially for short linear block codes. However, its substantial computational complexity hinders practical applications. In this paper, we introduce an advanced hybrid OSD algorithm that fully utilizes the hard-decision algebraic decoding results to selectively activate soft-decision OSD operations, significantly mitigating computational complexity. Through a rigorous analysis of error-correction characteristics, we derive a theoretical condition under which the hybrid OSD algorithm guarantees superior error-correction performance over the baseline OSD. To apply the proposed hybrid algorithm to the emerging URLLC systems, we also present a novel decoder architecture that efficiently integrates hard-and soft-decision operations. For (127, 64) BCH codes, the prototype decoder in a 28-nm process achieves an average processing latency of 773 ns at a target block error rate of 10-5, improving information throughput by 4.2× and energy-efficiency by 35× and offering coding gain compared to previous OSD hardware designs. Jaehee Kim, Sangbu Yun, Dongyun Kam, Soonhyun Kwon, Yongjune Kim 0001, Youngjoo Lee 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | A Lightweight ML-Based ECG Classification System Using Self-Personalized Anomaly DetectorabstractTargeting the real-time arrhythmia diagnosis on resource-limited edge devices, in this paper, we present a lightweight electrocardiogram classification system using event-driven machine learning processing. A self-personalized anomaly detector based on signal processing is newly developed to dynamically update internal decision criteria from each patient's recent electrocardiogram history, that activates the following machine learning model only for the abnormal cases. A Siamese neural network is adopted to identify detailed arrhythmia classes by comparing features from the self-personalized normal data and the current abnormal input, increasing the classification accuracy. We also develop a simple version of our Siamese model to reduce the number of trainable parameters while preserving the end-to-end classification accuracy. Experimental results show that the proposed event-driven system reduces ML model activations by 74% for normal beats, achieving a classification accuracy of 96.9% comparable to leading solutions. Additionally, it consumes three times less energy and achieves 3.6 times faster processing latency compared to cost-aware method on a mobile GPU platform, enabling extended battery life and real-time analysis on edge devices. Sunwoo Yoo, Seungwoo Hong, Dongyun Kam, Youngjoo Lee 0002 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Constrained Sorter Design using Zero-One PrincipleabstractTo derive efficient sorting architectures constrained to application-specific input/output conditions, we present in this paper a systematic design methodology that can effectively prune dispensable compare-and-swap (CAS) units. Unlike the previous works resorting to heuristic approaches, the proposed framework exploits the zero-one principle to validate the pruning of a CAS unit at a time, generating the cost-optimized sorter architecture in an iterative manner with a reasonable complexity. In addition to the given input/output constraints, we newly develop the architecture options for the proposed framework, allowing more design spaces for finding the most attractive constrained-sorter design. For 8-list polar decoders, the proposed framework successfully reduces 70% of CAS units in the baseline full sorter, relaxing the area-time complexity by 35% compared with the state-of-the-art solutions. Sangil Han, Jaehee Kim, Dongyun Kam, Byeong Yong Kong, Mijung Kim, Young-Seok Kim, Youngjoo Lee 0002 |
ISCAS | 3 |
| 2024 | A Design Framework for Cost-Efficient Sorters With Arbitrary Input/Output ConstraintsabstractThe sorting operation plays a vital role in various signal processing applications. However, due to high hardware complexity resulting from a series of comparisons, designing the cost-efficient sorter is one of the crucial requisites for improving the overall system performance. To obtain the cost-efficient sorting architectures constrained to application-specific input/output conditions, this paper presents a systematic design methodology that effectively eliminates dispensable compare-and-swap (CAS) units. Unlike the previous heuristic approaches, the proposed framework iteratively prunes a CAS unit followed by the validation step. The zero-one principle is newly applied to reduce the validation time for the practical convergence time with massive searching iterations. To expand the search space of the proposed framework, furthermore, we introduce new architectural options and pruning methods, allowing the cost-efficient design results even compared to the state-of-the-art solutions. Targeting the constrained sorters for communication systems, numerous case studies show that the proposed framework successfully removes more than half of CAS units in the baseline sorter design, significantly relaxing the area-time complexity, e.g., 35% reduction compared to the state-of-the-art architecture for 16-input metric sorter in the SCL polar decoder. Jaehee Kim, Sangil Han, Dongyun Kam, Byeong Yong Kong, Youngjoo Lee 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | GROW: A Row-Stationary Sparse-Dense GEMM Accelerator for Memory-Efficient Graph Convolutional Neural NetworksabstractGraph convolutional neural networks (GCNs) have emerged as a key technology in various application domains where the input data is relational. A unique property of GCNs is that its two primary execution stages, aggregation and combination, exhibit drastically different dataflows. Consequently, prior GCN accelerators tackle this research space by casting the aggregation and combination stages as a series of sparse-dense matrix multiplication. However, prior work frequently suffers from inefficient data movements, leaving significant performance left on the table. We present GROW, a GCN accelerator based on Gustavson’s algorithm to architect a row-wise product based sparse-dense GEMM accelerator. GROW co-designs the software/ hardware that strikes a balance in locality and parallelism for GCNs, reducing the average memory traffic by 2×, and achieving an average 2.8× and 2.3× improvement in performance and energy-efficiency, respectively. Ranggi Hwang, Minhoo Kang, Dongyun Kam, Youngjoo Lee 0002, Minsoo Rhu |
HPCA | 4 |
| 2023 | Energy-Efficient RISC-V-Based Vector Processor for Cache-Aware Structurally-Pruned TransformersabstractBased on recent RISC-V designs, we present in this paper a low-power vector processor architecture for efficiently deploying vision transformer (ViT) models. To fairly measure the processing efficiency of different processor designs with instruction/data cache memories, we first develop the evaluation framework based on numerous design tools for jointly considering the algorithm, architecture, and circuit performances together, numerically revealing that the previous CSR-based data compression cannot accelerate pruned transformer models at all due to under-utilization of the vector-extended processing units. We then introduce a series of algorithm-hardware co-optimization approaches to greatly minimize cache misses by applying 1) the accuracy-preserved structured ViT pruning, 2) the vertical-CSR (vCSR) data storing format, and 3) vCSR-aware custom memory-accessing instructions. Experimental results show that the proposed optimization schemes eventually improve the processing efficiency of pruned transformers in resource-limited computing platforms, e.g., achieving 11 times lower energy consumption for handling the 0.7-pruned ViT model. Jung Gyu Min, Dongyun Kam, Younghoon Byun, Gunho Park, Youngjoo Lee 0002 |
ISLPED | 2 |
| 2023 | Low-Latency SCL Polar Decoder Architecture Using Overlapped Pruning OperationsabstractAllowing the superior error-correction performance even for short-length codewords, the successive-cancellation list (SCL) decoding algorithm has allowed the polar code to be adopted in 5G New Radio standard for control channel. However, existing SCL polar decoders still suffer from long processing latency caused by a number of serialized internal operations. In this work, to solve the latency problem, we present several parallel computing solutions for the serialized operations, i.e., simplified data dependencies and two overlapped pruning operations. To realize the proposed parallel computing, we also introduce internal circuit blocks including dual read-port buffers, an on-the-fly parity checker, and overlapped processing units. The proposed SCL polar decoders are precisely designed with optimal design parameters by analyzing trade-offs between the latency reduction and area overheads. Implemented in a 65-nm CMOS technology, the proposed list-8 SCL polar decoder requires only 374 ns to handle a (1024, 512) 5G codeword, improving the decoding efficiency by 34.7% compared to the previous designs. Dongyun Kam, Byeong Yong Kong, Youngjoo Lee 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | Design and Evaluation Frameworks for Advanced RISC-based Ternary ProcessorabstractIn this paper, we introduce the design and veri-fication frameworks for developing a fully-functional emerging ternary processor. Based on the existing compiling environments for binary processors, for the given ternary instructions, the software-level framework provides an efficient way to convert the given programs to the ternary assembly codes. We also present a hardware-level framework to rapidly evaluate the performance of a ternary processor implemented in arbitrary design technology. As a case study, the fully-functional 9-trit advanced RISC-based ternary (ART-9) core is newly developed by using the proposed frameworks. Utilizing 24 custom ternary instructions, the 5-stage ART-9 prototype architecture is successfully verified by a number of test programs including dhrystone benchmark in a ternary domain, achieving the processing efficiency of 57.8 DMIPS/W and$3.06\times 10^{6}$DMIPS/W in the FPGA-level ternary-logic emulations and the emerging CNTFET ternary gates, respectively. Dongyun Kam, Jung Gyu Min, Jongho Yoon 0001, Sunmean Kim, Seokhyeong Kang, Youngjoo Lee 0002 |
DATE | 1 |
| 2022 | Low-Complexity and Low-Latency SVC Decoding Architecture Using Modified MAP-SP AlgorithmabstractThe compressive sensing (CS) based sparse vector coding (SVC) method is one of the promising ways for the next-generation ultra-reliable and low-latency communications. In this paper, we present advanced algorithm-hardware co-optimization schemes for realizing a cost-effective SVC decoding architecture. The previous maximum a posteriori subspace pursuit (MAP-SP) algorithm is newly modified to relax the computational overheads by applying novel residual forwarding and LLR approximation schemes. A fully-pipelined parallel hardware is also developed to support the modified decoding algorithm, reducing the overall processing latency, especially at the support identification step. In addition, an advanced least-square-problem solver is presented by utilizing the parallel Cholesky decomposer design, further reducing the decoding latency with parallel updates of support values. The implementation results from a 22nm FinFET technology showed that the fully-optimized design is 9.6 times faster while improving the area efficiency by 12 times compared to the baseline realization. Seungwoo Hong, Dongyun Kam, Sangbu Yun, Jeongwon Choe, Namyoon Lee, Youngjoo Lee 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | Low-Latency Polar Decoder Using Overlapped SCL ProcessingabstractIn this paper, we present a novel scheduling method that reduces the latency of polar decoders significantly. Unlike the prior pruning-based successive cancellation list (SCL) decoding that suffers from a number of idle cycles, the proposed overlapped SCL scheme immediately begins node operations without waiting for the list to be sorted, being exempt from such unfavorable cycles. All possible candidates for the next node operations are precomputed in parallel with the pruning operations, and are readily selected to minimize the latency. For the 5G New Radio systems, the proposed method shortens the decoding latency of the state-of-the-art approaches by up to 22% without degrading the error-correcting performance. Dongyun Kam, Byeong Yong Kong, Youngjoo Lee 0002 |
ICASSP | 1 |
| 2021 | Ultralow-Latency Successive Cancellation Polar Decoding Architecture Using Tree-Level ParallelismabstractAchieving the attractive error-correcting capability with a simple decoder structure, the polar code using successive cancellation (SC) decoding is now expected to be installed at the resource-limited IoT or embedded communications. However, the existing SC decoders normally suffer from the long processing latency caused by the serialized processing steps, limiting the practical applications of polar codes. In this article, to solve this latency problem, we present a new low-complexity merging operation that can increase the number of parallel factors for realizing the tree-level parallelism. We also modify the previous pruning method to further reduce the number of visited nodes at the parallel SC decoding scenario. In addition, a novel parallel partial-sum calculator (PSC) architecture is introduced to update partial-sum registers with multiple decoded bits by taking only one processing cycle. Implementation results show that the proposed 8-parallel SC polar decoder in 28-nm CMOS requires only 0.140$\mu \text{s}$to decode a (1024, 512) codeword of 5G system, remarkably reducing the decoding latency when compared to the state-of-the-art designs. Dongyun Kam, Hoyoung Yoo, Youngjoo Lee 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2020 | Ultra-Low-Latency LDPC Decoding Architecture using Reweighted Offset Min-Sum AlgorithmabstractDue to an iterative nature, a low-density parity-check (LDPC) decoder is associated with a long latency, being a major bottleneck of the baseband processor in wireless communication systems. Based on the practical min-sum (MS) decoding method, in this paper, we present a cost-effective algorithm for reducing the processing latency of LDPC decoders. By checking the number of short-length cycles in the LDPC code structure, the proposed method dynamically changes the reweighting factor at the iterative operations, successfully reducing the average number of iterations. In addition, we present several optimization schemes to mitigate the hardware overheads resulting from the proposed reweighting scheme. In a 65-nm CMOS process, a prototype IEEE 802.11ay LDPC decoder optimized by the proposed schemes reduces the decoding latency by 1.7 times with negligible overheads compared with the contemporary designs. Sangbu Yun, Dongyun Kam, Jeongwon Choe, Byeong Yong Kong, Youngjoo Lee 0002 |
ISCAS | 2 |
| 2019 | Ultra-Low-Latency Parallel SC Polar Decoding Architecture for 5G Wireless CommunicationsabstractIn this paper, we newly present a novel parallel polar decoding architecture that significantly reduces the processing latency for 5G wireless communications. Based on the original decoding tree, the proposed scheme first constructs the small trees that generate multiple soft-decision messages in parallel, potentially reducing the decoding latency compared to the previous serialized schemes. The hard-decision estimates are then calculated at the following merging step to decide the decoded outputs and to update the parallel trees. For each parallel tree, the parallel pruning scheme is newly utilized to further optimize the processing latency. In addition, we introduce an efficient parallel decoder architecture, successfully supporting the proposed low-latency algorithm. Implementation results show that the proposed 8-parallel polar decoder in 65nm CMOS uses only 267ns to decode a (1024, 512) polar codeword of 5G system, which is 1.67 times faster than the state-of-the-art design. Dongyun Kam, Youngjoo Lee 0002 |
ISCAS | 1 |