Youngjoo Lee 0002

dblp:24/3466-2 · DBLP profile ↗
← Back
45ranked-venue papers
4as first author
32since 2021 · last 2026
0000-0002-2467-8276ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 38 · 3 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 IterQuant: Iterative Quantization Framework for Mixed-Precision LLM Compression
abstract
Mixed-precision quantization is a promising approach for compressing large language models (LLMs) while maintaining output quality. However, the design space for selecting resolutions of different layers makes exhaustive search intractable. Existing methods either rely on rigid bit-width allocation schemes or require extensive hyperparameter tuning, often based on inaccurate layer-wise sensitivity metrics. In this work, we propose IterQuant, an iterative quantization framework that efficiently explores the mixed-precision space without requiring exhaustive enumeration. By incorporating momentum-based scoring to reflect historical performance trends and parameter grouping to balance quantization granularity, IterQuant achieves favorable trade-offs between compression and accuracy. Unlike prior approaches that assume bit allocation sensitivity from full-precision models directly transfers to quantized models, IterQuant dynamically updates its quantization decisions as the model evolves, better capturing inter-layer dependencies. Experimental results demonstrate that IterQuant significantly outperforms state-of-the-art mixed-precision quantization approaches by 2.8% near 4 bits in preserving token-level output quality across various LLM benchmarks.
Hyungyo Jeong, Hyeokjun Kwon, Jaeho Lee 0001, Youngjoo Lee 0002
DATE5
2026 Low-Latency Software-Defined 5G NR PUSCH Receiver with Mixed-Precision SIMD Acceleration
Jaehee Kim, Jaewha Kim, Nam-Il Kim, Youngjoo Lee 0002
ISCAS5
2026 Memory-Efficient Partially Self-Corrected Min-Sum LDPC Decoder for 5G NR Applications
Sangbu Yun, Jeongwon Choe, Youngjoo Lee 0002
ISCAS4
2026 A 3.3 Gb/s/mm2 Area-Efficient Non-Binary LDPC Decoder Using Column-Layered Processing
abstract
Non-binary low-density parity-check (NB-LDPC) codes are a prominent class of error-correction codes, offering superior error-correcting performance compared to their binary counterparts. However, previous NB-LDPC decoders suffer from high processing complexity and significant memory overhead when supporting high-order Galois fields and long codeword lengths. To address these challenges, the proposed decoder leverages the trellis min-max algorithm and adopts a column-layered decoding schedule with on-the-fly message computation to reduce memory requirements. Additionally, the proposed column-layered algorithm shares up-to-date information among columns, enhancing the convergence speed of the baseline design. Considering the structure of high-rate NB-LDPC codes, we introduce multi-column processing with an optimized banked memory architecture while minimizing parallel processing overhead through submodule optimization. Fabricated using a 28-nm CMOS technology, the prototype 4KB 0.9-rate decoder achieves a 1.42-fold improvement in area efficiency compared to state-of-the-art designs. While the proposed design is motivated by the requirements of storage applications, its modular organization and scalable parallelism also allow adaptation to diverse domains such as wireless and optical communications.
Jeongwon Choe, Youngjoo Lee 0002
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 RISC-V Driven Orchestration of Vector Processing Units and eFlash Compute-in-Memory Arrays for Fast and Accurate Keyword Spotting
abstract
In this paper, we propose a computationally efficient keyword spotting (KWS) model, named hybrid reparameterized FSMN (HRepFSMN), by carefully examining the impact of binarization on the accuracy. In particular, we found that binarizing depthwise convolution (DW-Conv) within the previous binarized KWS model, i.e., BiFSMNv2, does not lead to a significant reduction in FLOPs. Therefore, we allow floating-point (FP) operations on less computation-intensive DW-Conv layers while the remaining layers are computed in a binary fashion (hybrid data type). In addition, we remove skip connections, which require data fetching in full precision, by applying a reparameterization technique. More importantly, to efficiently compute the proposed HRepFSMN, we present a RISC-V controlled hardware accelerator that consists of reconfigurable vector processing units for FP operations and eFlash compute-in-memory arrays for binary operations. We extend RISC-V instructions so that the core can efficiently manage both computing fabrics. As a result, our HRepFSMN improves accuracy by 2.57%/4.98% with 24.02×/3.66× speed-up compared to BiFSMNv2/BiFSMNv2_small. By shrinking down our HRepFSMN, we achieve 0.95% higher accuracy with 20.87× speed-up compared to BiFSMNv2_small.
Gunil Kang, Dahoon Park, Sangwoo Jung 0001, Jung Gyu Min, Youngjoo Lee 0002, Jaeha Kung 0001
ASP-DAC7
2025 Panacea: Novel DNN Accelerator using Accuracy-Preserving Asymmetric Quantization and Energy-Saving Bit-Slice Sparsity
abstract
Low bit-precisions and their bit-slice sparsity have recently been studied to accelerate general matrix-multiplications (GEMM) during large-scale deep neural network (DNN) inferences. While the conventional symmetric quantization facilitates low-resolution processing with bit-slice sparsity for both weight and activation, its accuracy loss caused by the activation’s asymmetric distributions cannot be acceptable, especially for largescale DNNs. In efforts to mitigate this accuracy loss, recent studies have actively utilized asymmetric quantization for activations without requiring additional operations. However, the cuttingedge asymmetric quantization produces numerous nonzero slices that cannot be compressed and skipped by recent bit-slice GEMM accelerators, naturally consuming more processing energy to handle the quantized DNN models.To simultaneously achieve high accuracy and hardware efficiency for large-scale DNN inferences, this paper proposes an Asymmetrically-Quantized bit-Slice GEMM (AQS-GEMM) for the first time. In contrast to the previous bit-slice computing, which only skips operations of zero slices, the AQS-GEMM compresses frequent nonzero slices, generated by asymmetric quantization, and skips their operations. To increase the slicelevel sparsity of activations, we also introduce two algorithm-hardware co-optimization methods: a zero-point manipulation and a distribution-based bit-slicing. To support the proposed AQS-GEMM and optimizations at the hardware-level, we newly introduce a DNN accelerator, Panacea, which efficiently handles sparse/dense workloads of the tiled AQS-GEMM to increase data reuse and utilization. Panacea supports a specialized dataflow and run-length encoding to maximize data reuse and minimize external memory accesses, significantly improving its hardware efficiency. Numerous benchmark evaluations show that Panacea outperforms existing DNN accelerators, e.g., $1.97 \times$ and $3.26 \times$ higher energy efficiency, and $1.88 \times$ and $2.41 \times$ higher throughput than the recent bit-slice accelerator Sibia and the SIMD design, respectively, on OPT-2.7B, while providing better algorithm performance with asymmetric quantization.
Dongyun Kam, Myeongji Yun, Sunwoo Yoo, Seungwoo Hong, Zhengya Zhang, Youngjoo Lee 0002
HPCA6
2025 Cost-efficient Processing-in-Memory Architecture with Training-free and Universal Error Compensation
abstract
Despite the energy efficiency of memory-centric deep neural network (DNN) computations, the nonlinearities inherent in existing processing-in-memory (PIM) architectures cause severe accuracy drops. These imperfections necessitate additional methods to correct inaccurate vector-matrix multiplication (VMM) results. To address this issue without modifying DNN weights, we first propose an input sparsity-based error compensation method. This approach dynamically corrects accumulated errors along the column direction of the non-volatile memory (NVM) array using pre-collected errors and input characteristics. We then present a new PIM architecture along with the proposed compensation scheme by slightly modifying the existing analog-to-digital converter (ADC) or adding a few extra rows to the NVM array. Experimental results show that the proposed work mitigates the nonlinear effects of various emerging memory cells, achieving near-ideal DNN accuracy with negligible hardware overheads.
Myeongji Yun, Jung Gyu Min, Sein Oh, Jiwoung Choi, Jang-Sik Lee, Minkyu Je, Youngjoo Lee 0002
ISLPED7
2025 Hybrid Ordered Statistics Decoding of Short-Length BCH Codes for URLLC Systems: Theoretical Analysis and Decoder Implementation
abstract
The ordered statistics decoding (OSD) algorithm has been gaining popularity for ultra-reliable and low-latency communication (URLLC) scenarios due to its near-maximum likelihood decoding performance, especially for short linear block codes. However, its substantial computational complexity hinders practical applications. In this paper, we introduce an advanced hybrid OSD algorithm that fully utilizes the hard-decision algebraic decoding results to selectively activate soft-decision OSD operations, significantly mitigating computational complexity. Through a rigorous analysis of error-correction characteristics, we derive a theoretical condition under which the hybrid OSD algorithm guarantees superior error-correction performance over the baseline OSD. To apply the proposed hybrid algorithm to the emerging URLLC systems, we also present a novel decoder architecture that efficiently integrates hard-and soft-decision operations. For (127, 64) BCH codes, the prototype decoder in a 28-nm process achieves an average processing latency of 773 ns at a target block error rate of 10-5, improving information throughput by 4.2× and energy-efficiency by 35× and offering coding gain compared to previous OSD hardware designs.
Jaehee Kim, Sangbu Yun, Dongyun Kam, Soonhyun Kwon, Yongjune Kim 0001, Youngjoo Lee 0002
IEEE Trans. Circuits Syst. I Regul. Pap.7
2025 Distinguishing Pathologic Gait in Older Adults Using Instrumented Insoles and Deep Neural Networks
abstract
Gait abnormalities are common in the older population owing to aging- and disease-related changes in physical and neurological functions. Differentiating the causes of gait abnormalities is challenging because various abnormal gaits share a similar pattern in older patients. Herein, we propose a deep neural network (DNN) model to classify disease-specific gait patterns in older adults using commercialized instrumented insoles. This study included 150 patients aged ≥ 65 years, divided into the following five groups (N = 30 in each group): healthy older individuals (HI), patients with Parkinson's disease (PD), patients with spastic hemiplegic gait due to stroke (SH), patients with normal-pressure hydrocephalus (NPH), and patients with knee osteoarthritis (OA). Participants performed the timed up and go test (TUGT) wearing the commercialized instrumented insole, GDCA-MD (Gilon, Republic of Korea). Seven data streams were collected from each insole using a 3-axis accelerometer and four pressure sensors and were analyzed. First, the statistical differences among groups in spatiotemporal features during TUGT, such as step count, step length, velocity, acceleration, regularity, and symmetricity, were examined. Second, a two-stage DNN model was developed that distinguishes HI from others in the first network and classifies the pathologic groups in the second network. The areas under the curve were 0.96, 0.88, 0.98, 0.96, and 0.97 for identifying HI, PD, OA, SH, and NPH, respectively. We demonstrated that the proposed DNN model can reliably classify gait abnormalities in an older population using simple instrumented insoles and a test.
Jin Hyun, Seung-Ick Choi, Sangbu Yun, Kwangho Chung, Seok Jong Chung, Jun Kyu Hwang, Eun Joo Yang, Youngjoo Lee 0002
IEEE J. Biomed. Health Informatics9
2025 mDARTS: Searching ML-Based ECG Classifiers Against Membership Inference Attacks
abstract
This paper addresses the critical need for elctrocardiogram (ECG) classifier architectures that balance high classification performance with robust privacy protection against membership inference attacks (MIA). We introduce a comprehensive approach that innovates in both machine learning efficacy and privacy preservation. Key contributions include the development of a privacy estimator to quantify and mitigate privacy leakage in neural network architectures used for ECG classification. Utilizing this privacy estimator, we propose mDARTS (searching ML-based ECG classifier against MIA), integrating MIA's attack loss into the architecture search process to identify architectures that are both accurate and resilient to MIA threats. Our method achieves significant improvements, with an ECG classification accuracy of 92.1% and a lower privacy score of 54.3%, indicating reduced potential for sensitive information leakage. Heuristic experiments refine architecture search parameters specifically for ECG classification, enhancing classifier performance and privacy scores by up to 3.0% and 1.0%, respectively. The framework's adaptability supports user customization, enabling the extraction of architectures that meet specific criteria such as optimal classification performance with minimal privacy risk. By focusing on the intersection of high-performance ECG classification and the mitigation of privacy risks associated with MIA, our study offers a pioneering solution addressing the limitations of previous approaches.
Eun-Bin Park, Youngjoo Lee 0002
IEEE J. Biomed. Health Informatics2
2025 A Lightweight ML-Based ECG Classification System Using Self-Personalized Anomaly Detector
abstract
Targeting the real-time arrhythmia diagnosis on resource-limited edge devices, in this paper, we present a lightweight electrocardiogram classification system using event-driven machine learning processing. A self-personalized anomaly detector based on signal processing is newly developed to dynamically update internal decision criteria from each patient's recent electrocardiogram history, that activates the following machine learning model only for the abnormal cases. A Siamese neural network is adopted to identify detailed arrhythmia classes by comparing features from the self-personalized normal data and the current abnormal input, increasing the classification accuracy. We also develop a simple version of our Siamese model to reduce the number of trainable parameters while preserving the end-to-end classification accuracy. Experimental results show that the proposed event-driven system reduces ML model activations by 74% for normal beats, achieving a classification accuracy of 96.9% comparable to leading solutions. Additionally, it consumes three times less energy and achieves 3.6 times faster processing latency compared to cost-aware method on a mobile GPU platform, enabling extended battery life and real-time analysis on edge devices.
Sunwoo Yoo, Seungwoo Hong, Dongyun Kam, Youngjoo Lee 0002
IEEE J. Biomed. Health Informatics4
2024 Partially-Structured Transformer Pruning with Patch-Limited XOR-Gate Compression for Stall-Free Sparse-Model Access
abstract
The pruning-based model compression is regarded as an essential technique to deploy the recent large-size transformer models in practical services; however, accessing sparse transformer models cannot reach the ideal speed at all due to the frequent memory stalls for the irregular memory-accessing patterns. Based on the recent XOR-gate compression relaxing the amount of irregular accesses, this work presents a novel partially-structured transformer pruning method dedicated to the interface-friendly compression format. The stall-free memory access is firstly derived by limiting the number of patches per weight, introducing a new trade-off between model quality and effective memory bandwidth. Then, the partially-structured pruning patterns are deployed to provide better accuracy-bandwidth trade-off by significantly reducing the number of correction patches. Adjusting the patch distribution per weight in an aggressive way, the number of limited patches can be even smaller than that of weight bits, further increasing the effective bandwidth for achieving the similar model accuracy. We demonstrate the proposed stall-free XOR-gate compression schemes at pruned DeiT/BERT models on ImageNet/SQuAD datasets, presenting the highest effective bandwidth for accessing sparse transformers compared to the existing stall-based solutions.
Younghoon Byun, Youngjoo Lee 0002
DAC2
2024 Constrained Sorter Design using Zero-One Principle
abstract
To derive efficient sorting architectures constrained to application-specific input/output conditions, we present in this paper a systematic design methodology that can effectively prune dispensable compare-and-swap (CAS) units. Unlike the previous works resorting to heuristic approaches, the proposed framework exploits the zero-one principle to validate the pruning of a CAS unit at a time, generating the cost-optimized sorter architecture in an iterative manner with a reasonable complexity. In addition to the given input/output constraints, we newly develop the architecture options for the proposed framework, allowing more design spaces for finding the most attractive constrained-sorter design. For 8-list polar decoders, the proposed framework successfully reduces 70% of CAS units in the baseline full sorter, relaxing the area-time complexity by 35% compared with the state-of-the-art solutions.
Sangil Han, Jaehee Kim, Dongyun Kam, Byeong Yong Kong, Mijung Kim, Young-Seok Kim, Youngjoo Lee 0002
ISCAS7
2024 A Dual-Precision and Low-Power CNN Inference Engine Using a Heterogeneous Processing-in-Memory Architecture
abstract
In this article, we present an energy-scalable CNN model that can adapt to different hardware resource constraints. Specifically, we propose a dual-precision network, named DualNet, that leverages two independent bit-precision paths (INT4 and ternary-binary). DualNet achieves both high accuracy and low complexity by balancing the ratio between two paths. We also present an evolutionary algorithm that allows the automatic search of the optimal ratios. In addition to the novel CNN architecture design, we develop a heterogeneous processing-in-memory (PIM) hardware that integrates SRAM-and eDRAM-based PIMs to efficiently compute two precision paths in parallel. To verify the energy efficiency of DualNet computed on the heterogeneous PIM, we prototyped a test chip in 28nm CMOS technology. To maximize the hardware efficiency, we utilize an improved data mapping scheme achieving the most effective deployment of DualNets on multiple PIM arrays. With the proposed SW-HW co-optimization, we can obtain the most energy-efficient DualNet model operating on the actual PIM hardware. Compared to the other quantized networks with a single bit-precision, DualNet reduces the energy consumption, memory footprint, and latency by 29.0%, 49.5%, 47.3% on average, respectively, for CIFAR-10/100 and ImageNet datasets.
Sangwoo Jung 0001, Dahoon Park, Youngjoo Lee 0002, Jong-Hyeok Yoon, Jaeha Kung 0001
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 A Design Framework for Cost-Efficient Sorters With Arbitrary Input/Output Constraints
abstract
The sorting operation plays a vital role in various signal processing applications. However, due to high hardware complexity resulting from a series of comparisons, designing the cost-efficient sorter is one of the crucial requisites for improving the overall system performance. To obtain the cost-efficient sorting architectures constrained to application-specific input/output conditions, this paper presents a systematic design methodology that effectively eliminates dispensable compare-and-swap (CAS) units. Unlike the previous heuristic approaches, the proposed framework iteratively prunes a CAS unit followed by the validation step. The zero-one principle is newly applied to reduce the validation time for the practical convergence time with massive searching iterations. To expand the search space of the proposed framework, furthermore, we introduce new architectural options and pruning methods, allowing the cost-efficient design results even compared to the state-of-the-art solutions. Targeting the constrained sorters for communication systems, numerous case studies show that the proposed framework successfully removes more than half of CAS units in the baseline sorter design, significantly relaxing the area-time complexity, e.g., 35% reduction compared to the state-of-the-art architecture for 16-input metric sorter in the SCL polar decoder.
Jaehee Kim, Sangil Han, Dongyun Kam, Byeong Yong Kong, Youngjoo Lee 0002
IEEE Trans. Circuits Syst. I Regul. Pap.5
2024 A 43.9 μs IRS Controller SoC With Grid-Based Phase-Shift Optimization in 28 nm CMOS Technology for Next- Generation Communication
abstract
Intelligent reflecting surface (IRS) is one of the promising technologies for next-generation communication systems. Although IRS can enhance the integrity of the transmitted signal, however, it requires additional phase-shift optimization task, potentially increasing the total processing latency of the system. To implement a phase-shift optimization algorithm with low-latency, for the world-first, IRS controller SoC for saving the base station (BS) transmit (TX) antenna power is presented with novel latency reduction techniques; 1) grid-based phase-shift elements with a coarse-grained update, 2) quality-aware approximate processing element with low-resolution saturation arithmetic, 3) optimized data-flow scheduling with double-buffer architecture, and 4) novel low-latency weight update tracking technique. By adopting the proposed techniques, the fully-optimized SoC architecture can enhance the processing latency and energy consumption by 93% and 68%, respectively, compared with the unoptimized SoC architecture. The proposed system is designed and fabricated in 28 nm CMOS technology, where the implementation results show that the fully-optimized SoC system can achieve 43.9$\mu$s of processing latency while saving 14.5 dB of TX antenna power for optimizing the$16\times 16$grid-based IRS architecture with$128\times 8$MU-MIMO configuration.
Seungsik Moon, Youngjoo Lee 0002
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 Multi-Group Multicasting Systems Using Multiple RISs
abstract
In this paper, practical utilization of multiple distributed reconfigurable intelligent surfaces (RISs), which are able to conduct group-specific operations, for multi-group multicasting systems is investigated. To tackle the inter-group interference issue in the multi-group multicasting systems, the block diagonalization (BD)-based beamforming is considered first. Without any inter-group interference after the BD operation, the multiple distributed RISs are operated to maximize the minimum rate for each group. Since the computational complexity of the BD-based beamforming can be too high, a multicasting tailored zero-forcing (MTZF) beamforming technique is proposed to efficiently suppress the inter-group interference, and the novel design for the multiple RISs that makes up for the inevitable loss of MTZF beamforming is also described. Effective closed-form solutions for the loss minimizing RIS operations are obtained with basic linear operations, making the proposed MTZF beamforming-based RIS design highly practical. Numerical results show that the BD-based approach has ability to achieve high sum-rate, but it is useful only when the base station deploys large antenna arrays. Even with the small number of antennas, the MTZF beamforming-based approach outperforms the other schemes in terms of the sum-rate while the technique requires low computational complexity. The results also prove that the proposed techniques can work with the minimum rate requirement for each group.
Hyeongtaek Lee, Seungsik Moon, Youngjoo Lee 0002, Jaeky Oh, Jaehoon Chung, Junil Choi
IEEE Trans. Wirel. Commun.3
2023 GROW: A Row-Stationary Sparse-Dense GEMM Accelerator for Memory-Efficient Graph Convolutional Neural Networks
abstract
Graph convolutional neural networks (GCNs) have emerged as a key technology in various application domains where the input data is relational. A unique property of GCNs is that its two primary execution stages, aggregation and combination, exhibit drastically different dataflows. Consequently, prior GCN accelerators tackle this research space by casting the aggregation and combination stages as a series of sparse-dense matrix multiplication. However, prior work frequently suffers from inefficient data movements, leaving significant performance left on the table. We present GROW, a GCN accelerator based on Gustavson’s algorithm to architect a row-wise product based sparse-dense GEMM accelerator. GROW co-designs the software/ hardware that strikes a balance in locality and parallelism for GCNs, reducing the average memory traffic by 2×, and achieving an average 2.8× and 2.3× improvement in performance and energy-efficiency, respectively.
Ranggi Hwang, Minhoo Kang, Dongyun Kam, Youngjoo Lee 0002, Minsoo Rhu
HPCA5
2023 Energy-Efficient RISC-V-Based Vector Processor for Cache-Aware Structurally-Pruned Transformers
abstract
Based on recent RISC-V designs, we present in this paper a low-power vector processor architecture for efficiently deploying vision transformer (ViT) models. To fairly measure the processing efficiency of different processor designs with instruction/data cache memories, we first develop the evaluation framework based on numerous design tools for jointly considering the algorithm, architecture, and circuit performances together, numerically revealing that the previous CSR-based data compression cannot accelerate pruned transformer models at all due to under-utilization of the vector-extended processing units. We then introduce a series of algorithm-hardware co-optimization approaches to greatly minimize cache misses by applying 1) the accuracy-preserved structured ViT pruning, 2) the vertical-CSR (vCSR) data storing format, and 3) vCSR-aware custom memory-accessing instructions. Experimental results show that the proposed optimization schemes eventually improve the processing efficiency of pruned transformers in resource-limited computing platforms, e.g., achieving 11 times lower energy consumption for handling the 0.7-pruned ViT model.
Jung Gyu Min, Dongyun Kam, Younghoon Byun, Gunho Park, Youngjoo Lee 0002
ISLPED5
2023 Low-Latency SCL Polar Decoder Architecture Using Overlapped Pruning Operations
abstract
Allowing the superior error-correction performance even for short-length codewords, the successive-cancellation list (SCL) decoding algorithm has allowed the polar code to be adopted in 5G New Radio standard for control channel. However, existing SCL polar decoders still suffer from long processing latency caused by a number of serialized internal operations. In this work, to solve the latency problem, we present several parallel computing solutions for the serialized operations, i.e., simplified data dependencies and two overlapped pruning operations. To realize the proposed parallel computing, we also introduce internal circuit blocks including dual read-port buffers, an on-the-fly parity checker, and overlapped processing units. The proposed SCL polar decoders are precisely designed with optimal design parameters by analyzing trade-offs between the latency reduction and area overheads. Implemented in a 65-nm CMOS technology, the proposed list-8 SCL polar decoder requires only 374 ns to handle a (1024, 512) 5G codeword, improving the decoding efficiency by 34.7% compared to the previous designs.
Dongyun Kam, Byeong Yong Kong, Youngjoo Lee 0002
IEEE Trans. Circuits Syst. I Regul. Pap.3
2023 A Scalable Precoding Processor for Large-Scale MU-MIMO Systems
abstract
The number of devices served by baseband stations is constantly increasing due to the rising data traffic in modern communication systems. In order to support large-scale multi-user multiple-input multiple-output (MU-MIMO) systems and achieve their capacity, it is necessary to consider power allocation and user selection along with precoding. This paper introduces a scalable MU-MIMO precoding processor that solves the joint optimization problem for precoding, power allocation, and user selection. We define custom vector instructions and dedicate vector arithmetic operators based on the RV32IM instruction set architecture to efficiently support various MU-MIMO baseband processing scenarios. The proposed vector operators include parallel dual-precision multipliers to enable energy-efficient processing by adjusting the computing resolution of each step without degrading the algorithm-level quality. The proposed processor is fabricated using 28nm CMOS technology and is capable of solving the state-of-the-art joint optimization problem in only 0.51ms for the$64\times 64$large-scale MU-MIMO configuration. Our processor achieves up to 17.4 times higher processing efficiency compared to previous precoder design, even when supporting the largest number of users and the most complicated algorithm.
Seungsik Moon, Namyoon Lee, Youngjoo Lee 0002
IEEE Trans. Circuits Syst. I Regul. Pap.3
2023 Block Orthogonal Sparse Superposition Codes for Ultra-Reliable Low-Latency Communications
abstract
Low-rate and short-packet transmissions are important for ultra-reliable low-latency communications (URLLC). In this paper, we put forth a new family of sparse superposition codes for URLLC, called block orthogonal sparse superposition (BOSS) codes. We first present a code construction method for the efficient encoding of BOSS codes. The key idea is to construct codewords by the superposition of the orthogonal columns of a dictionary matrix with a sequential bit mapping strategy. We also propose an approximate maximum a posteriori probability (MAP) decoder with two stages. The approximate MAP decoder reduces the decoding latency significantly via a parallel decoding structure while maintaining a comparable decoding complexity to the successive cancellation list (SCL) decoder of polar codes. Furthermore, to gauge the code performance in the finite-blocklength regime, we derive an exact analytical expression for block-error rates (BLERs) of single-layered BOSS codes in terms of relevant code parameters. Lastly, we present a cyclic redundancy check aided-BOSS (CA-BOSS) code with simple list decoding to boost the code performance. Our experiments verify that CA-BOSS codes with the simple list decoder outperform CA-polar codes with SCL decoding in the low-rate and finite-blocklength regimes while achieving the finite-blocklength capacity upper bound within one dB of signal-to-noise ratio.
Donghwa Han, Jeonghun Park, Youngjoo Lee 0002, H. Vincent Poor, Namyoon Lee
IEEE Trans. Commun.3
2022 Design and Evaluation Frameworks for Advanced RISC-based Ternary Processor
abstract
In this paper, we introduce the design and veri-fication frameworks for developing a fully-functional emerging ternary processor. Based on the existing compiling environments for binary processors, for the given ternary instructions, the software-level framework provides an efficient way to convert the given programs to the ternary assembly codes. We also present a hardware-level framework to rapidly evaluate the performance of a ternary processor implemented in arbitrary design technology. As a case study, the fully-functional 9-trit advanced RISC-based ternary (ART-9) core is newly developed by using the proposed frameworks. Utilizing 24 custom ternary instructions, the 5-stage ART-9 prototype architecture is successfully verified by a number of test programs including dhrystone benchmark in a ternary domain, achieving the processing efficiency of 57.8 DMIPS/W and$3.06\times 10^{6}$DMIPS/W in the FPGA-level ternary-logic emulations and the emerging CNTFET ternary gates, respectively.
Dongyun Kam, Jung Gyu Min, Jongho Yoon 0001, Sunmean Kim, Seokhyeong Kang, Youngjoo Lee 0002
DATE6
2022 Convolutional Neural Networks With Discrete Cosine Transform Features
abstract
The Discrete Cosine Transform (DCT) exposes features of an image that are not evident in the original image's spatial domain. This brief contribution proposes a Convolutional Neural Network architecture that combines features from the spatial domainandthe DCT domain to improve image classification performance with negligible overhead.
Sanghyeon Ju, Youngjoo Lee 0002, Sunggu Lee
IEEE Trans. Computers2
2022 Low-Complexity and Low-Latency SVC Decoding Architecture Using Modified MAP-SP Algorithm
abstract
The compressive sensing (CS) based sparse vector coding (SVC) method is one of the promising ways for the next-generation ultra-reliable and low-latency communications. In this paper, we present advanced algorithm-hardware co-optimization schemes for realizing a cost-effective SVC decoding architecture. The previous maximum a posteriori subspace pursuit (MAP-SP) algorithm is newly modified to relax the computational overheads by applying novel residual forwarding and LLR approximation schemes. A fully-pipelined parallel hardware is also developed to support the modified decoding algorithm, reducing the overall processing latency, especially at the support identification step. In addition, an advanced least-square-problem solver is presented by utilizing the parallel Cholesky decomposer design, further reducing the decoding latency with parallel updates of support values. The implementation results from a 22nm FinFET technology showed that the fully-optimized design is 9.6 times faster while improving the area efficiency by 12 times compared to the baseline realization.
Seungwoo Hong, Dongyun Kam, Sangbu Yun, Jeongwon Choe, Namyoon Lee, Youngjoo Lee 0002
IEEE Trans. Circuits Syst. I Regul. Pap.6
2022 CHAMP: Channel Merging Process for Cost-Efficient Highly-Pruned CNN Acceleration
abstract
This paper presents an advanced offline scheduling scheme to improve the accelerator efficiency, especially for the highly-pruned convolutional neural networks (HP-CNNs). Based on the existing outlier-aware accelerator design, we demonstrate the efficiency drop of HP-CNN processing for the first time, and element-wise channel merging is proposed to make a dense processing sequence even for the highly-pruned model. The dedicated hardware architecture is also presented to process the merged channels with the minimum hardware-level overheads, improving the energy efficiency for handling HP-CNNs by preserving the hardware utilization. We further investigate the optimal size of accumulator and multiplexer, in addition to the number of merged channels, exploiting the attractive energy-performance trade-offs. As a result, unlike the practical HP-CNNs for the on-device solutions, the proposed method enhances the overall efficiency by up to 33% compared to the state-of-the-art schemes.
Hyeokjun Kwon, Younghoon Byun, Seokhyeong Kang, Youngjoo Lee 0002
IEEE Trans. Circuits Syst. I Regul. Pap.4
2021 Approach to Improve the Performance Using Bit-level Sparsity in Neural Networks
abstract
This paper presents a convolutional neural network (CNN) accelerator that can skip zero weights and handle outliers, which are few but have a significant impact on the accuracy of CNNs, to achieve speedup and increase the energy efficiency of CNN. We propose an offline weight-scheduling algorithm which can skip zero weights and combine two non-outlier weights simultaneously using bit-level sparsity of CNNs. We use a reconfigurable multiplier-and-accumulator (MAC) unit for two purposes; usually used to compute combined two non-outliers and sometimes to compute outliers. We further improve the speedup of our accelerator by clipping some of the outliers with negligible accuracy loss. Compared to DaDianNao [7] and Bit-Tactical [16] architectures, our CNN accelerator can improve the speed by 3.34 and 2.31 times higher and reduce energy consumption by 29.3% and 30.2%, respectively.
Yesung Kang, Eunji Kwon, Seunggyu Lee, Younghoon Byun, Youngjoo Lee 0002, Seokhyeong Kang
DATE5
2021 Low-Latency Polar Decoder Using Overlapped SCL Processing
abstract
In this paper, we present a novel scheduling method that reduces the latency of polar decoders significantly. Unlike the prior pruning-based successive cancellation list (SCL) decoding that suffers from a number of idle cycles, the proposed overlapped SCL scheme immediately begins node operations without waiting for the list to be sorted, being exempt from such unfavorable cycles. All possible candidates for the next node operations are precomputed in parallel with the pruning operations, and are readily selected to minimize the latency. For the 5G New Radio systems, the proposed method shortens the decoding latency of the state-of-the-art approaches by up to 22% without degrading the error-correcting performance.
Dongyun Kam, Byeong Yong Kong, Youngjoo Lee 0002
ICASSP3
2021 Rapid Design Space Exploration of Near-Optimal Memory-Reduced DCNN Architecture Using Multiple Model Compression Techniques
abstract
In spite of the attractive accuracy, it is hard to use a deep convolutional neural network (DCNN) directly at the resource-limited devices due to the energy-consuming memory overheads, and thus the aggressive compression schemes are essentially utilized in practice to reduce the DCNN model size. As the recent methods have been individually developed, however, it is inevitable to exhaustively find the optimal combination of different approaches, requiring an enormous amount of search time. Given the complex baseline network, in this work, we introduce a rapid and systematic way to find the near- optimal memory-reduced DCNN option using multiple compression schemes together. We first precisely observe the accuracy-size trade-off of each method and make a novel interpolating scheme to speculate the accuracy of an arbitrary combination. We then present an iterative search algorithm to minimize the number of network evaluations for finding the memory-efficient DCNN structure satisfying the required accuracy. Experimental results reveal that our framework provides a similar compression level to the naive full-search strategy with three popular optimization methods while saving the search time by 7.35 times.
Younghoon Byun, Youngjoo Lee 0002
ISCAS2
2021 Layerwise Buffer Voltage Scaling for Energy-Efficient Convolutional Neural Network
abstract
In order to effectively reduce buffer energy consumption, which constitutes a significant part of the total energy consumption in a convolutional neural network (CNN), it is useful to apply different amounts of energy conservation effort to the different levels of a CNN as the buffer energy to total energy usage ratios can differ quite substantially across the layers of a CNN. This article proposes layerwise buffer voltage scaling as an effective technique for reducing buffer access energy. Error-resilience analysis, including interlayer effects, conducted during design-time is used to determine the specific buffer supply voltage to be used for each layer of a CNN. Then these layer-specific buffer supply voltages are used in the CNN for image classification inference. Error injection experiments with three different types of CNN architectures show that, with this technique, the buffer access energy and overall system energy can be reduced by up to 68.41% and 33.68%, respectively, without sacrificing image classification accuracy.
Minho Ha, Younghoon Byun, Seungsik Moon, Youngjoo Lee 0002, Sunggu Lee
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2021 Design and Analysis of Approximate Compressors for Balanced Error Accumulation in MAC Operator
abstract
In this paper, we present a novel approximate computing scheme suitable for realizing the energy-efficient multiply-accumulate (MAC) processing. In contrast to the prior works that suffer from the error accumulation limiting the approximate range, we utilize different approximate multipliers in an interleaved way to compensate errors in the opposite direction during accumulate operations. For the balanced error accumulation, we first design the approximate 4-2 compressors generating errors in the opposite direction while minimizing the computational costs. Based on the probabilistic analysis, positive and negative multipliers are then carefully developed to provide a similar error distance. Simulation results on various practical applications reveal that the proposed MAC processing offers the energy-efficient computing scenario by extending the range of approximate parts. Even compared to the state-of-the-art solutions, for example, the proposed interleaving scheme relaxes the core-level energy consumption of the recent CNN accelerator by more than 35% without degrading the recognition accuracy.
Gunho Park, Jaeha Kung 0001, Youngjoo Lee 0002
IEEE Trans. Circuits Syst. I Regul. Pap.3
2021 Ultralow-Latency Successive Cancellation Polar Decoding Architecture Using Tree-Level Parallelism
abstract
Achieving the attractive error-correcting capability with a simple decoder structure, the polar code using successive cancellation (SC) decoding is now expected to be installed at the resource-limited IoT or embedded communications. However, the existing SC decoders normally suffer from the long processing latency caused by the serialized processing steps, limiting the practical applications of polar codes. In this article, to solve this latency problem, we present a new low-complexity merging operation that can increase the number of parallel factors for realizing the tree-level parallelism. We also modify the previous pruning method to further reduce the number of visited nodes at the parallel SC decoding scenario. In addition, a novel parallel partial-sum calculator (PSC) architecture is introduced to update partial-sum registers with multiple decoded bits by taking only one processing cycle. Implementation results show that the proposed 8-parallel SC polar decoder in 28-nm CMOS requires only 0.140$\mu \text{s}$to decode a (1024, 512) codeword of 5G system, remarkably reducing the decoding latency when compared to the state-of-the-art designs.
Dongyun Kam, Hoyoung Yoo, Youngjoo Lee 0002
IEEE Trans. Very Large Scale Integr. Syst.3
2020 Ultra-Low-Latency LDPC Decoding Architecture using Reweighted Offset Min-Sum Algorithm
abstract
Due to an iterative nature, a low-density parity-check (LDPC) decoder is associated with a long latency, being a major bottleneck of the baseband processor in wireless communication systems. Based on the practical min-sum (MS) decoding method, in this paper, we present a cost-effective algorithm for reducing the processing latency of LDPC decoders. By checking the number of short-length cycles in the LDPC code structure, the proposed method dynamically changes the reweighting factor at the iterative operations, successfully reducing the average number of iterations. In addition, we present several optimization schemes to mitigate the hardware overheads resulting from the proposed reweighting scheme. In a 65-nm CMOS process, a prototype IEEE 802.11ay LDPC decoder optimized by the proposed schemes reduces the decoding latency by 1.7 times with negligible overheads compared with the contemporary designs.
Sangbu Yun, Dongyun Kam, Jeongwon Choe, Byeong Yong Kong, Youngjoo Lee 0002
ISCAS5
2019 Low-Complexity Dynamic Channel Scaling of Noise-Resilient CNN for Intelligent Edge Devices
abstract
In this paper, we present a novel channel scaling scheme for convolutional neural networks (CNNs), which can improve the recognition accuracy for the practical distorted images without increasing the network complexity. During the training phase, the proposed work first prepares multiple filters under the same CNN architecture by taking account of different noise models and strengths. We then newly introduce an FFT-based noise classifier, which determines the noise property in the received input image by calculating the partial sum of the frequency-domain values. Based on the detected noise class, we dynamically change the filters of each CNN layer to provide the dedicated recognition. Furthermore, we propose a channel scaling technique to reduce the number of active filter parameters if the input data is relatively clean. Experimental results show that the proposed dynamic channel scaling reduces the computational complexity as well as the energy consumption, still providing the acceptable accuracy for intelligent edge devices.
Younghoon Byun, Minho Ha, Sunggu Lee, Youngjoo Lee 0002
DATE5
2019 Ultra-Low-Latency Parallel SC Polar Decoding Architecture for 5G Wireless Communications
abstract
In this paper, we newly present a novel parallel polar decoding architecture that significantly reduces the processing latency for 5G wireless communications. Based on the original decoding tree, the proposed scheme first constructs the small trees that generate multiple soft-decision messages in parallel, potentially reducing the decoding latency compared to the previous serialized schemes. The hard-decision estimates are then calculated at the following merging step to decide the decoded outputs and to update the parallel trees. For each parallel tree, the parallel pruning scheme is newly utilized to further optimize the processing latency. In addition, we introduce an efficient parallel decoder architecture, successfully supporting the proposed low-latency algorithm. Implementation results show that the proposed 8-parallel polar decoder in 65nm CMOS uses only 267ns to decode a (1024, 512) polar codeword of 5G system, which is 1.67 times faster than the state-of-the-art design.
Dongyun Kam, Youngjoo Lee 0002
ISCAS2
2019 Similarity-Based LSTM Architecture for Energy-Efficient Edge-Level Speech Recognition
abstract
Targeting the resource-limited edge devices, we present a novel processing architecture of long short-term memory (LSTM) networks for low-power speech recognition. The proposed scheme newly defines the similarity score between two inputs of adjacent LSTM cells, and then the processing mode of the current LSTM cell is dynamically determined to reduce the energy while providing the accurate recognition. If the similarity is high, more precisely, the current cell is disabled and the outputs are directly copied from the prior vectors, totally eliminating complex LSTM operations. To maximize the skipping ratio without degrading the accuracy, for the first time, we analyze the effects of skipping the consecutive cells and set the upper limit of the number of consecutive skips. When two adjacent inputs are weakly similar, in addition, we modify the concept of the previous delta-computing, which approximately activate the LSTM cell with low computational resolution, further reducing the energy consumption. Compared to the previous state-of-the-art solutions, as a result, the proposed LSTM architecture remarkably saves the energy consumed for the accurate speech recognition, which is suitable to the resource-limited embedded edges.
Junseo Jo, Jaeha Kung 0001, Sunggu Lee, Youngjoo Lee 0002
ISLPED4
2019 WMixNet: An Energy-Scalable and Computationally Lightweight Deep Learning Accelerator
abstract
In this paper, we present a lightweight CNN model named MixNet which is easily scalable to different energy requirements in embedded platforms. The MixNet model uses two extreme bit-precisions that efficiently balances the model accuracy and energy consumption. The energy consumption in processing MixNet is managed by controlling the ratio between high-precision (16bit) and low-precision (1bit) paths. Since only two bit-precisions are required in designing a hardware accelerator, the control logic becomes simpler compared to other multi-precision accelerators. In addition, a reconfigurable multiplier is proposed to enable highly parallel MixNet computations for faster prediction and/or training. Overall, the energy efficiency in terms of run-time per unit power improves by 1.75 ~1.94 × over the recently proposed reduced-precision CNN model.
Sangwoo Jung 0001, Seungsik Moon, Youngjoo Lee 0002, Jaeha Kung 0001
ISLPED3
2018 A 2.4pJ/bit, 6.37Gb/s SPC-enhanced BC-BCH decoder in 65nm CMOS for NAND flash storage systems
abstract
This paper present an energy-efficient block-concatenated BCH (BC-BCH) decoder which can achieve superior decoding performance for NAND flash storage systems. To enhance the error-correcting capability, an additional decoding step with single parity-check (SPC) block is newly employed. A novel memory based syndrome updating method effectively improves the energy efficiency as well as the decoding latency. Using the proposed methods, a prototype chip is implemented to decode a (36443, 32768) BC-BCH code in 65nm CMOS process. The proposed decoder provides a decoding throughput of 6.37Gb/s and an efficiency of 2.4pJ/bit, being superior to the state-of-the-art hard-decision decoders for storages.
In-Cheol Park, Youngjoo Lee 0002
ASP-DAC3
2014 7.3 Gb/s universal BCH encoder and decoder for SSD controllers
abstract
This paper presents a universal BCH encoder and decoder that can support multiple error-correction capabilities. A novel encoding architecture and on-demand syndrome calculation technique is proposed to reduce both hardware complexity and power consumption. Based on the proposed methods, 32-parallel universal encoder and decoder are designed for BCH (8192+14t, 8192, t) codes, where the error-correction capability t is configurable to 8, 11, 16, 24, 32, and 64. The prototype chip achieves a throughput of 7.3 Gb/s and occupies 2.24 mm2in 0.13μπι CMOS technology.
Hoyoung Yoo, Youngjoo Lee 0002, In-Cheol Park
ASP-DAC2
2014 High-Throughput and Low-Complexity BCH Decoding Architecture for Solid-State Drives
abstract
This paper presents a high-throughput and low-complexity BCH decoder for NAND flash memory applications, which is developed to achieve a high data rate demanded in the recent serial interface standards. To reduce the decoding latency, a data sequence read from a flash memory channel is re-encoded by using the encoder that is idle at that time. In addition, several optimizing methods are proposed to relax the hardware complexity of a massive-parallel BCH decoder and increase the operating frequency. In a 130-nm CMOS process, a (8640, 8192, 32) BCH decoder designed as a prototype provides a decoding throughput of 6.4 Gb/s while occupying an area of 0.85${\rm mm}^{2}$.
Youngjoo Lee 0002, Hoyoung Yoo, Injae Yoo, In-Cheol Park
IEEE Trans. Very Large Scale Integr. Syst.1
2013 A 3Gb/s 2.08mm2 100b error-correcting BCH decoder in 0.13µm CMOS process
abstract
This paper presents a high-throughput BCH decoder that can correct 100 bit-errors. Several optimization methods are proposed to reduce the hardware complexity caused by the large error-correction capability. Based on the proposed methods, an 8-parallel decoder is designed for the (9592, 8192, 100) BCH code, which achieves a decoding throughput of 3Gb/s and occupies 2.08mm2in 0.13µm CMOS process.
Youngjoo Lee 0002, Hoyoung Yoo, In-Cheol Park
ASP-DAC1
2012 Small-area parallel syndrome calculation for strong BCH decoding
abstract
This paper presents a new optimization method to reduce the hardware complexity of syndrome calculation in strong BCH decoding. All the operations required in the parallel syndrome calculation are reformulated as a single matrix computation to enlarge the search area for common sub-expressions. The computational complexity of syndrome calculation is significantly reduced by finding and sharing common terms in the single matrix computation. Implementation results show that the proposed architecture saves 55% of area overheads compared to the conventional structure.
Youngjoo Lee 0002, Hoyoung Yoo, In-Cheol Park
ICASSP1
2010 Capacitor array structure and switching control scheme to reduce capacitor mismatch effects for SAR analog-to-digital converters
abstract
This paper presents a new capacitor array structure developed for SAR analog-to-digital converters and its switching algorithm that can alleviate capacitor mismatch effects. The capacitor mismatch can induce many missing codes. The proposed capacitor array structure is based on the junction-splitting method as is efficient in terms of power consumption. To reduce the capacitor mismatch effects, two capacitor arrays are employed to enable redundant search. Simulation results show that the proposed method significantly reduces the number of missing codes caused by the capacitance mismatch, and reduces the energy consumption by more than 70% compared to the conventional charge redistribution method.
Youngjoo Lee 0002, In-Cheol Park
ISCAS1
2010 Design of a Scalable and Programmable Sound Synthesizer
abstract
Sound synthesis employed in many multimedia systems is a useful method to generate the sound of musical instruments. Although it has a long history of development, a few researches have been devoted to deriving efficient VLSI architectures. In this paper, we analyze the inherent dataflow of sound synthesis methods and propose a programmable VLSI architecture suitable for a scalable sound synthesizer. The sound quality and the level of polyphony can be enhanced only by increasing the operating speed and enlarging the memory. A fully integrated sound synthesis system is implemented as a prototype to verify the proposed architecture. The prototype chip fabricated in a 0.18- m CMOS process occupies 1.5 mm 1.5 mm, and can synthesize a 64-polyphonic sound in real time. The power consumption ranges from 2.05 to 13.8 mW depending on the level of polyphony and the sound quality.
Youngjoo Lee 0002, In-Cheol Park
IEEE Trans. Very Large Scale Integr. Syst.2
2009 A Scalable and Programmable Sound Synthesizer
abstract
Sound synthesis employed in many multimedia systems is a useful method to generate sounds of musical instruments. In this paper, we propose a new VLSI architecture suitable for a scalable sound synthesizer based on a programmable data-flow. The sound quality and the level of polyphony can be enhanced only by increasing operating speed and enlarging memory. A fully integrated sound synthesis system is implemented as a prototype to verify the proposed architecture. The prototype chip fabricated in a 0.18-mum CMOS process occupies 1.5 times 1.5 mm2, and can synthesize up to 64-polyphonic sound. The power consumption ranges from 2.05 mW to 13.8 mW depending on the quality of the synthesized sound.
Youngjoo Lee 0002, In-Cheol Park
ISCAS2