Kangjoon Choi

dblp:377/1760 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0004-4137-7688ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Hardware-Efficient Unified Approximation for Implementing Diverse Smooth Activation Functions
Jeongmin Kim 0001, Kangjoon Choi, In-Cheol Park
IEEE Trans. Computers2
2026 General-Purpose QC-LDPC Processor for Error-Correction Performance Evaluation
abstract
This paper introduces a general-purpose quasi-cyclic (QC)-low-density parity-check (LDPC) processor (GPLP) developed to evaluate LDPC decoding performance, which is a highly flexible and programmable field programmable gate array (FPGA)-based simulation platform. The GPLP enables to configure LDPC decoding processes, depending on the decoding algorithms and design choices to be simulated and verified. In contrast to traditional FPGA-based approaches, which typically rely on a fixed dataflow, the GPLP offers a reconfigurable, processor-like environment that can be programmed for various design configurations. This paper defines essential atomic operations required for QC-LDPC decoding and proposes the GPLP architecture supporting those operations with vector-based processing. The proposed GPLP significantly enhances simulation speed, providing a performance improvement of 220x to 574x over CPU-based simulations. The platform provides a unique combination of simulation speed, flexibility, and reconfigurability, making it an invaluable tool to rapidly evaluate and optimize LDPC decoding performance.
Kangjoon Choi, Jongmin Baek, In-Cheol Park
IEEE Trans. Circuits Syst. I Regul. Pap.1
2026 Enhanced-Throughput Antithetic Sample Generation Architecture for Monte Carlo Simulation
abstract
Monte Carlo simulation imposes significant computational and hardware costs due to extensive sample generation. While antithetic variates can reduce sample requirements, direct hardware application doubles transformation paths. Based on a point-symmetric inverse cumulative distribution, this brief introduces a hardware architecture that generates antithetic pairs using a single transformation path and a sign inversion, guaranteeing strong negative correlation and effective variance reduction. Implemented on a Xilinx ZCU104 field-programmable gate array (FPGA), the architecture doubles throughput with minimal hardware overhead. In Monte Carlo bit error rate (BER) simulation for an 802.11n LDPC code, the proposed architecture maintains accuracy while achieving a$2.14\times $to$3.62\times $speed gain in random-variable generation.
Kangjoon Choi, In-Cheol Park
IEEE Trans. Very Large Scale Integr. Syst.1
2025 Hardware-Efficient Architecture for Multiple Quantized Gaussian Noise Generation
abstract
This paper presents two novel architectures to generate a number of quantized Gaussian noises. The first architecture exploits inversion through uniform segmentation, enabling a uniform look up table (LUT) splitting technique to efficiently generate quantized Gaussian noise while maintaining reasonable tail quality in Gaussian noise generation. The second architecture utilizes inversion through hierarchical segmentation and a probability-based LUT selection, significantly reducing the total LUT size while preserving the tail quality of the generated Gaussian noise. Both designs generate multiple uniform random numbers by cascading combinational circuits, which improves Gaussian noise generation efficiency compared to the conventional linear feedback shift register-based method. Compared to the previous architecture based on inversion through hierarchical segmentation, the proposed uniform segmentation architecture achieves a 6.02x improvement, when implemented on a field-programmable gate array device, in terms of throughput per configurable logic block, and the proposed hierarchical segmentation architecture achieves a 2.71x improvement.
Kangjoon Choi, In-Cheol Park
IEEE Trans. Circuits Syst. I Regul. Pap.1
2025 Multiple-Resolution Decoding Architecture for QC-LDPC Codes
abstract
This paper proposes a hardware-efficient multiple-resolution decoding architecture for quasi-cyclic low-density parity-check (QC-LDPC) codes, specifically designed to meet the stringent requirements of the 5G New Radio (NR) standard. The proposed architecture adopts a single-instruction multiple-data (SIMD) approach to dynamically adjust the bitwidth of log-likelihood ratio (LLR) values based on theEb/N0condition, significantly reducing hardware complexity and improving throughput area ratio. Unlike conventional single-resolution decoders, the architecture processes 2-bit LLR values in highEb/N0regions and scales up to 4-bit or 8-bit LLR values for moderate and lowEb/N0conditions, maintaining robust error-correcting performance. Key innovations include SIMD-based design for variable-node units (VNUs), check-node units (CNUs), and quasi-cyclic shifting networks (QSNs), as well as optimized memory access scheduling to support all 51 lifting sizes defined in the 5G NR standard. Designed in a 65-nm CMOS process, the decoder achieves a peak throughput of 27.24 Gbps under error-free conditions with a throughput area ratio improvement of 2.07× compared to the state-of-the-art designs. Furthermore, the proposed architecture demonstrates superior throughput-area ratio and flexibility, supporting all code rates and lifting sizes specified in the 5G NR standard. Simulation results confirm that the proposed decoder meets the peak throughput requirement under error-free conditions, while maintaining robust performance in challenging channel environments.
In-Cheol Park, Kangjoon Choi, Hyejung Jang, Jongmin Baek
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 Area-Efficient QC-LDPC Decoding Architecture With Thermometer Code-Based Sorting and Relative Quasi-Cyclic Shifting
abstract
The 5G New-Radio (NR) communication standard requires high throughput and low latency, so low-density parity-check (LDPC) codes, which have higher inherent parallelism and lower decoding complexity than turbo codes, were adopted as the main coding method for data channels. In traditional LDPC min-sum decoders, the check node unit was realized using a sorting unit based on the min-tree structure. However, this structure resulted in high hardware complexity and long latency. To address this issue, we propose a new sorting method based on the thermometer code-based number system. Additionally, we introduce a new LDPC decoding architecture that reduces the number of QSN stages from two to one, significantly lowering the shifting logic complexity needed to support different lifting sizes. This is achieved by using relative shift amounts instead of absolute shift amounts specified in the parity check matrix. The proposed decoder implemented using a partially parallel structure in a 65nm CMOS technology satisfies the various operation modes and the throughput requirements of the 5G NR standard, and boasts a higher normalized throughput than state-of-the-art LDPC decoders.
Boseon Jang, Hyejung Jang, Kangjoon Choi, In-Cheol Park
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 Hardware-Efficient SoftMax Architecture With Bit-Wise Exponentiation and Reciprocal Calculation
abstract
The SoftMax function is one of the activation functions used in deep neural networks (DNN) to normalize input values to the range of (0,1). With the advent of DNN models including the Transformer, operations utilizing SoftMax have gained significant attention, and the efficient hardware implementation of such operations has become a prominent issue in hardware realization. Implementing SoftMax often involves exponential and division operations, which can be a significant bottleneck in terms of hardware cost and performance. Various efforts have been made to address this challenge, and this paper introduces a novel approach to efficiently implement SoftMax. In most previous works, the maximum input value is subtracted from all the input values to ensure numerical stability. In the proposed approach, the maximum value is replaced with a different value to reduce the hardware complexity with ensuring numerical stability. Additionally, in exponential operations, simple Look-Up Tables (LUTs) with only one entry each are used for bit-wise calculations, and the reciprocal of the total exponential sum is computed to replace division with multiplication. Applying the proposed methods reduces the computational complexity significantly compared to the previous log-sum-exp approach. As a result, the proposed 8-bit SoftMax accelerator achieves a high operating frequency of 3.12GHz and a high throughput of 25G inputs/s. It also improves area efficiency and power consumption by at least 2 times. From an accuracy perspective, furthermore, it is associated with similar or even better accuracy compared to previous works.
Jeongmin Kim 0001, Kangjoon Choi, In-Cheol Park
IEEE Trans. Circuits Syst. I Regul. Pap.3