Byeong Yong Kong

dblp:119/0553 · DBLP profile ↗
← Back
18ranked-venue papers
12as first author
8since 2021 · last 2026
0000-0001-5823-5505ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 8 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 Fault-Tolerant IDMA Multiuser Detector Based on Fault Injection Analysis of Internal Memories
abstract
In this article, a fault-tolerant architecture (FTA) is presented for the multiuser detector in interleave division multiple access (IDMA). The detector is inherently prone to soft errors, as its chip area is predominantly occupied by memories, which are easily exposed to high-energy particles and hostile interferences in harsh environments. One of the most widespread ways to protect memories is to encode their entries with error-correcting codes (ECCs). However, naïvely encoding all bits in an entry is likely to be costly and unnecessary. Accordingly, to sort out performance-critical bits and determine the priority of protection, we extensively scrutinize how vulnerable respective bits in the memories of the detector are too soft errors. Based on the analysis, in addition, an efficient FTA that selectively encodes only a subset of the bits in order of the identified vulnerability is developed. Furthermore, the proposed FTA implements the state-of-the-art multiuser detection (MUD) scheme called on-the-fly despreading (OD) and showcases a new feature named purification, which repeatedly replaces erroneous entries with corrected ones to keep them error-free. Complicated memory accesses to concurrently perform the OD as well as the purification are enabled by remodeling both the datapath and the control path of the baseline OD architecture (ODA). Implementation results demonstrate that, unlike the prior arts that fail to sustain near-optimal performances and become impractical even for a very low probability of soft error, the proposed FTA may operate robustly in a wide range of harsh conditions without incurring much overhead.
Byeong Yong Kong
IEEE Trans. Very Large Scale Integr. Syst.1
2025 Generalized Shifting-Based PAPR Reduction Architecture for Discrete Multitone Systems
abstract
In this paper, a generalized shifting-based peak-to-average power ratio (PAPR) reduction architecture is presented for discrete multitone (DMT) systems. The limitation of the state-of-the-art scheme is that it resorts only to the local extrema from subsets in calculating the compensation offset, being inherently prone to accuracy degradation. To overcome the drawback, the proposed generalized offset enables new design options that utilize multiple global extrema, mitigating the PAPR further. In addition to the generalization, a low-complexity bilateral sorter, which is the key component of the entire architecture to identify multiple extrema, is developed by simplifying and hybridizing mergesorters. Since the generalized offset not only offers the new options but also encompasses all the previous expressions of the offsets in the literature, it successfully unifies all dc shifting-based PAPR reduction schemes, and exempts us from comparing them and choosing one.
Byeong Yong Kong
ISCAS1
2024 Constrained Sorter Design using Zero-One Principle
abstract
To derive efficient sorting architectures constrained to application-specific input/output conditions, we present in this paper a systematic design methodology that can effectively prune dispensable compare-and-swap (CAS) units. Unlike the previous works resorting to heuristic approaches, the proposed framework exploits the zero-one principle to validate the pruning of a CAS unit at a time, generating the cost-optimized sorter architecture in an iterative manner with a reasonable complexity. In addition to the given input/output constraints, we newly develop the architecture options for the proposed framework, allowing more design spaces for finding the most attractive constrained-sorter design. For 8-list polar decoders, the proposed framework successfully reduces 70% of CAS units in the baseline full sorter, relaxing the area-time complexity by 35% compared with the state-of-the-art solutions.
Sangil Han, Jaehee Kim, Dongyun Kam, Byeong Yong Kong, Mijung Kim, Young-Seok Kim, Youngjoo Lee 0002
ISCAS4
2024 A Design Framework for Cost-Efficient Sorters With Arbitrary Input/Output Constraints
abstract
The sorting operation plays a vital role in various signal processing applications. However, due to high hardware complexity resulting from a series of comparisons, designing the cost-efficient sorter is one of the crucial requisites for improving the overall system performance. To obtain the cost-efficient sorting architectures constrained to application-specific input/output conditions, this paper presents a systematic design methodology that effectively eliminates dispensable compare-and-swap (CAS) units. Unlike the previous heuristic approaches, the proposed framework iteratively prunes a CAS unit followed by the validation step. The zero-one principle is newly applied to reduce the validation time for the practical convergence time with massive searching iterations. To expand the search space of the proposed framework, furthermore, we introduce new architectural options and pruning methods, allowing the cost-efficient design results even compared to the state-of-the-art solutions. Targeting the constrained sorters for communication systems, numerous case studies show that the proposed framework successfully removes more than half of CAS units in the baseline sorter design, significantly relaxing the area-time complexity, e.g., 35% reduction compared to the state-of-the-art architecture for 16-input metric sorter in the SCL polar decoder.
Jaehee Kim, Sangil Han, Dongyun Kam, Byeong Yong Kong, Youngjoo Lee 0002
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 Low-Latency PAPR Reduction Architecture for Discrete Multitone Based on Approximate Midrange
abstract
In this brief, a low-latency hardware architecture is presented for the peak-to-average power ratio (PAPR) reduction in discrete multitone (DMT) systems. The state-of-the-art scheme suffers from high latency due to a series of selections in calculating the midrange of the DMT frame. To overcome the drawback, an approximate calculation of the midrange that transforms the dominant operation from the selections to the summation is proposed. Grounded on the transformation, in addition, a low-latency architecture to calculate the approximate midrange is constituted. While the selections involve negation, carry propagation, and multiplexing in every stage, the summation can be promptly done by carry-save adders (CSAs) without such manipulations. Slightly relaxing the accuracy of the midrange, as a result, the overall latency can be effectively alleviated compared with the state-of-the-art counterpart.
Byeong Yong Kong
IEEE Trans. Very Large Scale Integr. Syst.1
2023 Low-Latency SCL Polar Decoder Architecture Using Overlapped Pruning Operations
abstract
Allowing the superior error-correction performance even for short-length codewords, the successive-cancellation list (SCL) decoding algorithm has allowed the polar code to be adopted in 5G New Radio standard for control channel. However, existing SCL polar decoders still suffer from long processing latency caused by a number of serialized internal operations. In this work, to solve the latency problem, we present several parallel computing solutions for the serialized operations, i.e., simplified data dependencies and two overlapped pruning operations. To realize the proposed parallel computing, we also introduce internal circuit blocks including dual read-port buffers, an on-the-fly parity checker, and overlapped processing units. The proposed SCL polar decoders are precisely designed with optimal design parameters by analyzing trade-offs between the latency reduction and area overheads. Implemented in a 65-nm CMOS technology, the proposed list-8 SCL polar decoder requires only 374 ns to handle a (1024, 512) 5G codeword, improving the decoding efficiency by 34.7% compared to the previous designs.
Dongyun Kam, Byeong Yong Kong, Youngjoo Lee 0002
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 Low-Latency Polar Decoder Using Overlapped SCL Processing
abstract
In this paper, we present a novel scheduling method that reduces the latency of polar decoders significantly. Unlike the prior pruning-based successive cancellation list (SCL) decoding that suffers from a number of idle cycles, the proposed overlapped SCL scheme immediately begins node operations without waiting for the list to be sorted, being exempt from such unfavorable cycles. All possible candidates for the next node operations are precomputed in parallel with the pruning operations, and are readily selected to minimize the latency. For the 5G New Radio systems, the proposed method shortens the decoding latency of the state-of-the-art approaches by up to 22% without degrading the error-correcting performance.
Dongyun Kam, Byeong Yong Kong, Youngjoo Lee 0002
ICASSP2
2021 Real-Time SSDLite Object Detection on FPGA
abstract
Deep neural network (DNN)-based object detection has been investigated and applied to various real-time applications. However, it is hard to employ the DNNs in embedded systems due to their high computational complexity and deep-layered structure. Although several field-programmable gate array (FPGA) implementations have been presented recently for real-time object detection, they suffer from either low throughput or low detection accuracy. In this article, we propose an efficient computing system for real-time SSDLite object detection on FPGA devices, which includes novel hardware architecture and system optimization techniques. In the proposed hardware architecture, a neural processing unit (NPU) that consists of heterogeneous units, such as band processing, scaling, and accumulating, and data fetching and formatting units is designed to accelerate the DNNs efficiently. In addition, system optimization techniques are presented to improve the throughput further. A task control unit is employed to balance the workload and increase the utilization of heterogeneous units in the NPU, and the object detection algorithm is refined accordingly. The proposed architecture is realized on an Intel Arria 10 FPGA and enhances the throughput by up to 13.6× compared to the state-of-the-art FPGA implementation.
Suchang Kim, Seungho Na, Byeong Yong Kong, Jaewoong Choi, In-Cheol Park
IEEE Trans. Very Large Scale Integr. Syst.3
2020 Improved Parallel-IDMA Architecture with Low-Complexity Elementary Signal Estimators
abstract
This paper presents a low-complexity parallel multiuser detector architecture for interleave division multiple access systems. To facilitate efficient hardware implementation, the formulae associated with the elementary signal estimator (ESE) are reorganized. Grounded on the reorganization, the proposed ESE circumvents redundant computations, and takes advantage of carry-save additions. The resulting datapath is further simplified by approximations that do not deteriorate the error rate noticeably. Prior to integrating with such ESEs, the state-of-the-art parallel architecture for the user-specific processing block (UPB) is also simplified by rescheduling the memory-access pattern. As a result, a prototype 2-parallel 16-user detector that incorporates the proposed ESEs and UPBs in a 65-nm CMOS occupies 23% less silicon area, dissipates 21% less power, and takes 50% less latency than the state-of-the-art serial detector.
Byeong Yong Kong
ISCAS1
2020 A Low-Latency Multi-Touch Detector Based on Concurrent Processing of Redesigned Overlap Split and Connected Component Analysis
abstract
A low-latency multi-touch detector architecture for locating numerous touches in large-panel devices is presented in this paper. Two respective processors for the overlap split and the connected component analysis (CCA) are the key components in a multi-touch detector. Exploiting the simplicity of typical intensity maps in practice, the two processors are first redesigned under a principle that every pixel is processed only once in a raster-scan order. More specifically, the concept of valley point is introduced, and a simple yet effective overlap-split scheme called valley-point division (VPD) is newly developed based on the concept. In addition, equivalent labels in the CCA are handled on the fly instead of being kept to be processed later, and the logics and memories for corner cases that never occur in practice are unloaded. Enabled by the redesign, subsequently, the two processors are integrated into one to concurrently conduct both the VPD and the CCA during the same raster scan. As a result, the proposed detector completes the whole detection procedure with a single raster scan of a map, and is exempted from a large memory required to hold an entire map. Implementation results in a 65-nm CMOS for a large panel of 400 × 250 sensors demonstrate that the proposed detector takes less than a half latency of the existing ones while occupying only 17% silicon area and consuming 52% power on average.
Byeong Yong Kong, Jooseung Lee, In-Cheol Park
ISCAS1
2020 A 120-mW 0.16-ms-Latency Connectivity-Scalable Multiuser Detector for Interleave Division Multiple Access
abstract
To facilitate the massive connectivity of the 5G New Radio, this brief proposes a connectivity-scalable multiuser detector (MUD) architecture for interleave division multiple access (IDMA). A regular inter-MUD interface makes the proposed MUD not only operable alone but also scalable to serve a massive number of users by connecting multiple MUDs. Besides, the memory subsystem is simplified to mitigate the power consumption and the silicon area. A prototype 16-MUD in a 65-nm CMOS consumes 120 mW, occupies 8.21 mm2, and takes 0.16 ms to process one 8192-chip frame. Such numbers represent 29% less silicon area and 59% less power dissipation than those of the state-of-the-art MUD.
Byeong Yong Kong, In-Cheol Park
ISCAS1
2020 Ultra-Low-Latency LDPC Decoding Architecture using Reweighted Offset Min-Sum Algorithm
abstract
Due to an iterative nature, a low-density parity-check (LDPC) decoder is associated with a long latency, being a major bottleneck of the baseband processor in wireless communication systems. Based on the practical min-sum (MS) decoding method, in this paper, we present a cost-effective algorithm for reducing the processing latency of LDPC decoders. By checking the number of short-length cycles in the LDPC code structure, the proposed method dynamically changes the reweighting factor at the iterative operations, successfully reducing the average number of iterations. In addition, we present several optimization schemes to mitigate the hardware overheads resulting from the proposed reweighting scheme. In a 65-nm CMOS process, a prototype IEEE 802.11ay LDPC decoder optimized by the proposed schemes reduces the decoding latency by 1.7 times with negligible overheads compared with the contemporary designs.
Sangbu Yun, Dongyun Kam, Jeongwon Choe, Byeong Yong Kong, Youngjoo Lee 0002
ISCAS4
2019 Parallel IDMA Architecture Based on Interleaving with Replicated Subpatterns
abstract
This paper presents a parallel multiuser detector architecture for low-latency interleave division multiple access. To enable P-parallel processing, an interleaving pattern is divided into P disjoint subpatterns, and all the subpatterns are designed to be identical without degrading error-rate performance noticeably. Since the subpatterns are all disjoint, they can be processed in parallel. Besides, by exploiting that they access the same address of separate memory banks at the same time, the banks are integrated into one to minimize the silicon area and the power consumption. As a result, the proposed architecture reduces the latency by a factor of P at the expense of a little hardware overhead. A prototype 2-parallel 16-user detector in a 65-nm CMOS completes the entire detection procedure two times earlier than the state-of-the-art nonparallel detector, while occupying only 12% more silicon area and dissipating 20% more power.
Byeong Yong Kong, In-Cheol Park
ICC1
2016 Low-complexity symbol detection for massive MIMO uplink based on Jacobi method
abstract
In this paper, we propose a low-complexity symbol detection algorithm for massive multiple-input multiple-output (MIMO) uplink. Grounded on the fact that a primary property of the massive MIMO systems guarantees the convergence of the Jacobi method, the method is exploited in the linear detection so that the estimate of transmitted symbols can be obtained without employing the computationally intensive matrix inversion. In addition, we propose a multiplication-free initial estimate for the Jacobi method in order to lessen the computational complexity further. Owing to the elimination of matrix inversion and the efficient initial estimate, the proposed algorithm achieves near-optimal error-rate performance with fewer computations than the state-of-the-art schemes.
Byeong Yong Kong, In-Cheol Park
PIMRC1
2015 Narrow-range frequency estimation based on comprehensive optimization of DFT and interpolation
abstract
An efficient procedure for frequency estimation is proposed in this paper to alleviate the computational complexity. Grounded on the fact that the frequency of a target signal usually lies in a known range in practical applications, two fundamental steps in the frequency estimation, i.e., the discrete Fourier transform (DFT) and the interpolation of the DFT samples, are modified accordingly. Unlike the previous works focusing on either the DFT or the interpolation, this paper does not decouple the two steps but optimizes the whole procedure comprehensively by considering the interrelationship between the two steps. As a result, the number of operations required for the estimation is remarkably diminished while the performance remains competitive with the recent works.
Byeong Yong Kong, In-Cheol Park
ICASSP1
2014 Low-Complexity Low-Latency Architecture for Matching of Data Encoded With Hard Systematic Error-Correcting Codes
abstract
A new architecture for matching the data protected with an error-correcting code (ECC) is presented in this brief to reduce latency and complexity. Based on the fact that the codeword of an ECC is usually represented in a systematic form consisting of the raw data and the parity information generated by encoding, the proposed architecture parallelizes the comparison of the data and that of the parity information. To further reduce the latency and complexity, in addition, a new butterfly-formed weight accumulator (BWA) is proposed for the efficient computation of the Hamming distance. Grounded on the BWA, the proposed architecture examines whether the incoming data matches the stored data if a certain number of erroneous bits are corrected. For a (40, 33) code, the proposed architecture reduces the latency and the hardware complexity by ~32% and 9%, respectively, compared with the most recent implementation.
Byeong Yong Kong, Jihyuck Jo, Hyewon Jeong, Mina Hwang, Soyoung Cha, Bongjin Kim, In-Cheol Park
IEEE Trans. Very Large Scale Integr. Syst.1
2013 Adaptive Metric Calculation for Improving Detection Capability of MIMO Detectors
abstract
A simple yet effective scheme is proposed to improve the detection capability of multiple-input multiple-output (MIMO) symbol detectors in wireless communication systems at low signal-to-noise ratio (SNR). The proposed scheme is to change the metric calculation adaptively to the channel SNR, grounded on the fundamental relationship between the detection capability and the bit-error rate (BER) performance with respect to the channel SNR. An efficient hardware architecture implementing the proposed scheme is also presented to show its applicability in a practical sense and it is proven that the architecture induces only a negligible hardware overhead compared to the state-of-the-art MIMO symbol detector. Experimental results show that the proposed scheme can indeed increase the detection capability effectively without degrading the BER performance noticeably.
Byeong Yong Kong, In-Cheol Park
VTC Spring1
2012 FIR Filter Synthesis Based on Interleaved Processing of Coefficient Generation and Multiplier-Block Synthesis
abstract
An efficient filter synthesis algorithm is proposed to minimize the number of adders required in the design of finite-impulse response filters. Given a specification, a filter can be synthesized by conducting two main steps: coefficient generation and multiplier-block synthesis. While most of previous works have focused on only one of the steps, the proposed algorithm integrates the two steps in an interleaved manner so as to take into account the effect of multiplier-block synthesis in generating coefficients. In addition, the concept of sensitivity is developed to reduce the complexity of computing the variable ranges of coefficients. Experimental results show that the proposed algorithm outperforms previous algorithms in terms of adder cost and takes a relatively short computation time.
Byeong Yong Kong, In-Cheol Park
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1