EDBT 2026 Demo / reviewers in the wild / expert
Byeong Yong Kong
dblp:119/0553
· DBLP profile ↗
18ranked-venue papers
12as first author
8since 2021 · last 2026
0000-0001-5823-5505ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 8 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fault-Tolerant IDMA Multiuser Detector Based on Fault Injection Analysis of Internal MemoriesabstractIn this article, a fault-tolerant architecture (FTA) is presented for the multiuser detector in interleave division multiple access (IDMA). The detector is inherently prone to soft errors, as its chip area is predominantly occupied by memories, which are easily exposed to high-energy particles and hostile interferences in harsh environments. One of the most widespread ways to protect memories is to encode their entries with error-correcting codes (ECCs). However, naïvely encoding all bits in an entry is likely to be costly and unnecessary. Accordingly, to sort out performance-critical bits and determine the priority of protection, we extensively scrutinize how vulnerable respective bits in the memories of the detector are too soft errors. Based on the analysis, in addition, an efficient FTA that selectively encodes only a subset of the bits in order of the identified vulnerability is developed. Furthermore, the proposed FTA implements the state-of-the-art multiuser detection (MUD) scheme called on-the-fly despreading (OD) and showcases a new feature named purification, which repeatedly replaces erroneous entries with corrected ones to keep them error-free. Complicated memory accesses to concurrently perform the OD as well as the purification are enabled by remodeling both the datapath and the control path of the baseline OD architecture (ODA). Implementation results demonstrate that, unlike the prior arts that fail to sustain near-optimal performances and become impractical even for a very low probability of soft error, the proposed FTA may operate robustly in a wide range of harsh conditions without incurring much overhead. Byeong Yong Kong |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | Generalized Shifting-Based PAPR Reduction Architecture for Discrete Multitone SystemsabstractIn this paper, a generalized shifting-based peak-to-average power ratio (PAPR) reduction architecture is presented for discrete multitone (DMT) systems. The limitation of the state-of-the-art scheme is that it resorts only to the local extrema from subsets in calculating the compensation offset, being inherently prone to accuracy degradation. To overcome the drawback, the proposed generalized offset enables new design options that utilize multiple global extrema, mitigating the PAPR further. In addition to the generalization, a low-complexity bilateral sorter, which is the key component of the entire architecture to identify multiple extrema, is developed by simplifying and hybridizing mergesorters. Since the generalized offset not only offers the new options but also encompasses all the previous expressions of the offsets in the literature, it successfully unifies all dc shifting-based PAPR reduction schemes, and exempts us from comparing them and choosing one. Byeong Yong Kong |
ISCAS | 1 |
| 2024 | Constrained Sorter Design using Zero-One PrincipleabstractTo derive efficient sorting architectures constrained to application-specific input/output conditions, we present in this paper a systematic design methodology that can effectively prune dispensable compare-and-swap (CAS) units. Unlike the previous works resorting to heuristic approaches, the proposed framework exploits the zero-one principle to validate the pruning of a CAS unit at a time, generating the cost-optimized sorter architecture in an iterative manner with a reasonable complexity. In addition to the given input/output constraints, we newly develop the architecture options for the proposed framework, allowing more design spaces for finding the most attractive constrained-sorter design. For 8-list polar decoders, the proposed framework successfully reduces 70% of CAS units in the baseline full sorter, relaxing the area-time complexity by 35% compared with the state-of-the-art solutions. Sangil Han, Jaehee Kim, Dongyun Kam, Byeong Yong Kong, Mijung Kim, Young-Seok Kim, Youngjoo Lee 0002 |
ISCAS | 4 |
| 2024 | A Design Framework for Cost-Efficient Sorters With Arbitrary Input/Output ConstraintsabstractThe sorting operation plays a vital role in various signal processing applications. However, due to high hardware complexity resulting from a series of comparisons, designing the cost-efficient sorter is one of the crucial requisites for improving the overall system performance. To obtain the cost-efficient sorting architectures constrained to application-specific input/output conditions, this paper presents a systematic design methodology that effectively eliminates dispensable compare-and-swap (CAS) units. Unlike the previous heuristic approaches, the proposed framework iteratively prunes a CAS unit followed by the validation step. The zero-one principle is newly applied to reduce the validation time for the practical convergence time with massive searching iterations. To expand the search space of the proposed framework, furthermore, we introduce new architectural options and pruning methods, allowing the cost-efficient design results even compared to the state-of-the-art solutions. Targeting the constrained sorters for communication systems, numerous case studies show that the proposed framework successfully removes more than half of CAS units in the baseline sorter design, significantly relaxing the area-time complexity, e.g., 35% reduction compared to the state-of-the-art architecture for 16-input metric sorter in the SCL polar decoder. Jaehee Kim, Sangil Han, Dongyun Kam, Byeong Yong Kong, Youngjoo Lee 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | Low-Latency PAPR Reduction Architecture for Discrete Multitone Based on Approximate MidrangeabstractIn this brief, a low-latency hardware architecture is presented for the peak-to-average power ratio (PAPR) reduction in discrete multitone (DMT) systems. The state-of-the-art scheme suffers from high latency due to a series of selections in calculating the midrange of the DMT frame. To overcome the drawback, an approximate calculation of the midrange that transforms the dominant operation from the selections to the summation is proposed. Grounded on the transformation, in addition, a low-latency architecture to calculate the approximate midrange is constituted. While the selections involve negation, carry propagation, and multiplexing in every stage, the summation can be promptly done by carry-save adders (CSAs) without such manipulations. Slightly relaxing the accuracy of the midrange, as a result, the overall latency can be effectively alleviated compared with the state-of-the-art counterpart. Byeong Yong Kong |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | Low-Latency SCL Polar Decoder Architecture Using Overlapped Pruning OperationsabstractAllowing the superior error-correction performance even for short-length codewords, the successive-cancellation list (SCL) decoding algorithm has allowed the polar code to be adopted in 5G New Radio standard for control channel. However, existing SCL polar decoders still suffer from long processing latency caused by a number of serialized internal operations. In this work, to solve the latency problem, we present several parallel computing solutions for the serialized operations, i.e., simplified data dependencies and two overlapped pruning operations. To realize the proposed parallel computing, we also introduce internal circuit blocks including dual read-port buffers, an on-the-fly parity checker, and overlapped processing units. The proposed SCL polar decoders are precisely designed with optimal design parameters by analyzing trade-offs between the latency reduction and area overheads. Implemented in a 65-nm CMOS technology, the proposed list-8 SCL polar decoder requires only 374 ns to handle a (1024, 512) 5G codeword, improving the decoding efficiency by 34.7% compared to the previous designs. Dongyun Kam, Byeong Yong Kong, Youngjoo Lee 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | Low-Latency Polar Decoder Using Overlapped SCL ProcessingabstractIn this paper, we present a novel scheduling method that reduces the latency of polar decoders significantly. Unlike the prior pruning-based successive cancellation list (SCL) decoding that suffers from a number of idle cycles, the proposed overlapped SCL scheme immediately begins node operations without waiting for the list to be sorted, being exempt from such unfavorable cycles. All possible candidates for the next node operations are precomputed in parallel with the pruning operations, and are readily selected to minimize the latency. For the 5G New Radio systems, the proposed method shortens the decoding latency of the state-of-the-art approaches by up to 22% without degrading the error-correcting performance. Dongyun Kam, Byeong Yong Kong, Youngjoo Lee 0002 |
ICASSP | 2 |
| 2021 | Real-Time SSDLite Object Detection on FPGAabstractDeep neural network (DNN)-based object detection has been investigated and applied to various real-time applications. However, it is hard to employ the DNNs in embedded systems due to their high computational complexity and deep-layered structure. Although several field-programmable gate array (FPGA) implementations have been presented recently for real-time object detection, they suffer from either low throughput or low detection accuracy. In this article, we propose an efficient computing system for real-time SSDLite object detection on FPGA devices, which includes novel hardware architecture and system optimization techniques. In the proposed hardware architecture, a neural processing unit (NPU) that consists of heterogeneous units, such as band processing, scaling, and accumulating, and data fetching and formatting units is designed to accelerate the DNNs efficiently. In addition, system optimization techniques are presented to improve the throughput further. A task control unit is employed to balance the workload and increase the utilization of heterogeneous units in the NPU, and the object detection algorithm is refined accordingly. The proposed architecture is realized on an Intel Arria 10 FPGA and enhances the throughput by up to 13.6× compared to the state-of-the-art FPGA implementation. Suchang Kim, Seungho Na, Byeong Yong Kong, Jaewoong Choi, In-Cheol Park |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | Improved Parallel-IDMA Architecture with Low-Complexity Elementary Signal EstimatorsabstractThis paper presents a low-complexity parallel multiuser detector architecture for interleave division multiple access systems. To facilitate efficient hardware implementation, the formulae associated with the elementary signal estimator (ESE) are reorganized. Grounded on the reorganization, the proposed ESE circumvents redundant computations, and takes advantage of carry-save additions. The resulting datapath is further simplified by approximations that do not deteriorate the error rate noticeably. Prior to integrating with such ESEs, the state-of-the-art parallel architecture for the user-specific processing block (UPB) is also simplified by rescheduling the memory-access pattern. As a result, a prototype 2-parallel 16-user detector that incorporates the proposed ESEs and UPBs in a 65-nm CMOS occupies 23% less silicon area, dissipates 21% less power, and takes 50% less latency than the state-of-the-art serial detector. Byeong Yong Kong |
ISCAS | 1 |
| 2020 | A Low-Latency Multi-Touch Detector Based on Concurrent Processing of Redesigned Overlap Split and Connected Component AnalysisabstractA low-latency multi-touch detector architecture for locating numerous touches in large-panel devices is presented in this paper. Two respective processors for the overlap split and the connected component analysis (CCA) are the key components in a multi-touch detector. Exploiting the simplicity of typical intensity maps in practice, the two processors are first redesigned under a principle that every pixel is processed only once in a raster-scan order. More specifically, the concept of valley point is introduced, and a simple yet effective overlap-split scheme called valley-point division (VPD) is newly developed based on the concept. In addition, equivalent labels in the CCA are handled on the fly instead of being kept to be processed later, and the logics and memories for corner cases that never occur in practice are unloaded. Enabled by the redesign, subsequently, the two processors are integrated into one to concurrently conduct both the VPD and the CCA during the same raster scan. As a result, the proposed detector completes the whole detection procedure with a single raster scan of a map, and is exempted from a large memory required to hold an entire map. Implementation results in a 65-nm CMOS for a large panel of 400 × 250 sensors demonstrate that the proposed detector takes less than a half latency of the existing ones while occupying only 17% silicon area and consuming 52% power on average. Byeong Yong Kong, Jooseung Lee, In-Cheol Park |
ISCAS | 1 |
| 2020 | A 120-mW 0.16-ms-Latency Connectivity-Scalable Multiuser Detector for Interleave Division Multiple AccessabstractTo facilitate the massive connectivity of the 5G New Radio, this brief proposes a connectivity-scalable multiuser detector (MUD) architecture for interleave division multiple access (IDMA). A regular inter-MUD interface makes the proposed MUD not only operable alone but also scalable to serve a massive number of users by connecting multiple MUDs. Besides, the memory subsystem is simplified to mitigate the power consumption and the silicon area. A prototype 16-MUD in a 65-nm CMOS consumes 120 mW, occupies 8.21 mm2, and takes 0.16 ms to process one 8192-chip frame. Such numbers represent 29% less silicon area and 59% less power dissipation than those of the state-of-the-art MUD. Byeong Yong Kong, In-Cheol Park |
ISCAS | 1 |
| 2020 | Ultra-Low-Latency LDPC Decoding Architecture using Reweighted Offset Min-Sum AlgorithmabstractDue to an iterative nature, a low-density parity-check (LDPC) decoder is associated with a long latency, being a major bottleneck of the baseband processor in wireless communication systems. Based on the practical min-sum (MS) decoding method, in this paper, we present a cost-effective algorithm for reducing the processing latency of LDPC decoders. By checking the number of short-length cycles in the LDPC code structure, the proposed method dynamically changes the reweighting factor at the iterative operations, successfully reducing the average number of iterations. In addition, we present several optimization schemes to mitigate the hardware overheads resulting from the proposed reweighting scheme. In a 65-nm CMOS process, a prototype IEEE 802.11ay LDPC decoder optimized by the proposed schemes reduces the decoding latency by 1.7 times with negligible overheads compared with the contemporary designs. Sangbu Yun, Dongyun Kam, Jeongwon Choe, Byeong Yong Kong, Youngjoo Lee 0002 |
ISCAS | 4 |
| 2019 | Parallel IDMA Architecture Based on Interleaving with Replicated SubpatternsabstractThis paper presents a parallel multiuser detector architecture for low-latency interleave division multiple access. To enable P-parallel processing, an interleaving pattern is divided into P disjoint subpatterns, and all the subpatterns are designed to be identical without degrading error-rate performance noticeably. Since the subpatterns are all disjoint, they can be processed in parallel. Besides, by exploiting that they access the same address of separate memory banks at the same time, the banks are integrated into one to minimize the silicon area and the power consumption. As a result, the proposed architecture reduces the latency by a factor of P at the expense of a little hardware overhead. A prototype 2-parallel 16-user detector in a 65-nm CMOS completes the entire detection procedure two times earlier than the state-of-the-art nonparallel detector, while occupying only 12% more silicon area and dissipating 20% more power. Byeong Yong Kong, In-Cheol Park |
ICC | 1 |
| 2016 | Low-complexity symbol detection for massive MIMO uplink based on Jacobi methodabstractIn this paper, we propose a low-complexity symbol detection algorithm for massive multiple-input multiple-output (MIMO) uplink. Grounded on the fact that a primary property of the massive MIMO systems guarantees the convergence of the Jacobi method, the method is exploited in the linear detection so that the estimate of transmitted symbols can be obtained without employing the computationally intensive matrix inversion. In addition, we propose a multiplication-free initial estimate for the Jacobi method in order to lessen the computational complexity further. Owing to the elimination of matrix inversion and the efficient initial estimate, the proposed algorithm achieves near-optimal error-rate performance with fewer computations than the state-of-the-art schemes. Byeong Yong Kong, In-Cheol Park |
PIMRC | 1 |
| 2015 | Narrow-range frequency estimation based on comprehensive optimization of DFT and interpolationabstractAn efficient procedure for frequency estimation is proposed in this paper to alleviate the computational complexity. Grounded on the fact that the frequency of a target signal usually lies in a known range in practical applications, two fundamental steps in the frequency estimation, i.e., the discrete Fourier transform (DFT) and the interpolation of the DFT samples, are modified accordingly. Unlike the previous works focusing on either the DFT or the interpolation, this paper does not decouple the two steps but optimizes the whole procedure comprehensively by considering the interrelationship between the two steps. As a result, the number of operations required for the estimation is remarkably diminished while the performance remains competitive with the recent works. Byeong Yong Kong, In-Cheol Park |
ICASSP | 1 |
| 2014 | Low-Complexity Low-Latency Architecture for Matching of Data Encoded With Hard Systematic Error-Correcting CodesabstractA new architecture for matching the data protected with an error-correcting code (ECC) is presented in this brief to reduce latency and complexity. Based on the fact that the codeword of an ECC is usually represented in a systematic form consisting of the raw data and the parity information generated by encoding, the proposed architecture parallelizes the comparison of the data and that of the parity information. To further reduce the latency and complexity, in addition, a new butterfly-formed weight accumulator (BWA) is proposed for the efficient computation of the Hamming distance. Grounded on the BWA, the proposed architecture examines whether the incoming data matches the stored data if a certain number of erroneous bits are corrected. For a (40, 33) code, the proposed architecture reduces the latency and the hardware complexity by ~32% and 9%, respectively, compared with the most recent implementation. Byeong Yong Kong, Jihyuck Jo, Hyewon Jeong, Mina Hwang, Soyoung Cha, Bongjin Kim, In-Cheol Park |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | Adaptive Metric Calculation for Improving Detection Capability of MIMO DetectorsabstractA simple yet effective scheme is proposed to improve the detection capability of multiple-input multiple-output (MIMO) symbol detectors in wireless communication systems at low signal-to-noise ratio (SNR). The proposed scheme is to change the metric calculation adaptively to the channel SNR, grounded on the fundamental relationship between the detection capability and the bit-error rate (BER) performance with respect to the channel SNR. An efficient hardware architecture implementing the proposed scheme is also presented to show its applicability in a practical sense and it is proven that the architecture induces only a negligible hardware overhead compared to the state-of-the-art MIMO symbol detector. Experimental results show that the proposed scheme can indeed increase the detection capability effectively without degrading the BER performance noticeably. Byeong Yong Kong, In-Cheol Park |
VTC Spring | 1 |
| 2012 | FIR Filter Synthesis Based on Interleaved Processing of Coefficient Generation and Multiplier-Block SynthesisabstractAn efficient filter synthesis algorithm is proposed to minimize the number of adders required in the design of finite-impulse response filters. Given a specification, a filter can be synthesized by conducting two main steps: coefficient generation and multiplier-block synthesis. While most of previous works have focused on only one of the steps, the proposed algorithm integrates the two steps in an interleaved manner so as to take into account the effect of multiplier-block synthesis in generating coefficients. In addition, the concept of sensitivity is developed to reduce the complexity of computing the variable ranges of coefficients. Experimental results show that the proposed algorithm outperforms previous algorithms in terms of adder cost and takes a relatively short computation time. Byeong Yong Kong, In-Cheol Park |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |