VLDB 2026 Research / reviewers in the wild / expert
Sunwoong Kim
dblp:60/9847
· DBLP profile ↗
19ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-1471-0228ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | StreamNTT: A High-Throughput, HLS-Based Streaming NTT Accelerator for HBM-Equipped FPGAsabstractLattice-based post-quantum cryptography (PQC) relies extensively on the number-theoretic transform (NTT), and large-scale deployments require throughput-optimized accelerators capable of computing thousands of NTTs per session. To address this need, we present StreamNTT, a high-throughput NTT accelerator developed using high-level synthesis (HLS) for field-programmable gate arrays (FPGAs) equipped with high bandwidth memory (HBM). StreamNTT effectively leverages the parallelism inherent in the NTT by overcoming several obstacles that limit the scalability of streaming dataflow designs. Specifically, we introduce HLS-friendly butterfly units that integrate butterfly computation and reorder buffering into a single pipelined loop. In addition, we propose a butterfly merging strategy that reduces first-in, first-out channels and buffers. Finally, we present a placement-aware instance-level parallelism scheme for FPGA platforms with multiple HBM channels and super logic regions. On the Alveo U280 platform, the proposed optimizations improve throughput by a factor of 7.2 and achieve a speedup of more than 2.7 times compared to state-of-the-art FPGA-based NTT accelerators. As an open-source design, StreamNTT establishes a scalable foundation for advancing high-performance PQC acceleration. Young-Kyu Choi, Sunwoong Kim |
DATE | 4 |
| 2026 | SPEED: Structured kernel block pruning with filter groups for efficient and elastic SW-HW co-design in FPGA-based CNN accelerators
Kwanghyun Koo, Sunwoong Kim, Hyun Kim 0001 |
Neurocomputing | 2 |
| 2024 | Accelerating Homomorphic Comparison Operations for Thresholding Using an Asymmetric Input Range and Input ScalingabstractIn a cyber-physical system (CPS), the interconnection of cyber and physical components occurs through a network. This structure, particularly cyber components and networks, makes it susceptible to malicious attacks. One of the solutions to this CPS security issue is to employ end-to-end homomorphic encryption (HE) that allows direct computations on encrypted data. Despite its promise, HE only supports basic operations, such as addition and multiplication, which limits its application areas. Numerical methods have been presented to perform a comparison operation in the HE domain. However, they suffer from a slow processing speed due to an inherently high number of iterations. To accelerate a homomorphic comparison operation, this paper introduces a novel approach that scales inputs using an asymmetric input range in thresholding. Additionally, parallelism in HE-based multilevel thresholding is explored and exploited through the use of a parallel processing application programming interface for further acceleration. Compared to a previous comparison operation method, the proposed method achieves comparable accuracy with fewer iterations, resulting in a 48% reduction in execution time on an edge computing device. Furthermore, employing an additional thread using parallelism increases this reduction to 63%. Sunwoong Kim, Wonhee Cho 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2024 | Area-Efficient Iterative Logarithmic Approximate Multipliers for IEEE 754 and Posit NumbersabstractThe IEEE 754 standard for floating-point (FP) arithmetic is widely used for real numbers. Recently, a variant called posit was proposed to improve the precision around 1 and −1. Since FP multiplication requires high computational complexity, various algorithmic approaches and hardware accelerator solutions have been explored. In this context, this article proposes a novel area-efficient logarithmic multiplier architecture for different real number formats, which also provides a significant and useful accuracy/latency tradeoff at runtime. To reduce the logic area in field-programmable gate arrays (FPGAs), this article offers two innovations: applying logarithm to only a single operand and mitigating the accuracy drop caused by this modification with advanced error converging and operand selection schemes. Our multiplier design for single-precision FP (SPFP) numbers uses 58% fewer hardware resources than the iterative Mitchell’s multiplier (IMM) design of Babić et al. extended for SPFP numbers. The error falls within 0.5% when the number of iterations reaches 5. In JPEG, our SPFP multiplier with four iterations produces nearly identical image quality results to the conventional exact multiplier. We further show how to merge two SPFP multipliers for double-precision FP (DPFP) multiplication. This DPFP multiplier design reduces the hardware resources of the IMM design extended for DPFP numbers by 60%. Finally, we demonstrate how our SPFP multiplier design can be slightly modified for 32-bit posit multiplication. It achieves a significantly higher accuracy by increasing the number of iterations compared to state-of-the-art approximate posit multiplier designs. Sunwoong Kim, Cameron James Norris, James Oelund, Rob A. Rutenbar |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | ILAFD: Accuracy-Configurable Floating-Point Divider Using an Approximate Reciprocal and an Iterative Logarithmic MultiplierabstractApproximate computing provides benefits in the logic area, latency, and/or power without significantly affecting the outputs of some applications. Various approximate computing techniques have been applied for floating-point (FP) dividers that are expensive compared to other FP operators. Since the accuracy requirement varies depending on the application, this paper proposes an accuracy-configurable approximate FP divider design. This design computes the reciprocal of a divisor in a hardware-friendly way. To correct approximation errors, multiple error biases are calculated based on error analysis and stored in a lookup table. The calculated reciprocal is multiplied by a dividend. For accuracy-configurability, an iterative logarithmic FP multiplier is adopted. Compared to previous approximate designs, our proposed design achieves a significantly higher level of accuracy by increasing the number of iterations, at the cost of a slight increase in hardware resources. Compared to the state-of-art Xilinx LogiCORE IP design for accurate FP division, the proposed approximate design reduces the number of lookup table blocks by 53% and the number of flip-flops by 90% in a field-programmable gate array. In JPEG compression, the conventional exact FP multiplier/divider and the proposed design with the number of iterations of 4 show almost the same peak signal-to-noise ratio values. James Oelund, Sunwoong Kim |
ACM Great Lakes Symposium on VLSI | 2 |
| 2023 | A Use Case of Iterative Logarithmic Floating-Point Multipliers: Accelerating Histogram Stretching on Programmable SoCabstractProgrammable system-on-chip (SoC) platforms often have a hard-core processor and programmable logic for heterogeneous computing. A custom hardware design implemented with programmable logic accelerates specific operations in an application. This paper applies an area-efficient iterative logarithmic floating-point multiplier to histogram stretching on a programmable SoC platform. Specifically, the floating-point multiplier is implemented with programmable logic, and the hard-core processor on the same platform determines the accuracy and latency of the multiplier by changing the number of iterations used for error correction at runtime. To deploy many multiplier cores working in parallel and have the iterative architecture process streaming data, a higher-level pipelined architecture is proposed. The number of multipliers in this architecture is configurable by changing the maximum number of iterations. For$352\times 288$resolution videos, with 250 frames per video, the proposed histogram stretching design reduces the multiplication time by 78% and the total execution time by 13% compared to a design performing exact floating-point multiplications on a hard-core processor. Cameron James Norris, Sunwoong Kim |
ISCAS | 2 |
| 2023 | HEBGS: Homomorphic Encryption-based Background Subtraction Using a Fast-Converging Numerical MethodabstractRecent advances in cloud services provide greater computing ability to edge devices on cyber-physical systems (CPS) and internet of things (IoT) but cause security issues in cloud servers and networks. This paper applies homomorphic encryption (HE) to background subtraction (BGS) in CPS/IoT. Cheon et al. 's numerical methods are adopted to implement the non-linear functions of BGS in the HE domain. In particular, square- and square root-based HE-based BGS (HEBGS) designs are proposed for the input condition of the numerical comparison operation. In addition, a fast-converging method is proposed so that the numerical comparison operation outputs more accurate results with lower iterations. Although the outer loop of the numerical comparison operation is removed, the proposed square-based HEBGS with the fast-converging method shows an average peak signal-to-noise ratio value of 20dB and an average structural similarity index measure value of 0.89 compared to the non-HE-based conventional BGS. On a PC, the execution time of the proposed design for each$128\times 128$-sized frame is 0.34 seconds. Justin Shyi, Sunwoong Kim |
ISCAS | 2 |
| 2022 | HELPSE: Homomorphic Encryption-based Lightweight Password Strength Estimation in a Virtual Keyboard SystemabstractRecently, cyber-physical systems are actively using cloud servers to overcome the limitations of power and processing speed of edge devices. When passwords generated on a client device are evaluated on a server, the information is exposed not only on networks but also on the server-side. To solve this problem, we move the previous lightweight password strength estimation (LPSE) algorithm to a homomorphic encryption (HE) domain. Our proposed method adopts numerical methods to perform the operations of the LPSE algorithm, which is not provided in HE schemes. In addition, the LPSE algorithm is modified to increase the number of iterations of the numerical methods given depth constraints. Our proposed HE-based LPSE (HELPSE) method is implemented as a client-server model. As a client-side, a virtual keyboard system is implemented on an embedded development board with a camera sensor. A password is obtained from this system, encrypted, and sent over a network to a resource-rich server-side. The proposed HELPSE method is performed on the server. Using depths of about 20, our proposed method shows average error rates of less than 1% compared to the original LPSE algorithm. For a polynomial degree of 32K, the execution time on the server-side is about 5 seconds. Michael Cho, Keewoo Lee, Sunwoong Kim |
ACM Great Lakes Symposium on VLSI | 3 |
| 2022 | HEKWS: Privacy-Preserving Convolutional Neural Network-based Keyword Spotting with a Ciphertext Packing TechniqueabstractKeyword spotting (KWS) is a key technology in smart devices. However, privacy issues in these devices have been constantly raised. To solve this problem, this paper applies homomorphic encryption (HE) to a previous small-footprint convolutional neural network (CNN)-based KWS algorithm. This allows for a trustless system in which a command word can be securely identified by a remote cloud server without exposing client data. To alleviate the burden on an edge device of a client, a novel packing technique is proposed that reduces the number of ciphertexts for an input keyword to one. Our HE-based KWS shows a prediction accuracy of 72% for Google's Speech Commands Dataset with 12 labels. This is almost identical to the accuracy of the non-HE-based implementation that has the same CNN layers and approximates a rectified linear unit in the same manner. On a workstation, it takes 19 seconds to process one keyword on average, which can be improved in the future through parallelization, HE parameter optimization, and/or the use of custom hardware accelerators. Daniel L. Elworth, Sunwoong Kim |
MMSP | 2 |
| 2022 | HEMTH: Small Depth Multilevel Thresholding for a Homomorphically Encrypted ImageabstractOne of the image segmentation techniques, multilevel thresholding, is widely used in many computer vision applications because of its low computational complexity and efficient data representation. When it is used in cyber-physical systems and internet-of-things, a special technique is required to protect the sensitive information in an image. This paper proposes a novel homomorphic encryption (HE)-based multilevel thresholding method. To implement a comparison operation in the HE domain, which is not a basic homomorphic operation, a numerical method is adopted. Our proposed method executes comparison operations in parallel to perform more iterations and increase accuracy. When the number of iterations in the numerical comparison operation is (5, 3), the proposed three-level thresholding method shows an average peak signal-to-noise ratio of 28 dB compared to a conventional non-HE-based method and takes 3 minutes on a PC. Paul Nam, Justin Shyi, Sunwoong Kim |
MMSP | 3 |
| 2022 | Virtual Keyboards With Real-Time and Robust Deep Learning-Based Gesture RecognitionabstractIn head-mounted display devices for augmented reality and virtual reality, external signals are often entered using a virtual keyboard (VKB). Among various user interfaces for VKBs, hand gestures are widely used because they are fast and intuitive. This work proposes a gesture-recognition (GR)-based VKB algorithm that is accurate in any environment and operates in real time. Specifically, the proposed ambidextrous VKB layouts reduce the total finger travel distance on one-hand VKB layouts. Additionally, a fast typing action is proposed to use characteristics when previous and current keys are adjacent. To be robust in any environment, we utilize a deep learning (DL)-based GR method in the proposed VKB algorithm. To train DL networks, seven classes are defined and an automated dataset generation method is proposed to reduce the necessary time and effort. The proposed one-hand VKB layout with the fast typing action shows a 1.5× faster typing speed than the popular ABC keyboard layout. Furthermore, the proposed ambidextrous VKB layout brings an additional 52% improvement compared with the proposed one-hand VKB layout. The proposed DL-based GR method implemented on the well-known YOLOv3 machine learning framework shows a mean average precision rate of 95% for images including background colors similar to skin color. The proposed DL-based GR method for one-hand and ambidextrous VKBs achieves around 41 frames per second on a software platform, which allows real-time processing. Sunwoong Kim |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2021 | An Approximate and Iterative Posit Multiplier Architecture for FPGAsabstractRecently, many applications have demanded cheaper and faster arithmetic while providing a wider dynamic range than the popular IEEE 754 floating-point (FP) arithmetic. As a result, a new number format called posit was proposed. As in a variety of number systems, multiplication in posit arithmetic is one of the most frequently used but expensive operations. To reduce the number of hardware resources in a posit multiplier, this paper applies an iterative approach to posit multiplication. To exploit the features of the posit format, the number of truncated bits in the fraction component is dynamically changed depending on the number of regime bits. In addition, architectures for fast parser and packer are proposed. Thanks to the posit format, the proposed design supports about 60 decades wider dynamic range than a single-precision FP multiplier design. Using the iterative approach, the proposed design also reduces the number of lookup tables used in a previous exact posit multiplier design by 44% while achieving a 51% higher maximum frequency, and the appropriate balance between latency and accuracy can be finely determined at runtime. To the best of the authors' knowledge, this paper is the first work that proposes a hardware architecture of an approximate and iterative posit multiplier. Cameron James Norris, Sunwoong Kim |
ISCAS | 2 |
| 2021 | A Homomorphic Encryption-based Adaptive Image Filter Using Division Over Encrypted DataabstractHomomorphic encryption (HE) is an important cryptographic technique that allows one to directly perform computation on encrypted data without decryption. In HE-based applications using digital images, a user often encrypts a private image captured on a local device. This image can contain noise that negatively affects the results of HE-based applications. To solve this problem, this paper proposes an HE-based adaptive image filter. For small-sized encrypted input data, pixels that have no dependency when sliding a window are encoded into the same ciphertext. For division in the adaptive filter, which is not supported by conventional HE schemes, a numerical approach is adopted. To the best of the authors’ knowledge, this paper is the first work that applies division over encrypted data to an image processing algorithm. We implemented the proposed HE-based adaptive filter as a proof-of-concept client-server model. The proposed design can address important privacy issues in image processing applications in internet-of-things and cyber-physical systems, where many devices are connected through a vulnerable network. Sharmila Devi Kannivelu, Sunwoong Kim |
RTCSA | 2 |
| 2020 | Hardware Architecture of a Number Theoretic Transform for a Bootstrappable RNS-based Homomorphic Encryption SchemeabstractHomomorphic encryption (HE) is one of the most promising solutions to secure cloud computing. The number theoretic transform (NTT) that is widely used for convolution operations in HE requires a large amount of computation and has high parallelism, and therefore it has been a good candidate for hardware acceleration. Nevertheless, prior NTT hardware solutions for HE-based applications are impractical in most applications because they do not seriously consider the critical bootstrapping procedure that allows unlimited homomorphic operations on encrypted data. In this paper, we suggest practical bootstrappable parameters, specifically for an established residue number system (RNS)based HE scheme, and apply them to our NTT hardware design. In addition, to limit the size of internal memory for roots of unity increased by the bootstrappable parameters, only a few roots of unity are stored and others are generated on the fly. In our NTT hardware architecture, multiple NTT butterfly units (BUs) are efficiently deployed for high throughput and high resource utilization. In particular, several groups of BUs for respective moduli work in a parallel and pipelined manner, which is effective in an RNS-based HE scheme with a number of moduli. Our implementation on a Xilinx UltraScale FPGA with the bootstrappable parameters achieves a $118 \times$ faster processing speed than a software implementation, and it further provides various trade-off choices such as the number of DSP slices against BRAMs based on available FPGA resources. Sunwoong Kim, Keewoo Lee, Wonhee Cho 0001, Yujin Nam, Jung Hee Cheon, Rob A. Rutenbar |
FCCM | 1 |
| 2019 | An Area-Efficient Iterative Single-Precision Floating-Point Multiplier Architecture for FPGAabstractApproximate multipliers have been widely used in critical applications, such as machine learning and multimedia, which are tolerant to approximation errors. This paper proposes a novel single-precision floating-point (SPFP) multiplication algorithm and its architecture. The proposed work approximates only one of the operands to reduce the number of logic blocks and iteratively compensates the approximation error to achieve acceptable error ranges in applications. To reduce the accuracy degradation by the single operand approximation, a rounding scheme and an operand selection scheme are additionally introduced. Compared with the widely-known previous iterative Mitchell design, our proposed SPFP multiplier design decreases the numbers of look up tables (LUTs) and flip flops (FFs) by 55% and 59% respectively, and shows two cycles shorter latency. The accuracy of our design becomes close to that of the iterative Mitchell design as the number of iterations increases, and it always meets the error tolerance of 1% when the number of iterations is four. Sunwoong Kim, Rob A. Rutenbar |
ACM Great Lakes Symposium on VLSI | 1 |
| 2018 | Accelerator Design with Effective Resource Utilization for Binary Convolutional Neural Networks on an FPGAabstractIn binary convolutional neural networks (BCNN), arithmetic operations are replaced by bitwise operations and the required memory size is greatly reduced, which is a good opportunity to accelerate training or inference on FPGAs. This paper proposes a BCNN architecture with a single engine that achieves high resource utilization. The proposed design deploys a large number of processing elements in parallel to increase throughput, and a forwarding scheme to increase resource utilization on the existing engine. In addition, we demonstrate a novel reuse scheme to make fully-connected layers exploit the same engine. The proposed design is combined with an inference environment for comparison and implemented on a Xilinx XCVU190 FPGA. The implemented design uses 61k look-up tables (LUTs), 45k flip-flops (FFs), and 13.9Mbit block RAM (BRAM). In addition, it achieves 61.6 GOPS/kLUT at 240MHz, which is 1.16 times higher than that of the best prior BCNN design, even though it uses a single engine without optimal configurations on each layer. Sunwoong Kim, Rob A. Rutenbar |
FCCM | 1 |
| 2016 | Fixed-length Golomb-Rice coding by quantization level estimationabstractGolomb-Rice coding is one of the popular variable-length codings which require quantization of input data to meet the target compression ratio. In order to obtain the optimal quantization level, a conventional iterative approach increases the quantization level one by one until the target ratio is achieved. This iterative approach makes it difficult to implement in hard ware because the number of iterations cannot be estimated at hardware design time. This paper proposes a non-iterative algorithm for Golomb-Rice coding to estimate a near-optimal quantization level. To this end, the algorithm performs Golomb-Rice coding without any quantization of input data and then uses this coding result to estimate the codeword length with quantization. Based on the estimation, a near-optimal quantization level that meets the target length is selected. For the case when the selected level is not optimal, the algorithm performs additional Golomb-Rice codings with modified quantization levels which guarantee the codeword to meet the target length. Experimental results with twenty-four Kodak images show that the proposed coders practically cover all the optimal quantization levels. Moonsoo Kim, Sunwoong Kim |
ISCAS | 2 |
| 2016 | A High-Throughput Hardware Design of a One-Dimensional SPIHT AlgorithmabstractVideo display systems include frame memory, which stores video data for display. To reduce system cost, video data are often compressed for storage in frame memory. A desirable characteristic for display memory compression is support for the raster-scan processing order and the fixed target compression ratio. Set partitioning in hierarchical trees (SPIHT) is an efficient two-dimensional compression algorithm that guarantees a fixed target compression ratio, but its one-dimensional (1D) variation has received little attention, even though its 1D nature supports the raster-scan processing order. This paper proposes a novel hardware design for 1D SPIHT. The algorithm is modified to exploit parallelism for effective hardware implementation. For the encoder, dependences that prohibit parallel execution are resolved and a pipelined schedule is proposed. For the parallel execution of the decoder, the algorithm is modified to enable estimation of the bitstream length of each pass prior to decoding. This modification allows parallel and pipelined decoding operations, leading to a high-throughput design for both encoder and decoder. Although the modifications slightly decrease compression efficiency, additional optimizations are proposed to improve such efficiency. As a result, the peak signal-to-noise ratio drop is reduced from 1.40 dB to 0.44 dB. The throughputs of the proposed encoder and decoder are 7.04 Gbps and 7.63 Gbps, respectively, and their respective gate counts are 37.2 K and 54.1 K. Sunwoong Kim |
IEEE Trans. Multim. | 1 |
| 2011 | Power-aware design with various low-power algorithms for an H.264/AVC encoderabstractH.264/AVC video compression standard provides high coding efficiency, but requires a considerable amount of complexity and power consumption. This paper presents advanced low-power algorithms for an H.264/AVC encoder and a power-aware design composed of low-power algorithms. Power reduction algorithms with frame memory compression and early skip mode decision are presented, and the search range for motion estimation is reduced for further power reduction. The proposed power aware design controls the power consumption depending on the remaining energy by controlling the operation condition of the proposed low-power algorithms. In order to estimate the power reduction by the proposed algorithms, the power consumed by external memory as well as the bus between an H.264 encoder and an external DRAM is considered. Simulation results show that up to 49.9% of the power consumed by bus and external memory is reduced and that the power consumption from 0% to 41.56% is achieved with a reasonably small degradation of R-D performance. Hyun Kim 0001, Chae-Eun Rhee, Sunwoong Kim |
ISCAS | 4 |