Jaewoo Park 0006

dblp:35/3306-6 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-6477-9813ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EP-HDC: Hyperdimensional Computing with Encrypted Parameters for High-Throughput Privacy-Preserving Inference
abstract
While homomorphic encryption (HE) provides strong privacy protection, its high computational cost has restricted its application to simple tasks. Recently, hyperdimensional computing (HDC) applied to HE has shown promising performance for privacy-preserving machine learning (PPML). However, when applied to more realistic scenarios such as batch inference, the HDC-based HE has still very high compute time as well as high encryption and data transmission overheads. To address this problem, we propose HDC with encrypted parameters (EP-HDC), which is a novel PPML approach featuring client-side HE, i.e., inference is performed on a client using a homomorphically encrypted model. Our EP-HDC can effectively mitigate the encryption and data transmission overhead, as well as providing high scalability with many clients while providing strong protection for user data and model parameters. In addition to application examples for our client-side PPML, we also present design space exploration involving quantization, architecture, and HE-related parameters. Our experimental results using the BFV scheme and the Face/Emotion datasets demonstrate that our method can improve throughput and latency of batch inference by orders of magnitude over previous PPML methods ($36.52 \sim 1068 \times$ and $6.45 \sim 733 \times$, respectively) with <1% accuracy degradation.
Jaewoo Park 0006, Chenghao Quan, Jongeun Lee
ASP-DAC1
2025 SPIMA: Scalable and Cost-Efficient Sparse Matrix Multiplication via Processing in DRAM Array
abstract
Sparse matrix multiplication (SpMM) is a critical kernel used in a wide range of applications, but irregular memory access patterns and memory bandwidth bottleneck as well as load imbalance make the efficient and scalable processing on parallel architectures a significant challenge. Motivated by the memory-bound nature of SpMM computation, we propose a cost-effective SpMM accelerator based on a DRAM processing-in-memory (PIM) approach. Our design introduces a novel dataflow to exploit high bank-level parallelism and reuse both input and output data even for highly sparse matrices. Our proposed architecture, SPIMA, features multiple input buffers for scheduling DRAM access, a output buffer and vector register files working holistically, co-designed to maximize the performance of our novel dataflow. Our experimental results using various sparse matrices demonstrate that our proposed dataflow and architecture are robust in terms of matrix size and sparsity. Compared with the state-of-the-art accelerators implemented on PIM, ASIC, and FPGA, we estimate that our PIM architecture can yield competitive performance with highly sparse matrices.
Tairali Assylbekov, Minsang Yu, Jaewoo Park 0006, Mingon Kim, Seungsu Kim, Jongeun Lee
ICCAD3
2023 NTT-PIM: Row-Centric Architecture and Mapping for Efficient Number-Theoretic Transform on PIM
abstract
Recently DRAM-based PIMs (processing-in-memories) with unmodified cell arrays have demonstrated impressive performance for accelerating AI applications. However, due to the very restrictive hardware constraints, PIM remains an accelerator for simple functions only. In this paper we propose NTT-PIM, which is based on the same principles such as no modification of cell arrays and very restrictive area budget, but shows state-of-the-art performance for a very complex application such as NTT, thanks to features optimized for the application’s characteristics, such as in-place update and pipelining via multiple buffers. Our experimental results demonstrate that our NTT-PIM can outperform previous best PIM-based NTT accelerators in terms of runtime by 1.7 ∼ 17× while having negligible area and power overhead.
Jaewoo Park 0006, Sugil Lee, Jongeun Lee
DAC1
2023 Hyperdimensional Computing as a Rescue for Efficient Privacy-Preserving Machine Learning-as-a-Service
abstract
Machine learning models are often provisioned as a cloud-based service where the clients send their data to the service provider to obtain the result. This setting is commonplace due to the high value of the models, but it requires the clients to forfeit the privacy that the query data may contain. Homomorphic encryption (HE) is a promising technique to address this adversity. With HE, the service provider can take encrypted data as a query and run the model without decrypting it. The result remains encrypted, and only the client can decrypt it. All these benefits come at the cost of computational cost because HE turns simple floating-point arithmetic into the computation between long (of degree ≥ 1024) polynomials. Previous work has proposed to tailor deep neural networks for efficient computation over encrypted data, but already high computational cost is again amplified by HE, hindering performance improvement. In this paper we show hyperdimensional computing can be a rescue for privacy-preserving machine learning over encrypted data. We find that the advantage of hyperdimensional computing in performance is amplified when working with HE. This observation led us to design HE-HDC, a machine-learning inference system that uses hyperdimensional computing with HE. We carefully structure the machine learning service so that the server will perform only the HE-friendly computation. Moreover, we adapt the computation and HE parameters to expedite computation while preserving accuracy and security. Our experimental result based on real measurements shows that HE-HDC outperforms existing systems by 26 ~ 3000 x times with comparable classification accuracy.
Jaewoo Park 0006, Chenghao Quan, Hyungon Moon, Jongeun Lee
ICCAD1
2022 Centered Symmetric Quantization for Hardware-Efficient Low-Bit Neural Networks
Faaiz Asim, Jaewoo Park 0006, Azat Azamat, Jongeun Lee
BMVC2
2022 Squeezing Accumulators in Binary Neural Networks for Extremely Resource-Constrained Applications
abstract
The cost and power consumption of BNN (Binarized Neural Network) hardware is dominated by additions. In particular, accumulators account for a large fraction of hardware overhead, which could be effectively reduced by using reduced-width accumulators. However, it is not straightforward to find the optimal accumulator width due to the complex interplay between width, scale, and the effect of training. In this paper we present algorithmic and hardware-level methods to find the optimal accumulator size for BNN hardware with minimal impact on the quality of result. First, we present partial sum scaling, a top-down approach to minimize the BNN accumulator size based on advanced quantization techniques. We also present an efficient, zero-overhead hardware design for partial sum scaling. Second, we evaluate a bottom-up approach that is to use saturating accumulator, which is more robust against overflows. Our experimental results using CIFAR-10 dataset demonstrate that our partial sum scaling along with our optimized accumulator architecture can reduce the area and power consumption of datapath by 15.50% and 27.03%, respectively, with little impact on inference performance (less than 2%), compared to using 16-bit accumulator.
Azat Azamat, Jaewoo Park 0006, Jongeun Lee
ICCAD2