VLDB 2026 Research / reviewers in the wild / expert
Bokyung Kim 0001
dblp:122/9032-1
· DBLP profile ↗
9ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0001-5578-5237ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 8 first-author · 9 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Privacy-Preserving Machine Learning Accelerators: Balancing Efficiency and Security
Bokyung Kim 0001 |
ISCAS | 1 |
| 2026 | DPIMA: A Differentially Private Processing-in-DRAM AcceleratorabstractThe rapid success of machine learning (ML) has accelerated its deployment in a wide range of sensitive domains, raising critical concerns about data privacy. Privacy-preserving ML (PPML) has, therefore, emerged as a fundamental requirement, with differential privacy (DP) providing a rigorous mathematical framework to guarantee individual privacy while enabling meaningful learning. However, applying DP to deep learning imposes substantial computational and memory overhead due to noise injection, privacy accounting, and additional data manipulation steps, which are not well supported by existing ML hardware accelerators. To address this challenge, we proposeDPIMA, a novel DP-guaranteed processing-in-memory (PIM) accelerator designed specifically to support efficient DP training. DPIMA leverages the intrinsic parallelism and bandwidth of DRAM-based PIM while introducing architectural innovations, including a full-operation adder tree (FOAT) and a system-levelprivacy-aware dataflowthat jointly optimize the execution of DP-specific operations. Through a comprehensive set of experiments across diverse deep learning models and training configurations, DPIMA demonstrates up to$13.8\times $speedup and$123.8\times $energy efficiency improvement over the baseline design for DP training. These results highlight the effectiveness of codesigning privacy mechanisms with memory-centric hardware, paving the way for scalable and efficient privacy-preserving learning systems. Bokyung Kim 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | A DRAM-Based Processing-in-Memory Accelerator for Privacy-Protecting Machine LearningabstractThe sensational outcomes of machine learning (ML) are witnessed. As the big success relies on continual training with massive data encompassing sensitive information, deep neural network (DNN) models easily leak private information. For instance, large pre-trained language models contain a substantial volume of private information, which can be acquired by querying with appropriate prompts. This raises concerns regarding privacy in ML, and DNN models are evolving for privacy-preserving ML (PPML). Bokyung Kim 0001 |
DATE | 1 |
| 2025 | DPIMA: A DRAM-Based Processing-in-Memory Accelerator for Privacy-Preserving Machine LearningabstractThe unprecedented success of machine learning (ML) increases the adoption of ML in sensitive domains. Consequently, privacy-preserving ML (PPML) becomes inevitable for data privacy. Differential privacy (DP) provides a formal framework to protect individual data, allowing meaningful learning. However, DP methods, particularly in deep learning, introduce significant computational overhead with their unique computing phases. Existing hardware accelerators optimized for standard ML are not specifically designed for DP workloads, leading to inefficiencies in privacy preservation computation. Driven by the motivation, this paper proposes a novel differential private processing-in-memory accelerator, DPIMA. Our specialized accelerator efficiently supports differentially private training, minimizing performance degradation and maintaining privacy guarantees. Leveraging the inherent advantages of DRAM and PIM, we offer novelties on micro-architecture with full operation adder tree (FOAT) and systemic design with dataflow. Our experiments with various learning models demonstrate that DPIMA achieves an average improvement of 13.8× and 123.8× in performance and energy efficiency compared to baseline. Bokyung Kim 0001 |
ISLPED | 1 |
| 2025 | Efficient and Robust Edge AI: Software, Hardware, and the Co-designabstractArtificial intelligence (AI) provides versatile capabilities in applications such as image classification and voice recognition that are most useful in edge or mobile computing settings. Shrinking these sophisticated algorithms into small form factors with minimal computing resources and power budgets requires innovation at several layers of abstraction: software, algorithmic, architectural, circuit, and device-level innovations. However, improvements to system efficiency may impact robustness and vice-versa. Therefore, a co-design framework is often necessary to customize a system for its given application. A system that prioritizes efficiency might use circuit-level innovations that introduce process variations or signal noise into the system, which may use software-level redundancy in order to compensate. In this tutorial, we will first examine various methods of improving efficiency and robustness in edge AI and their tradeoffs at each level of abstraction. Then, we will outline co-design techniques for designing efficient and robust edge AI systems, using federated learning as a specific example to illustrate the effectiveness of co-design. Bokyung Kim 0001, Shiyu Li 0001, Brady Taylor, Yiran Chen 0001 |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2024 | Processing-in-Memory Designs Based on Emerging Technology for Efficient Machine Learning AccelerationabstractThe unprecedented success of artificial intelligence (AI) enriches machine learning (ML)-based applications. The availability of big data and compute-intensive algorithms empowers versatility and high accuracy in ML approaches. However, the data processing and innumerable computations burden conventional hardware systems with high power consumption and low performance. Breaking away from the traditional hardware design, non-conventional accelerators exploiting emerging technology have gained significant attention with a leap forward since the emerging devices enable processing-in-memory (PIM) designs of dramatic improvement in efficiency. This paper presents a summary of state-of-the-art PIM accelerators over a decade. The PIM accelerators have been implemented for diverse models and advanced algorithm techniques across diverse neural networks in language processing and image recognition to expedite inference and training. We will provide the implemented designs, methodologies, and results, following the development in the past years. The promising direction of the PIM accelerators, vertically stacking for More than Moore, is also discussed. Bokyung Kim 0001, Hai Li 0001, Yiran Chen 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2023 | INCA: Input-stationary Dataflow at Outside-the-box Thinking about Deep Learning AcceleratorsabstractThis paper first presents an input-stationary (IS) implemented crossbar accelerator (INCA), supporting inference and training for deep neural networks (DNNs). Processing-in-memory (PIM) accelerators for DNNs have been actively researched, specifically, with resistive random-access memory (RRAM), due to RRAM’s computing and memorizing capabilities and device merits. To the best of our knowledge, all previous PIM accelerators have saved weights into RRAMs and inputs (activations) into conventional memories—it naturally forms weight-stationary (WS) dataflow. WS has generally been considered the most optimized choice for high parallelism and data reuse. How-ever, WS-based PIM accelerators show fundamental limitations: first, remaining high dependency on DRAM and buffers for fetching and saving inputs (activations); second, a remarkable number of extra RRAMs for transposed weights and additional computational intermediates in training; third, coarse-grained arrays demanding high-bit analog-to-digital converters (ADCs) and introducing poor utilization in depthwise and pointwise convolution; last, degraded accuracy due to its sensitivity to weights which are affected by RRAM’s nonideality. On the other hand, we observe that IS dataflow, where RRAMs retain inputs (activations), can effectively address the limitations of WS, because of low dependency by only loading weights, no need for extra RRAMs, feasibility of fine-grained accelerator design, and less impact of input (activation) variance on accuracy. But IS dataflow is hardly achievable by the existing crossbar structure because it is difficult to implement kernel sliding and preserve the high parallelism. To support kernel movement, we constitute a cell structure with two-transistor-one-RRAM (2T1R). Based on the 2T1R cell, we design a novel three-dimensional (3D) architecture for high parallelism in batch training. Our experiment results prove the potential of INCA. Compared to the WS accelerator, INCA achieves up to 20.6× and 260× energy efficiency improvement in inference and training, respectively; 4.8× (inference) and 18.6× (training) speedup as well. While accuracy in WS drops to 15% in our high-noise simulation, INCA presents an even more robust result as 86% accuracy. Bokyung Kim 0001, Shiyu Li 0001, Hai Li 0001 |
HPCA | 1 |
| 2023 | Invited Paper: Towards the Efficiency, Heterogeneity, and Robustness of Edge AIabstractOver the past decade, there has been a persistent trend in edge computing, driving the migration of intelligence closer to the edge. The increasing need to process data locally has fueled the deployment of highly efficient computing hardware and artificial intelligence (AI) models onto edge devices. The performance and robustness of edge computing systems are significantly influenced by the heterogeneity of computing systems and the diverse nature of data to be processed by each edge device. This paper aims to explore the principles of software/hardware co-design for edge computing systems in AI applications. We will delve into the robustness concerns faced by edge AI due to the inherent heterogeneity of systems and data. Furthermore, we will present various solutions that effectively mitigate these adverse effects and enhance the resilience of edge AI systems. Bokyung Kim 0001, Zhixu Du, Jingwei Sun 0002, Yiran Chen 0001 |
ICCAD | 1 |
| 2022 | Bionic Robust Memristor-Based Artificial Nociception System for RoboticsabstractNociception is an important ability for robots to interact safely with humans or work in hostile environments. By referring to previous research in mechanical receptors with ring oscillators and memristor-based nociceptors, we propose a complete robotic sensing system that senses forces and processes electrical signals at both the edge and the central. This artificial system mimics the nociception behavior of the human nervous system processing external damaging mechanical stimuli. Given mechanical receptors and nociceptors are subject to damage under harsh working conditions, we designed an error detection module that involves memristor crossbar arrays (CBA) to detect components’ potential failures. Furthermore, we propose solutions to non-idealities such as the quantization error and the memristance programming variations to boost the accuracy and the robustness of our failure detection memristor crossbar array. The precision CBA proposed reduces the quantization error for a single cell to at least 1/180 of the error under the common strategy. Our simulations show that using serial connection memristors cells and high resistance ratio cells reduce the output current variation due to programming variation to 62.2% and 71.52% respectively. The full-system simulation under a typical application scenario shows the system's ability to sense noxious stimuli and maintain system robustness at a high level. Guangyu Feng, Bokyung Kim 0001, Hai Li 0001 |
ISCAS | 2 |