EDBT 2026 Demo / reviewers in the wild / expert
Kevin Nam
dblp:58/6563
· DBLP profile ↗
12ranked-venue papers
6as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 4 since 2021Security and privacy · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HEPIC: Private Inference over Homomorphic Encryption with Client InterventionabstractHomomorphic Encryption (HE) enables Private Inference (PI) in Machine Learning as a Service (MLaaS), protecting both client inputs and server-side neural network (NN) parameters. Existing PI techniques are predominantly implemented as either HE-based fire-and-forget methods or MPC-based interactive methods. Recent HE-based PI systems improve the accuracy--performance trade-off via a layer-wise scheme and parameter switching, yet remain bottlenecked by fire-and-forget execution in which the server alone performs costly ciphertext management (e.g., bootstrapping and scheme/parameter conversions). We present HEPIC, an HE-based PI system that explores a different design point by leveraging client interventions for ciphertext managements. In a sense, HEPIC shares a common ground with MPC-based PI of being interactive with the client, but differs in that the client only intervenes for ciphertext managements required in HE operations. Because ciphertext management has identical semantics on the client and the server, HEPIC lets developers decide where and how often to execute it, enabling fine-grained trade-offs among computation, communication, and ciphertext configuration. HEPIC makes such execution practical by overlapping client re-encryption, server computation, and communication via dependency-aware pipelining and streaming-based transfers. We further enhance the performance with a cache-aware task allocator (CATA) and a cost-aware client intervention scheduler (CACIS) to exploit ciphertext-level parallelism and to mitigate stalls under client-server performance disparity. Our evaluation shows that HEPIC achieves up to 2.20--41.93× speedup over state-of-the-art fire-and-forget HE-based PI, while maintaining zero loss in inference accuracy. Kevin Nam, Youyeon Joo, Seungjin Ha, Hyungon Moon, Yunheung Paek |
ASPLOS (2) | 1 |
| 2025 | Affinity-based Optimizations for TFHE on Processing-in-DRAMabstractProcessing-in-memory (PIM) architectures are promising for accelerating intensive workloads due to their high internal bandwidth. This paper introduces a technique for accelerating Fully Homomorphic Encryption over the Torus (TFHE), a promising yet intensive application, on a realistic PIM system. Existing TFHE accelerators focus on exploiting parallelism, often overlooking data affinity, which leads to performance degradation in PIM due to excessive remote data accesses (RDAs). To address this, we present an affinity-based approach that optimizes the computation of TFHE on PIM. We apply algorithmic optimizations to TFHE, enabling PIM to effectively leverage its high internal bandwidth. We analyze the affinity patterns in the sub-tasks of TFHE and develop an offline scheduler that exploits our analysis to find optimal scheduling, minimizing RDAs while maintaining sufficient parallelism. To demonstrate the practicality of our work, we design a variant of an existing PIM-HBM device with minimal hardware modifications, and perform evaluations over a real FPGA-based PIM system. Our experiments demonstrate that our affinity-based optimizations outperform prior TFHE accelerators by 4.24-209× for real-world benchmarks. Kevin Nam, Heon Hui Jung, Hyunyoung Oh, Yunheung Paek |
ASPLOS (2) | 1 |
| 2025 | An Accelerator for Low-Computational Overhead Privacy-Preserving GNN InferenceabstractGraph Neural Networks (GNNs) are increasingly used in domains such as finance and bioinformatics, where both node features and edge structures can contain sensitive information. While Fully Homomorphic Encryption (FHE) offers a promising solution for privacy-preserving GNN inference, existing approaches such as PPGNN rely on costly Homomorphic Rotation and MUX operations for operand obfuscation, resulting in significant computational overhead. In this work, we propose a new obfuscation method that leverages the probabilistic nature of FHE to duplicate ciphertexts at the client side, thereby eliminating the need for runtime selection logic. To support this method efficiently, we design a pipelined hardware accelerator with a simplified CKKS datapath and parallel TFHE execution, avoiding the complexity of rotation-heavy designs. Despite reduced ciphertext reuse, our architecture mitigates memory pressure through buffer-aware PBS unit design. Experimental results demonstrate up to$8.8 \times$speedup and$7.69 \times$energy efficiency improvement over PPGNN, while also outperforming existing multi-scheme accelerators such as Trinity and UFC even when applying the same obfuscation strategy. Our approach offers a practical and scalable solution for efficient, privacy-preserving GNN inference. Heon Hui Jung, Whoi Ree Ha, Kevin Nam, Youyeon Joo, Lucas Oros, Yunheung Paek |
HiPC | 3 |
| 2025 | SLOTHE : Lazy Approximation of Non-Arithmetic Neural Network Functions over Encrypted Data
Kevin Nam, Youyeon Joo, Seungjin Ha, Yunheung Paek |
USENIX Security Symposium | 1 |
| 2025 | LOHEN: Layer-wise Optimizations for Neural Network Inferences over Encrypted Data with High Performance or Accuracy
Kevin Nam, Youyeon Joo, Dongju Lee, Seungjin Ha, Hyunyoung Oh, Hyungon Moon, Yunheung Paek |
USENIX Security Symposium | 1 |
| 2022 | Accelerating N-Bit Operations over TFHE on Commodity CPU-FPGAabstractTFHE is a fully homomorphic encryption (FHE) scheme that evaluates Boolean gates, which we will hereafter call Tgates, over encrypted data. TFHE is considered to have higher expressive power than many existing schemes in that it is able to compute not only N-bit Arithmetic operations but also Logical/Relational ones as arbitrary ALR operations can be represented by Tgate circuits. Despite such strength, TFHE has a weakness that like all other schemes, it suffers from colossal computational overhead. Incessant efforts to reduce the overhead have been made by exploiting the inherent parallelism of FHE operations on ciphertexts. Unlike other FHE schemes, the parallelism of TFHE can be decomposed into multilayers: one inside each FHE operation (equivalent to a single Tgate) and the other between Tgates. Unfortunately, previous works focused only on exploiting the parallelism inside Tgate. However, as each N-bit operation over TFHE corresponds to a Tgate circuit constructed from multiple Tgates, it is also necessary to utilize the parallelism between Tgates for optimizing an entire operation. This paper proposes an acceleration technique to maximize performance of a TFHE N-bit operation by simultaneously utilizing both parallelism comprising the operation. To fully profit from both layers of parallelism, we have implemented our technique on a commodity CPU-FPGA hybrid machine with parallel execution capabilities in hardware. Our implementation outperforms prior ones by 2.43× in throughput and 12.19× in throughput per watt when performing N-bit operations under the 128-bit quantum security parameters. Kevin Nam, Hyunyoung Oh, Hyungon Moon, Yunheung Paek |
ICCAD | 1 |
| 2017 | FISH: Linux system calls for FPGA acceleratorsabstractThis, paper presents the FISH (FPGA-Initiated Software-Handled) framework which allows FPGA accelerators to make system calls to the Linux operating system in CPU-FPGA systems. A special FISH Linux kernel module running on the CPU provides a system call interface for FPGA accelerators, much like the ABI which exists for software programs. We provide a proof-of-concept implementation of this framework running on the Intel Cyclone V SoC device, and show that an FPGA accelerator can seamlessly make system calls as if it were the host program. We see the FISH framework being especially useful for high-level synthesis (HLS) by making it possible to synthesize software code that contains system calls. Kevin Nam, Blair Fort, Stephen Brown 0003 |
FPL | 1 |
| 2014 | Visualization evaluation for cyber security: trends and future directionsabstractThe Visualization for Cyber Security research community (VizSec) addresses longstanding challenges in cyber security by adapting and evaluating information visualization techniques with application to the cyber security domain. This research effort has created many tools and techniques that could be applied to improve cyber security, yet the community has not yet established unified standards for evaluating these approaches to predict their operational validity. In this paper, we survey and categorize the evaluation metrics, components, and techniques that have been utilized in the past decade of VizSec research literature. We also discuss existing methodological gaps in evaluating visualization in cyber security, and suggest potential avenues for future research in order to help establish an agenda for advancing the state-of-the-art in evaluating cyber security visualizations. Diane Staheli, Tamara Yu, R. Jordan Crouser, Suresh Damodaran, Kevin Nam, B. David O'Gwynn, Sean McKenna, Lane Harrison |
VizSEC | 5 |
| 2012 | Impact of Cache Architecture and Interface on Performance and Area of FPGA-Based Processor/Parallel-Accelerator SystemsabstractWe describe new multi-ported cache designs suitable for use in FPGA-based processor/parallel-accelerator systems, and evaluate their impact on application performance and area. The baseline system comprises a MIPS soft processor and custom hardware accelerators with a shared memory architecture: on-FPGA L1 cache backed by off-chip DDR2 SDRAM. Within this general system model, we evaluate traditional cache design parameters (cache size, line size, associativity). In the parallel accelerator context, we examine the impact of the cache design and its interface. Specifically, we look at how the number of cache ports affects performance when multiple hardware accelerators operate (and access memory) in parallel, and evaluate two different hardware implementations of multi-ported caches using: 1) multi-pumping, and 2) a recently-published approach based on the concept of a live-value table. Results show that application performance depends strongly on the cache interface and architecture: for a system with 6 accelerators, depending on the cache design, speed up swings from 0.73× to 6.14×, on average, relative to a baseline sequential system (with a single accelerator and a direct-mapped, 2KB cache with 32B lines). Considering both performance and area, the best architecture is found to be a 4-port multi-pump direct-mapped cache with a 16KB cache size and a 128B line size. Jongsok Choi, Kevin Nam, Andrew Canis, Jason Helge Anderson, Stephen Brown 0003, Tomasz S. Czajkowski |
FCCM | 2 |
| 2012 | Impact of FPGA architecture on resource sharing in high-level synthesisabstractResource sharing is a key area-reduction approach in high-level synthesis (HLS) in which a single hardware functional unit is used to implement multiple operations in the high-level circuit specification. We show that the utility of sharing depends on the underlying FPGA logic element architecture and that different sharing trade-offs exist when 4-LUTs vs. 6-LUTs are used. We further show that certain multi-operator patterns occur multiple times in programs, creating additional opportunities for sharing larger composite functional units comprised of patterns of interconnected operators. A sharing cost/benefit analysis is used to inform decisions made in the binding phase of an HLS tool, whose RTL output is targeted to Altera commercial FPGA families: Stratix IV (dual-output 6-LUTs) and Cyclone II (4-LUTs). Stefan Hadjis, Andrew Canis, Jason Helge Anderson, Jongsok Choi, Kevin Nam, Stephen Brown 0003, Tomasz S. Czajkowski |
FPGA | 5 |
| 2010 | Finding the lost treasure: understanding reuse of used computing devicesabstractIn this paper, we report our findings on the adoption practices of used personal digital assistants (PDAs) to inform reuse of outdated computing products. Our interviews with 12 eBay users who bought used PDAs showed a variety of ways in which users indirectly supported sustainability. This allowed us to re-examine sustainability as something that is dynamically and arbitrarily shaped by the users and not just dependent on the sustainable feature of the product. We end with design implications for supporting users' shaping of sustainability. Jina Huh, Kevin Nam |
CHI | 2 |
| 2007 | Testing the technology: playing games with video conferencingabstractVideo connections can establish a media space in which games may be played, just as people play games while collocated. Experiments with participants playing the game 'Mafia' indicate that people in a video condition have similar levels of satisfaction, fun, and frustration, to those that play while collocated. This finding holds for both those with prior experience using video systems and those without, suggesting it is not merely a "novelty effect." Results differ about whether there exist differences in focus of attention, suspicion/trust, and pointing for people playing the game while using a video system. Implications for both fun and work uses of video are suggested. Archer L. Batcheller, Brian Hilligoss, Kevin Nam, Emilee Rader, Marta Rey-Babarro, Xiaomu Zhou |
CHI | 3 |