Hyunyoung Oh

dblp:236/7050 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Affinity-based Optimizations for TFHE on Processing-in-DRAM
abstract
Processing-in-memory (PIM) architectures are promising for accelerating intensive workloads due to their high internal bandwidth. This paper introduces a technique for accelerating Fully Homomorphic Encryption over the Torus (TFHE), a promising yet intensive application, on a realistic PIM system. Existing TFHE accelerators focus on exploiting parallelism, often overlooking data affinity, which leads to performance degradation in PIM due to excessive remote data accesses (RDAs). To address this, we present an affinity-based approach that optimizes the computation of TFHE on PIM. We apply algorithmic optimizations to TFHE, enabling PIM to effectively leverage its high internal bandwidth. We analyze the affinity patterns in the sub-tasks of TFHE and develop an offline scheduler that exploits our analysis to find optimal scheduling, minimizing RDAs while maintaining sufficient parallelism. To demonstrate the practicality of our work, we design a variant of an existing PIM-HBM device with minimal hardware modifications, and perform evaluations over a real FPGA-based PIM system. Our experiments demonstrate that our affinity-based optimizations outperform prior TFHE accelerators by 4.24-209× for real-world benchmarks.
Kevin Nam, Heon Hui Jung, Hyunyoung Oh, Yunheung Paek
ASPLOS (2)3
2025 LOHEN: Layer-wise Optimizations for Neural Network Inferences over Encrypted Data with High Performance or Accuracy
Kevin Nam, Youyeon Joo, Dongju Lee, Seungjin Ha, Hyunyoung Oh, Hyungon Moon, Yunheung Paek
USENIX Security Symposium5
2022 XTENSTORE: Fast Shielded In-memory Key-Value Store on a Hybrid x86-FPGA System
abstract
We propose XtenStore, a system that extends the existing SGX-based secure in-memory key-value store with an external hardware accelerator in order to ensure comparable security guarantees with lower performance degradation. The accelerator is implemented on a commodity FPGA card that is readily connected with the x86 CPU via PCIe interconnect to form a hybrid x86-FPGA system. In comparison to the prior SGX-based work, XtenStore improves the throughput by 4–33x, and exhibits considerably shorter tail latency (>23x, 99th-percentile).
Hyunyoung Oh, Dongil Hwang, Maja Malenko, Myunghyun Cho, Hyungon Moon, Marcel Baunach, Yunheung Paek
DATE1
2022 Accelerating N-Bit Operations over TFHE on Commodity CPU-FPGA
abstract
TFHE is a fully homomorphic encryption (FHE) scheme that evaluates Boolean gates, which we will hereafter call Tgates, over encrypted data. TFHE is considered to have higher expressive power than many existing schemes in that it is able to compute not only N-bit Arithmetic operations but also Logical/Relational ones as arbitrary ALR operations can be represented by Tgate circuits. Despite such strength, TFHE has a weakness that like all other schemes, it suffers from colossal computational overhead. Incessant efforts to reduce the overhead have been made by exploiting the inherent parallelism of FHE operations on ciphertexts. Unlike other FHE schemes, the parallelism of TFHE can be decomposed into multilayers: one inside each FHE operation (equivalent to a single Tgate) and the other between Tgates. Unfortunately, previous works focused only on exploiting the parallelism inside Tgate. However, as each N-bit operation over TFHE corresponds to a Tgate circuit constructed from multiple Tgates, it is also necessary to utilize the parallelism between Tgates for optimizing an entire operation. This paper proposes an acceleration technique to maximize performance of a TFHE N-bit operation by simultaneously utilizing both parallelism comprising the operation. To fully profit from both layers of parallelism, we have implemented our technique on a commodity CPU-FPGA hybrid machine with parallel execution capabilities in hardware. Our implementation outperforms prior ones by 2.43× in throughput and 12.19× in throughput per watt when performing N-bit operations under the 128-bit quantum security parameters.
Kevin Nam, Hyunyoung Oh, Hyungon Moon, Yunheung Paek
ICCAD2
2021 A metadata-driven approach to efficiently detect code-reuse attacks on ARM multiprocessors
Hyunyoung Oh, Yeongpil Cho, Yunheung Paek
J. Supercomput.1
2020 TRUSTORE: Side-Channel Resistant Storage for SGX using Intel Hybrid CPU-FPGA
abstract
Intel SGX is a security solution promising strong and practical security guarantees for trusted computing. However, recent reports demonstrated that such security guarantees of SGX are broken due to access pattern based side-channel attacks, including page fault, cache, branch prediction, and speculative execution. In order to stop these side-channel attackers, Oblivious RAM (ORAM) has gained strong attention from the security community as it provides cryptographically proven protection against access pattern based side-channels. While several proposed systems have successfully applied ORAM to thwart side-channels, those are severely limited in performance and its scalability due to notorious performance issues of ORAM. This paper presents TrustOre, addressing these issues that arise when using ORAM with Intel SGX. TrustOre leverages an external device, FPGA, to implement a trusted storage service within a completed isolated environment secure from side-channel attacks. TrustOre tackles several challenges in achieving such a goal: extending trust from SGX to FPGA without imposing architectural changes, providing a verifiably-secure connection between SGX applications and FPGA, and seamlessly supporting various access operations from SGX applications to FPGA.We implemented TrustOre on the commodity Intel Hybrid CPU-FPGA architecture. Then we evaluated with three state-of-the-art ORAM-based SGX applications, ZeroTrace, Obliviate, and Obfuscuro, as well as an end-to-end key-value store application. According to our evaluation, TrustOre-based applications outperforms ORAM-based original applications ranging from 10x to 43x, while also showing far better scalability than ORAM-based ones. We emphasize that since TrustOre can be deployed as a simple plug-in to SGX machine's PCIe slot, it is readily used to thwart side-channel attacks in SGX, arguably one of the most cryptic and critical security holes today.
Hyunyoung Oh, Adil Ahmad, Seonghyun Park 0001, Byoungyoung Lee, Yunheung Paek
CCS1
2020 Life Review Using a Life Metaphoric Game to Promote Intergenerational Communication
abstract
This study explores whether a life metaphoric game can effectively facilitate communication between senior participants and young adult partners during a life review activity. Life review provides older adults with opportunities to organize their past experiences, rediscover the meaning of life, and prepare for their future life and eventual death. We held workshops in which 33 senior participants (ages 51-85) co-played a commercial game, titled "Long Journey of Life," paired with young adult partners (ages 19-23). The young partners asked senior participants about their associated memories during the life review activity. Through inductive thematic analysis, we found seven communication themes during the life review activity: reminiscence, future and death preparation, life advice, appreciation and evaluation, small talks involving self-disclosure, interpretation of metaphors, and explanation of game control methods. We investigated the relationships between communication themes and game design elements that promote conversation. In addition, we identified how participants' responses to the game differ depending on the player's characteristics and generation. Game design suggestions for an effective intergenerational life review using digital games were offered.
Seyeon Lee, Hyunyoung Oh, Chungkon Shi, Young Yim Doh
Proc. ACM Hum. Comput. Interact.2
2019 Real-Time Anomalous Branch Behavior Inference with a GPU-inspired Engine for Machine Learning Models
abstract
Attacks on embedded devices are likely to occur any time in unexpected manners. Thus, the defense systems based on fixed sets of rules will easily be subverted by such unexpected, unknown attacks. Learning-based anomaly detection may potentially prevent new unknown zero-day attacks by leveraging the capability of machine learning (ML) to learn the intricate true nature of software hidden within raw information. This paper introduces our work to develop an MPSoC, called RTAD, which can efficiently support in hardware various ML models that run to detect anomalous behaviors on embedded devices in a real-time fashion, and thus enable the devices to counteract the anomalies in the field. In the IoT era, the importance of security for embedded devices cannot be exaggerated because they will become an enticing target for adversaries as they are being integrated into everyday life to provide users with various services. The above-mentioned potential of learning-based detection is believed to benefit those deployed devices under attacks occurring any time during their field operations in unexpected manners. We hereby assume that ML models are trained with runtime branch information as their data features since a sequence of branches serves as a record of control flow transfers during program execution. In fact, there have been numerous ML studies that examine various types of branches in order to infer (or detect) anomaly in branch behaviors that may be induced by diverse attacks that can cause deviant control flow in software. Our goal of real-time anomalous branch behavior inference poses two challenges to our development of RTAD. Firstly, RTAD must collect and transfer in a timely fashion a sequence of branches as the input to the ML model. Secondly, RTAD must be able to promptly process the delivered branch data with the ML model. To tackle these challenges, we have implemented in RTAD two core components: an input generation module and a GPU-inspired ML processing engine. According to our experiments, RTAD enables various ML models to infer anomaly instantly after the victim program behaves aberrantly as the result of attacks being injected into the system.
Hyunyoung Oh, Hayoon Yi, Hyeokjun Choe, Yeongpil Cho, Sungroh Yoon, Yunheung Paek
DATE1
2007 Collision-Free Interleavers Using Latin Squares for Parallel Decoding of Turbo Codes
abstract
In the parallel decoding of turbo codes, the constituent interleaver must avoid the memory collision. This paper proposes a collision-free interleaver structure which can be optimized easily over various information blocksizes. The performance of the proposed interleaver is almost the same as or 0.1dB loss against almost regular permutation (ARP) at FER 10-5region with information block sizes of 320 and 640 when the simulation environment is given by the 3GPP standard turbo codes with 4 parallelism in AWGN channel.
Hyunyoung Oh, Dae-Son Kim, Joon-Sung Kim, Hong-Yeop Song
VTC Spring1