Donghyun Kwon

dblp:130/9723 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-7507-3111ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 1 first-author · 3 since 2021Security and privacy · 9 · 2 first-author · 5 since 2021Computer networks · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 Beyond address spaces: In-process memory isolation for RISC-V
Hayoung Kang, Donghyun Kwon
Comput. Secur.3
2026 TzPy: Python Runtime for Edge Code Deployment on TrustZone
abstract
As edge devices continue to improve in performance, their role is expanding in real-world environments such as smart factories and smart homes, where executing server-provided code on edge devices is essential for real-time processing and operational flexibility. However, edge code deployment introduces several security challenges, including the exposure of sensitive data and operations to runtime attacks, the absence of remote attestation mechanisms to validate code integrity and execution correctness, and insufficient access control for sensitive peripherals. Existing solutions address only a subset of these challenges and therefore fail to provide comprehensive security guarantees for edge devices. In this article, we present TzPy , a Python runtime designed for secure edge code deployment on ARM TrustZone. TzPy enables the execution of Python code entirely within the secure world. In addition, we design a remote attestation mechanism that verifies both the static integrity of deployed Python code and its runtime control flow through control-flow tracing. Furthermore, TzPy enforces a manifest-driven, fine-grained peripheral access control mechanism to prevent unauthorised or unintended access to sensitive devices. We evaluate TzPy on a Raspberry Pi 3B+ using a smart farm workload and a blinky application, and the results show an end-to-end overhead of 13.4%.
Suhyeon Song, Jaeyeol Park, Chaewon Shin, Donghyun Kwon
ACM Trans. Embed. Comput. Syst.4
2025 ELK: Effective Lock-and-Key Technique for Temporal Memory Safety on Embedded Devices in ARMv8-M
abstract
Embedded devices built on microcontroller units (MCUs) power a vast array of safety- and performance-critical applications. Contrary to the long-held belief that resource-constrained embedded devices avoid dynamic allocations, recent studies show that heap use is actually common. Therefore, embedded software written in C/C++ is prone to temporal memory safety violations. The lock-and-key mechanism is a well-established defense for temporal memory safety. For high performance, recent studies embed keys directly in the unused high-order bits of 64-bit pointers, but this approach has been considered infeasible on MMU -less embedded devices. In this paper, we propose ELK, an effective lock-and-key technique for temporal memory safety on embedded devices in ARMv8-M. ELK observes that only the low-order bits of a 32-bit pointer are needed to access on-chip SRAM. Thus, ELK embeds keys in the unused middle-range bits of 32-bit pointers and stores locks inline with objects through size-aligned allocation that leverages ARMv8-M hardware features. Evaluating the ELK prototype shows geometric mean runtime overheads of 3.9% and 7.4 % on the BEEBs and MiBench2, respectively. The memory overhead, which represents the increase in peak heap usage, ranges from 0.1 % to 2.6 % across six heap-intensive applications. On the Juliet test suite, ELK detects all 180 temporal memory safety violations with zero false positives.
Jeonghwan Kang, Kyounghwan Kim, Donghyun Kwon
ACSAC3
2025 SMORE: Practical Redzone-Based Stack Memory Error Detection Mechanism for Embedded Systems
abstract
Stack memory errors, such as buffer overflows and over-reads, remain a critical attack vector in embedded systems, which rely heavily on stack memory due to their constrained resources and real-time requirements. Although conventional stack canary techniques have been widely adopted as a lightweight solution, they suffer from several limitations, including limited detection coverage, delayed error reporting, and inability to detect illegal reads. In this paper, we present SMORE, a practical stack memory error detection mechanism designed for ARMv8-M-based embedded systems. SMORE leverages the Memory Protection Unit to implement a hardware-assisted redzone, enabling immediate detection of both read and write violations at runtime. SMORE manages redzones efficiently at function boundaries using the Test Target instructions, minimizing performance impact. We evaluate SMORE on a real embedded platform, the Nucleo-L552ZE-Q running RIOT-OS, and show that it provides significantly improved detection coverage over stack canary while maintaining low performance overhead. These results demonstrate that SMORE provides a practical and effective solution for improving stack memory error detection in embedded software.
Jaeyeol Park, Yunju Gu, Donghyun Kwon
ACSAC3
2025 ROSec: Intra-Process Isolation for ROS Composition With Memory Protection Keys
abstract
Robot Operating System(ROS) is a software framework for robotic systems that includes various packages for developing robotic applications.Compositionis a package that combines multiple applications, namely,nodes, to be loaded and executed in a single process. However, permitting multiple nodes to share the address space could expand the attack surface such that vulnerabilities in a node are more likely to be exploited to subvert nodes running in the same space. We propose ROSec, an in-process isolation solution for ROS composition that utilizes Intel Memory Protection Keys. ROSecaims to enforce memory isolation between nodes within a process by preventing unauthorized access from one node to another. Unlike previous works that assume the number and sizes of nodes are statically defined and partitioned by developers, ROSecis designed to handle the dynamic nature of nodes that can be loaded and executed in a process at any time during execution. To achieve this, ROSecadopts a unique scheduling mechanism that utilizes theexecutor-centricexecution model of ROS to perform two main operations for MPK-based isolation:protection key assignmentandreassignment. Our evaluation shows that ROSeceffectively enforces in-process isolation while incurring a 6.4% performance overhead on a real-world application.Note to Practitioners—Cyber-Physical Systems are the core of modern applications, particularly robotics, as they integrate computing and physical processes. ROS necessitates real-time and security guarantees, which, unfortunately, trade-off with each other. While traditional ROS architecture relies on process isolation to separate various nodes, ROS2 introduces a feature called composition, which allows multiple nodes to run inside a single process, thus exposing various nodes to potential malicious compromises from others. This paper proposes a technique that utilizes Intel memory protection keys (MPK) to provide intraprocess isolation for ROS composition. Given that ROS nodes are dynamic, ROSEC provides key assignment and reassignment techniques to configure MPK dynamically.
Martin Kayondo, Jeonghwan Kang, Kyeongryong Lee, Donghyun Kwon, Yunheung Paek
IEEE Trans Autom. Sci. Eng.5
2024 Look Before You Access: Efficient Heap Memory Safety for Embedded Systems on ARMv8-M
abstract
Numerous embedded systems utilize firmware written in memory-unsafe C/C++. So, the firmware may exhibit spatial memory vulnerabilities, such as buffer overflows, which, if exploited by an attacker, can lead to various software attacks. While several studies have proposed defenses against these memory vulnerabilities, they often introduce significant performance and memory overhead or are impractical for application in embedded systems. In this paper, we introduce micro-fat pointer, a novel solution for heap memory safety for embedded systems. Notably, micro-fat pointer leverages TT instructions newly introduced in ARMv8-M to implement an efficient bounds-checking mechanism. Our evaluation results demonstrate that micro-fat pointer exhibits a 41% performance improvement compared to the existing state-of-the-art heap memory safety solution.
Jeonghwan Kang, Jaeyeol Park, Donghyun Kwon
DAC4
2023 Sfitag: Efficient Software Fault Isolation with Memory Tagging for ARM Kernel Extensions
abstract
As ARM is becoming more popular in today’s processor market, the OS kernel on ARM is gradually bloated to meet the market demand for more sophisticated services by absorbing diverse kernel extensions. Since this kernel bloating inevitably increases the attack surface, there has been a continuous effort to decrease the surface by dissociating or isolating untrusted extensions from the kernel. One approach in this effort is using software fault isolation (SFI) that instruments memory and control-transfer instructions to prevent isolated extensions from having unauthorized accesses to memory regions of the core kernel. Being implementable in pure software has been considered the greatest strength of SFI and thus popularly adopted by engineers to isolate kernel extensions, but software versions of SFI mostly suffer from high performance overhead, which can be a critical drawback for performance-sensitive mobile devices that overwhelmingly use ARM CPUs. The purpose of our work, named as Sfitag, is to make SFI for ARM kernel extensions more efficient by leveraging the hardware support from the latest ARM AArch64 architecture, called the ARM8.5-A memory tagging extension (MTE). For efficiency, Sfitag relies on MTE support when it allocates a tag value different from the core kernel for untrusted extensions and enforces extensions to use that value as a tag for pointers and memory objects. Consequently, in Sfitag, accessing the core kernel memory is legitimate only when the tag of a pointer matches the value of the kernel tag, which by means of MTE in effect enables us to safely confine unexpected and buggy behaviors of extensions within the space isolated from the kernel. Through our evaluation, we prove the effectiveness of Sfitag by showing that our MTE-supported SFI efficiently enforces isolation for extensions just with 1% slowdown on the throughput of a network driver and 5.7% on a block device driver.
Junseung You, Yungi Cho, Yeongpil Cho, Donghyun Kwon, Yunheung Paek
AsiaCCS5
2023 ZOMETAG: Zone-Based Memory Tagging for Fast, Deterministic Detection of Spatial Memory Violations on ARM
abstract
Against spatial memory violations threatening a vast amount of legacy software, various safety solutions have been suggested for decades. However, their practical uses have been impeded by diverse reasons, such as significant overheads and mandatory modifications of existing architectures. Accordingly, there has been a clear need for a practical safety solution that is fast enough and yet runs on commodity systems for its wide applicability in the field. As an effort to meet this need, a major processor vendor, ARM, recently announced a hardware extension, called Memory Tagging Extension (MTE), that helps engineers to implement efficient safety solutions. However, due to lack of hardware tags to isolate all data objects, MTE either resorts to a probabilistic memory safety guarantee, which is susceptible to a security loophole, or suffers from severe performance degradation to guarantee deterministic security. The aim of our work is to develop a MTE-based deterministic spatial safety solution, called ZOMETAG, with high efficiency by capitalizing on salient architectural features. Our key idea for fast, deterministic safety is to somehow assign permanently all objects unique tags throughout program execution. For this, ZOMETAG first divides the data memory into a number of small regions, called zones, and distributes data objects over the zones subject to certain constraints (to be discussed later). Then, we extend the notion of a tag in a way that each object stored with MTE tag$t$in zone$z$is uniquely assigned the zone-tag pair$z$,$t$> as a new tag. To work with this new tag assignment, we devise a novel mechanism, called two-layer isolation, that is basically a combination of MTE-based tagging (for one-layer of isolation) with zone-based tagging (for the other) both of which collaborate together to ensure spatial safety for all objects by preventing a pointer currently assigned one zone-tag pair from erroneously referring to objects assigned different pairs. Our experimental results are quite encouraging. ZOMETAG enforces deterministic spatial safety with overheads of 35% in SPEC CPU2006 and merely of 6% in real world applications like nginx.
Junseung You, Donghyun Kwon, Yeongpil Cho, Yunheung Paek
IEEE Trans. Inf. Forensics Secur.3
2023 Exploring effective uses of the tagged memory for reducing bounds checking overheads
Inyoung Bang, Yungi Cho, Jangseop Shin, Dongil Hwang, Donghyun Kwon, Yeongpil Cho, Yunheung Paek
J. Supercomput.6
2020 RIMI: Instruction-level Memory Isolation for Embedded Systems on RISC-V
abstract
With the advent of the Internet of Things, embedded systems have become widely used in various fields. Concurrently, the security of these systems has become a concern for many. However, security features that are already available for high-end systems have not been provided in low-end embedded systems due to its negative impact on cost and power consumption. Thus, to increase security with low overhead, many studies to implement the memory isolation approach to these systems have been conducted. However, existing techniques for this approach have suffered from problems in terms of scalability or performance. To mitigate such problems, we present RIMI, a new instruction extension to provide memory isolation in embedded systems. Thanks to instructions in RIMI, we can implement an instruction-level memory isolation where the access permission is bound to each memory and control transfer instructions. We implemented the RIMI prototype on a RISC-V architecture, which is a prominent open-source instruction set architecture (ISA). Our evaluation results show that existing security solutions, i.e., shadow stacks and in-process isolation, can be efficiently implemented with RIMI.
Haeyoung Kim, Jinjae Lee, Derry Pratama, Asep Muhamad Awaludin, Howon Kim 0001, Donghyun Kwon
ICCAD6
2020 PrOS: Light-Weight Privatized Se cure OSes in ARM TrustZone
abstract
TrustZone is a hardware security technique in ARM mobile devices. Using TrustZone, software components running within the secure world can be completely isolated from the normal world, which ensures hardware-enforced security access control over the underlying computing resources. In order to support multiple trusted applications, TrustZone runs its own operating system, called the secure OS, within the secure world. Unfortunately, attackers have been exploiting privilege escalation vulnerabilities in a secure OS, as reported in most of major secure OSes from product vendors including Samsung, Huawei, and Qualcomm. More critically, as all trusted applications are running on the same secure OS instance, compromising the secure OS leads to compromising all trusted applications, rendering the secure OS as a single point of failure endangering the entire TrustZone's security. This paper presents PrOS, our mechanism to privatize secure OSes through direct virtualization of TrustZone. PrOS allows each trusted application to run with its own secure OS such that the secure OS is no longer a single point of security failure. One particular challenge for PrOS lies in how efficiently to implement software-only virtualization for TrustZone for a practical deployment in real systems despite the condition that the current ARM architectures do not support hardware-assisted virtualization for TrustZone. As opposed to the common belief that software-only virtualization is inefficient and sluggish, we have found several common design features inherent in the secure OS to leverage for optimally tailoring the TrustZone virtualization scheme. We implemented PrOS on a 64-bit ARM development board. According to our evaluation, PrOS incurs 0.02 and 1.18 percent performance overheads on average in the normal and secure worlds, respectively, demonstrating its effectiveness in the field.
Donghyun Kwon, Yeongpil Cho, Byoungyoung Lee, Yunheung Paek
IEEE Trans. Mob. Comput.1
2019 RiskiM: Toward Complete Kernel Protection with Hardware Support
abstract
The OS kernel is typically the assumed trusted computing base in a system. Consequently, when they try to protect the kernel, developers often build their solutions in a separate secure execution environment externally located and protected by special hardware. Due to limited visibility into the host system, the external solutions basically all entail the semantic gap problem which can be easily exploited by an adversary to circumvent them. Thus, for complete kernel protection against such adversarial exploits, previous solutions resorted to aggressive techniques that usually come with various adverse side effects, such as high performance overhead, kernel code modifications and/or excessively complicated hardware designs. In this paper, we introduce RiskiM, our new hardware-based monitoring platform to ensure kernel integrity from outside the host system. To overcome the semantic gap problem, we have devised a hardware interface architecture, called PEMI, by which RiskiM is supplied with all internal states of the host system essential for fulfilling its monitoring task to protect the kernel even in the presence of attacks exploiting the semantic gap between the host and RiskiM. To empirically validate the security strength and performance of our monitoring platform in existing systems, we have fully implemented RiskiM in a RISC-V system. Our experiments show that RiskiM succeeds in the host kernel protection by detecting even the advanced attacks which could circumvent previous solutions, yet suffering from virtually no aforementioned side effects.
Dongil Hwang, Myonghoon Yang, Seongil Jeon, Younghan Lee 0001, Donghyun Kwon, Yunheung Paek
DATE5
2019 CRCount: Pointer Invalidation with Reference Counting to Mitigate Use-after-free in Legacy C/C++
Jangseop Shin, Donghyun Kwon, Yeongpil Cho, Yunheung Paek
NDSS2
2019 uXOM: Efficient eXecute-Only Memory on ARM Cortex-M
Donghyun Kwon, Jangseop Shin, Giyeol Kim, Byoungyoung Lee, Yeongpil Cho, Yunheung Paek
USENIX Security Symposium1
2019 Safe and Efficient Implementation of a Security System on ARM using Intra-level Privilege Separation
abstract
Security monitoring has long been considered as a fundamental mechanism to mitigate the damage of a security attack. Recently, intra-level security systems have been proposed that can efficiently and securely monitor system software without any involvement of more privileged entity. Unfortunately, there exists no full intra-level security system that can universally operate at any privilege level on ARM. However, as malware and attacks increase against virtually every level of privileged software including an OS, a hypervisor, and even the highest privileged software armored by TrustZone, we have been motivated to develop an intra-level security system, named Hilps . Hilps realizes true intra-level scheme in all these levels of privileged software on ARM by elaborately exploiting a new hardware feature of ARM’s latest 64-bit architecture, called TxSZ, that enables elastic adjustment of the accessible virtual address range. Furthermore, Hilps newly supports the sandbox mechanism that provides security tools with individually isolated execution environments, thereby minimizing security threats from untrusted security tools. We have implemented a prototype of Hilps on a real machine. The experimental results demonstrate that Hilps is quite promising for practical use in real deployments.
Donghyun Kwon, Hayoon Yi, Yeongpil Cho, Yunheung Paek
ACM Trans. Priv. Secur.1
2018 Hypernel: a hardware-assisted framework for kernel protection without nested paging
abstract
Large OS kernels always suffer from attacks due to their numerous inherent vulnerabilities. To protect the kernel, hypervisors have been employed by many security solutions. However, relying on a hypervisor has a detrimental impact on the system performance due mainly to nested paging. In this paper, we present Hypernel, a security framework combining hardware and software components to address this problem. Hypersec, the software component, provides an isolated execution environment for security solutions, and the hardware monitor component enables a word-granularity monitoring capability on the kernel memory. Our evaluation shows that Hypernel efficiently fulfills the role of a security framework, while imposing mere 3.1% of runtime overhead on the system.
Donghyun Kwon, Kuenwhee Oh, Junmo Park, Seungyong Yang, Yeongpil Cho, Brent ByungHoon Kang, Yunheung Paek
DAC1
2018 VM-CFI: Control-Flow Integrity for Virtual Machine Kernel Using Intel PT
Donghyun Kwon, Sehyun Baek, Giyeol Kim, Sunwoo Ahn, Yunheung Paek
ICCSA (5)1
2017 Instruction-Level Data Isolation for the Kernel on ARM
abstract
As more sophisticated services are increasingly offered by the OS kernel on mobile devices, the security and sensitivity of kernel data that they depend on are becoming a critical issue. Data isolation has emerged as a key technique that can address the issue by providing strong protection for sensitive kernel data. However, existing data isolation mechanisms for mobile devices all incur non-negligible performance overhead. We deem that such computational burden would be a serious problem for mobile devices which already suffer from resource poverty. To alleviate this problem, we have developed a new mechanism that enforces data isolation very efficiently on ARM-based machines backed by unique hardware instructions. For evaluation, this instruction-level data isolation mechanism has been implemented in the Android/Linux kernel running on ARM. According to the experiment, it provides a lightweight data isolation capability for security services installed in the kernel.
Yeongpil Cho, Donghyun Kwon, Yunheung Paek
DAC2
2017 Dynamic Virtual Address Range Adjustment for Intra-Level Privilege Separation on ARM
Yeongpil Cho, Donghyun Kwon, Hayoon Yi, Yunheung Paek
NDSS2
2017 Optimization techniques to enable execution offloading for 3D video games
Donghyun Kwon, Seungjun Yang, Yunheung Paek, Kwangman Ko
Multim. Tools Appl.1
2016 Hardware-Assisted On-Demand Hypervisor Activation for Efficient Security Critical Code Execution on Mobile Devices
Yeongpil Cho, Jun-Bum Shin, Donghyun Kwon, MyungJoo Ham, Yuna Kim, Yunheung Paek
USENIX ATC3
2016 Precise execution offloading for applications with dynamic behavior in mobile cloud computing
Yongin Kwon, Hayoon Yi, Donghyun Kwon, Seungjun Yang, Yeongpil Cho, Yunheung Paek
Pervasive Mob. Comput.3
2015 Mantis: Efficient Predictions of Execution Time, Energy Usage, Memory Usage and Network Usage on Smart Mobile Devices
abstract
We present Mantis, a framework for predicting the computational resource consumption (CRC) of Android applications on given inputs accurately, and efficiently. A key insight underlying Mantis is that program codes often contain features that correlate with performance and these features can be automatically computed efficiently. Mantis synergistically combines techniques from program analysis and machine learning. It constructs concise CRC models by choosing from many program execution features only a handful that are most correlated with the program's CRC metric yet can be evaluated efficiently from the program's input. We apply program slicing to reduce evaluation time of a feature and automatically generate executable code snippets for efficiently evaluating features. Our evaluation shows that Mantis predicts four CRC metrics of seven Android apps with estimation error in the range of 0-11.1 percent by executing predictor code spending at most 1.3 percent of their execution time on Galaxy Nexus.
Yongin Kwon, Hayoon Yi, Donghyun Kwon, Seungjun Yang, Byung-Gon Chun, Ling Huang 0001, Petros Maniatis, Mayur Naik, Yunheung Paek
IEEE Trans. Mob. Comput.4
2014 Techniques to Minimize State Transfer Costs for Dynamic Execution Offloading in Mobile Cloud Computing
abstract
In order to meet the increasing demand for high performance in smartphones, recent studies suggested mobile cloud computing techniques that aim to connect the phones to adjacent powerful cloud servers to throw their computational burden to the servers. These techniques often employ execution offloading schemes that migrate a process between machines during its execution. In execution offloading, code regions to be executed on the server are decided statically or dynamically based on the complex analysis on execution time and process state transfer costs of every region. Expectedly, the transfer cost is a deciding factor for the success of execution offloading. According to our analysis, it is dominated by the total size of heap objects transferred over the network. But previous work did not try hard to minimize this size. Thus in this paper, we introduce novel techniques based on compiler code analysis that effectively reduce the transferred data size by transferring only the essential heap objects and the stack frames actually referenced in the server. The experiments exhibit that the reduced size positively influences not only the transfer time itself but also the overall effectiveness of execution offloading, and ultimately, improves the performance of our mobile cloud computing significantly in terms of execution time and energy consumption.
Seungjun Yang, Donghyun Kwon, Hayoon Yi, Yeongpil Cho, Yongin Kwon, Yunheung Paek
IEEE Trans. Mob. Comput.2
2013 Fast dynamic execution offloading for efficient mobile cloud computing
abstract
In order to meet the increasing demand for high performance in smartphones, recent studies suggested mobile cloud computing techniques that aim to connect the phones to adjacent powerful cloud servers to throw their computational burden to the servers. These techniques often employ execution offloading schemes that migrate a process between machines during its execution. In execution offloading, code regions to be executed on the server are decided statically or dynamically based on the complex analysis on execution time and process state transfer time of every region. Expectedly, the transfer time is a deciding factor for the success of execution offloading. According to our analysis, it is dominated by the total size of heap objects transferred over the network. But previous work did not try hard to minimize this size. Thus in this paper, we introduce novel techniques based on compiler code analysis that effectively reduce the transferred data size by transferring only the essential heap objects. The experiments exhibit that the reduced size positively influences not only the transfer time itself but also the overall effectiveness of execution offloading, and ultimately, improves the performance of our mobile cloud computing significantly in terms of execution time and power consumption.
Seungjun Yang, Yongin Kwon, Yeongpil Cho, Hayoon Yi, Donghyun Kwon, Jonghee M. Youn, Yunheung Paek
PerCom5
2013 Mantis: Automatic Performance Prediction for Smartphone Applications
Yongin Kwon, Hayoon Yi, Donghyun Kwon, Seungjun Yang, Byung-Gon Chun, Ling Huang 0001, Petros Maniatis, Mayur Naik, Yunheung Paek
USENIX ATC4