VLDB 2026 Research / reviewers in the wild / expert
Yunheung Paek
dblp:65/3751
· DBLP profile ↗
118ranked-venue papers
5as first author
26since 2021 · last 2026
0000-0002-6412-2926ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 67 · 1 first-author · 10 since 2021Security and privacy · 26 · 14 since 2021Software engineering, systems software and programming languages · 24 · 3 first-author · 3 since 2021Computer networks · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HEPIC: Private Inference over Homomorphic Encryption with Client InterventionabstractHomomorphic Encryption (HE) enables Private Inference (PI) in Machine Learning as a Service (MLaaS), protecting both client inputs and server-side neural network (NN) parameters. Existing PI techniques are predominantly implemented as either HE-based fire-and-forget methods or MPC-based interactive methods. Recent HE-based PI systems improve the accuracy--performance trade-off via a layer-wise scheme and parameter switching, yet remain bottlenecked by fire-and-forget execution in which the server alone performs costly ciphertext management (e.g., bootstrapping and scheme/parameter conversions). We present HEPIC, an HE-based PI system that explores a different design point by leveraging client interventions for ciphertext managements. In a sense, HEPIC shares a common ground with MPC-based PI of being interactive with the client, but differs in that the client only intervenes for ciphertext managements required in HE operations. Because ciphertext management has identical semantics on the client and the server, HEPIC lets developers decide where and how often to execute it, enabling fine-grained trade-offs among computation, communication, and ciphertext configuration. HEPIC makes such execution practical by overlapping client re-encryption, server computation, and communication via dependency-aware pipelining and streaming-based transfers. We further enhance the performance with a cache-aware task allocator (CATA) and a cost-aware client intervention scheduler (CACIS) to exploit ciphertext-level parallelism and to mitigate stalls under client-server performance disparity. Our evaluation shows that HEPIC achieves up to 2.20--41.93× speedup over state-of-the-art fire-and-forget HE-based PI, while maintaining zero loss in inference accuracy. Kevin Nam, Youyeon Joo, Seungjin Ha, Hyungon Moon, Yunheung Paek |
ASPLOS (2) | 5 |
| 2026 | SECV: Securing Connected Vehicles with Hardware Trust Anchors
Martin Kayondo, Junseung You, Eunmin Kim, Yunheung Paek |
NDSS | 5 |
| 2026 | Towards an efficient dataflow-flexible accelerator by finding optimal dataflows of DNNsabstract• Leverage design space exploration tool to create an efficient accelerator for DNNs. • Reuse redundant hardware components to minimize chip area. • Schedule to minimize dataflow transitions to maximize efficiency. This paper proposes a new dataflow-flexible accelerator design that addresses the limitations of existing heterogeneous dataflow accelerator (HDA) for handling the computation of multiple deep neural network (DNN) models. The design offers increased dataflow flexibility and higher efficiency compared to existing works. The accelerator utilizes a fixed set of representative dataflows implemented as operating modes and switches between them dynamically. A design space exploration (DSE) tool is leveraged to evaluate the efficiency of candidate dataflows and determine the optimal number and types of operating modes. Each layer of the target DNN models is assessed with different operating modes to select the optimal mode for each layer. Also, two supplementary optimization techniques are adopted to reduce the overheads from supporting a multitude of dataflows. One optimizes to minimize the number of transitions of dataflows, which incur severe overheads. The other optimizes to maximize the reuse of hardware components associated with supporting multiple dataflows. By identifying the redundant hardware components, the proposed design minimizes the chip area, another aspect where dataflow-flexible accelerators suffer. Experimental results demonstrate that our algorithm achieves greater dataflow flexibility with high efficiency, Compared to HDA, our design is, on average, 34.6 % lower in latency at the cost of 6.4 % area and negligible energy overhead. Whoi Ree Ha, Dongju Lee, Deumji Woo, Jonghee Yoon, Yongin Kwon, Yunheung Paek |
Future Gener. Comput. Syst. | 10 |
| 2025 | Affinity-based Optimizations for TFHE on Processing-in-DRAMabstractProcessing-in-memory (PIM) architectures are promising for accelerating intensive workloads due to their high internal bandwidth. This paper introduces a technique for accelerating Fully Homomorphic Encryption over the Torus (TFHE), a promising yet intensive application, on a realistic PIM system. Existing TFHE accelerators focus on exploiting parallelism, often overlooking data affinity, which leads to performance degradation in PIM due to excessive remote data accesses (RDAs). To address this, we present an affinity-based approach that optimizes the computation of TFHE on PIM. We apply algorithmic optimizations to TFHE, enabling PIM to effectively leverage its high internal bandwidth. We analyze the affinity patterns in the sub-tasks of TFHE and develop an offline scheduler that exploits our analysis to find optimal scheduling, minimizing RDAs while maintaining sufficient parallelism. To demonstrate the practicality of our work, we design a variant of an existing PIM-HBM device with minimal hardware modifications, and perform evaluations over a real FPGA-based PIM system. Our experiments demonstrate that our affinity-based optimizations outperform prior TFHE accelerators by 4.24-209× for real-world benchmarks. Kevin Nam, Heon Hui Jung, Hyunyoung Oh, Yunheung Paek |
ASPLOS (2) | 4 |
| 2025 | BASTAG: Byte-level Access Control on Shared Memory using ARM Memory Tagging ExtensionabstractAs software grows in size and complexity, modular designs are increasingly adopted, leading to frequent interactions via shared memory between components. This design however increases the risk of vulnerabilities from uncontrolled memory access to shared memory. Enforcing byte-level access control can mitigate these risks by enabling byte-level permissions on complex shared objects and their sub-elements. However, existing approaches face performance limitations as they increase the granularity of control to byte level. In this paper, we present BASTAG, a novel system that leverages ARM's Memory Tagging Extension (MTE) to tack this challenge. Although MTE enforces tag-matching between pointers and memory, its hardware-defined granularity is too coarse to support byte-level control on its own. To address the inherent limitations of applying MTE for nuanced access control, BASTAG incorporates a technique known as shadow memory tagging that places separate, but associated MTE tags for the actual memory targets, allowing for more flexible and finer access control with efficiency. We implemented a BASTAG prototype on AArch64 hardware with MTE support and evaluated it on three real-world use cases. Our results demonstrate that BASTAG significantly outperforms existing byte-level access control mechanisms. Junseung You, Kyeongryong Lee, Yeongpil Cho, Yunheung Paek |
CCS | 5 |
| 2025 | An Accelerator for Low-Computational Overhead Privacy-Preserving GNN InferenceabstractGraph Neural Networks (GNNs) are increasingly used in domains such as finance and bioinformatics, where both node features and edge structures can contain sensitive information. While Fully Homomorphic Encryption (FHE) offers a promising solution for privacy-preserving GNN inference, existing approaches such as PPGNN rely on costly Homomorphic Rotation and MUX operations for operand obfuscation, resulting in significant computational overhead. In this work, we propose a new obfuscation method that leverages the probabilistic nature of FHE to duplicate ciphertexts at the client side, thereby eliminating the need for runtime selection logic. To support this method efficiently, we design a pipelined hardware accelerator with a simplified CKKS datapath and parallel TFHE execution, avoiding the complexity of rotation-heavy designs. Despite reduced ciphertext reuse, our architecture mitigates memory pressure through buffer-aware PBS unit design. Experimental results demonstrate up to$8.8 \times$speedup and$7.69 \times$energy efficiency improvement over PPGNN, while also outperforming existing multi-scheme accelerators such as Trinity and UFC even when applying the same obfuscation strategy. Our approach offers a practical and scalable solution for efficient, privacy-preserving GNN inference. Heon Hui Jung, Whoi Ree Ha, Kevin Nam, Youyeon Joo, Lucas Oros, Yunheung Paek |
HiPC | 6 |
| 2025 | Unified MEDS Accelerator
Sanjay Deshpande, Mamuri Nawan, Kashif Nawaz, Ruben Niederhagen, Yunheung Paek, Jakub Szefer |
SAC | 6 |
| 2025 | SLOTHE : Lazy Approximation of Non-Arithmetic Neural Network Functions over Encrypted Data
Kevin Nam, Youyeon Joo, Seungjin Ha, Yunheung Paek |
USENIX Security Symposium | 4 |
| 2025 | LOHEN: Layer-wise Optimizations for Neural Network Inferences over Encrypted Data with High Performance or Accuracy
Kevin Nam, Youyeon Joo, Dongju Lee, Seungjin Ha, Hyunyoung Oh, Hyungon Moon, Yunheung Paek |
USENIX Security Symposium | 7 |
| 2025 | ROSec: Intra-Process Isolation for ROS Composition With Memory Protection KeysabstractRobot Operating System(ROS) is a software framework for robotic systems that includes various packages for developing robotic applications.Compositionis a package that combines multiple applications, namely,nodes, to be loaded and executed in a single process. However, permitting multiple nodes to share the address space could expand the attack surface such that vulnerabilities in a node are more likely to be exploited to subvert nodes running in the same space. We propose ROSec, an in-process isolation solution for ROS composition that utilizes Intel Memory Protection Keys. ROSecaims to enforce memory isolation between nodes within a process by preventing unauthorized access from one node to another. Unlike previous works that assume the number and sizes of nodes are statically defined and partitioned by developers, ROSecis designed to handle the dynamic nature of nodes that can be loaded and executed in a process at any time during execution. To achieve this, ROSecadopts a unique scheduling mechanism that utilizes theexecutor-centricexecution model of ROS to perform two main operations for MPK-based isolation:protection key assignmentandreassignment. Our evaluation shows that ROSeceffectively enforces in-process isolation while incurring a 6.4% performance overhead on a real-world application.Note to Practitioners—Cyber-Physical Systems are the core of modern applications, particularly robotics, as they integrate computing and physical processes. ROS necessitates real-time and security guarantees, which, unfortunately, trade-off with each other. While traditional ROS architecture relies on process isolation to separate various nodes, ROS2 introduces a feature called composition, which allows multiple nodes to run inside a single process, thus exposing various nodes to potential malicious compromises from others. This paper proposes a technique that utilizes Intel memory protection keys (MPK) to provide intraprocess isolation for ROS composition. Given that ROS nodes are dynamic, ROSEC provides key assignment and reassignment techniques to configure MPK dynamically. Martin Kayondo, Jeonghwan Kang, Kyeongryong Lee, Donghyun Kwon, Yunheung Paek |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | SPHINCSLET: An Area-Efficient Accelerator for the Full SPHINCS+ Digital Signature AlgorithmabstractThis work presents SPHINCSLET, the first fully standard-compliant and area-efficient hardware implementation of the SLH-DSA algorithm, formerly known as SPHINCS+, a post-quantum digital signature scheme. SPHINCSLET is designed to be parameterizable across different security levels and hash functions, offering a balanced tradeoff between area efficiency and performance. Existing hardware implementations either feature a large area footprint to achieve fast signing and verification or adopt a coprocessor-based approach that significantly slows down these operations. SPHINCSLET addresses this gap by delivering a 4.7× reduction in area compared to high-speed designs while achieving a 2.5× to 5× improvement in signing time over the most efficient coprocessor-based designs for a SHAKE256-based SPHINCS+ implementation. The SHAKE256-based SPHINCS+ FPGA implementation targeting the AMD Artix-7 requires fewer than 10.8K LUTs for any security level of SLH-DSA. Furthermore, the SHA-2-based SPHINCS+ implementation achieves a 2× to 4× speedup in signature generation across various security levels compared to existing SLH-DSA hardware, all while maintaining a compact area footprint of 6K to 15K LUTs. This makes it the fastest SHA-2-based SLH-DSA implementation to date. With an optimized balance of area and performance, SPHINCSLET can assist resource-constrained devices in transitioning to post-quantum cryptography. Sanjay Deshpande, Cansu Karakuzu Aslan, Jakub Szefer, Yunheung Paek |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2024 | VFLIP: A Backdoor Defense for Vertical Federated Learning via Identification and Purification
Yungi Cho, Woorim Han, Miseon Yu, Younghan Lee 0001, Ho Bae, Yunheung Paek |
ESORICS (4) | 6 |
| 2024 | MetaSafe: Compiling for Protecting Smart Pointer Metadata to Ensure Safe Rust Integrity
Martin Kayondo, Inyoung Bang, Yeongjun Kwak, Hyungon Moon, Yunheung Paek |
USENIX Security Symposium | 5 |
| 2023 | Sfitag: Efficient Software Fault Isolation with Memory Tagging for ARM Kernel ExtensionsabstractAs ARM is becoming more popular in today’s processor market, the OS kernel on ARM is gradually bloated to meet the market demand for more sophisticated services by absorbing diverse kernel extensions. Since this kernel bloating inevitably increases the attack surface, there has been a continuous effort to decrease the surface by dissociating or isolating untrusted extensions from the kernel. One approach in this effort is using software fault isolation (SFI) that instruments memory and control-transfer instructions to prevent isolated extensions from having unauthorized accesses to memory regions of the core kernel. Being implementable in pure software has been considered the greatest strength of SFI and thus popularly adopted by engineers to isolate kernel extensions, but software versions of SFI mostly suffer from high performance overhead, which can be a critical drawback for performance-sensitive mobile devices that overwhelmingly use ARM CPUs. The purpose of our work, named as Sfitag, is to make SFI for ARM kernel extensions more efficient by leveraging the hardware support from the latest ARM AArch64 architecture, called the ARM8.5-A memory tagging extension (MTE). For efficiency, Sfitag relies on MTE support when it allocates a tag value different from the core kernel for untrusted extensions and enforces extensions to use that value as a tag for pointers and memory objects. Consequently, in Sfitag, accessing the core kernel memory is legitimate only when the tag of a pointer matches the value of the kernel tag, which by means of MTE in effect enables us to safely confine unexpected and buggy behaviors of extensions within the space isolated from the kernel. Through our evaluation, we prove the effectiveness of Sfitag by showing that our MTE-supported SFI efficiently enforces isolation for extensions just with 1% slowdown on the throughput of a network driver and 5.7% on a block device driver. Junseung You, Yungi Cho, Yeongpil Cho, Donghyun Kwon, Yunheung Paek |
AsiaCCS | 6 |
| 2023 | KVSEV: A Secure In-Memory Key-Value Store with Secure Encrypted VirtualizationabstractAMD's Secure Encrypted Virtualization (SEV) is a hardware-based Trusted Execution Environment (TEE) designed to secure tenants' data on the cloud, even against insider threats. The latest version of SEV, SEV-Secure Nested Paging (SEV-SNP), offers protection against most well-known attacks such as cold boot and hypervisor-based attacks. However, it remains susceptible to a specific type of attack known as Active DRAM Corruption (ADC), where attackers manipulate memory content using specially crafted memory devices. The in-memory key-value store (KVS) on SEV is a prime target for ADC attacks due to its critical role in cloud infrastructure and the predictability of its data structures. To counter this threat, we propose KVSEV, an in-memory KVS resilient to ADC attacks. KVSEV leverages SNP's Virtual Machine Management (VMM) and attestation mechanism to protect the integrity of key-value pairs, thereby securing the KVS from ADC attacks. Our evaluation shows that KVSEV secures in-memory KVSs on SEV with a performance overhead comparable to other secure in-memory KVS solutions. Junseung You, Kyeongryong Lee, Hyungon Moon, Yeongpil Cho, Yunheung Paek |
SoCC | 5 |
| 2023 | FLGuard: Byzantine-Robust Federated Learning via Ensemble of Contrastive Models
Younghan Lee 0001, Yungi Cho, Woorim Han, Ho Bae, Yunheung Paek |
ESORICS (4) | 5 |
| 2023 | Exploring Clustered Federated Learning's Vulnerability against Property Inference AttackabstractClustered federated learning (CFL) is an advanced technique in the field of federated learning (FL) that addresses the issue of catastrophic forgetting caused by non-independent and identically distributed (non-IID) datasets. CFL achieves this by clustering clients based on the similarity of their datasets and training a global model for each cluster. Despite the effectiveness of CFL in mitigating performance degradation resulting from non-IID datasets, the potential risk of privacy leakages in CFL has not been thoroughly studied. Previous work evaluated the risk of privacy leakages in FL using the property inference attack (PIA), which extracts information about unintended properties (i.e., attributes that differ from the target attribute of the global model’s main task). In this paper, we explore the potential risk of unintended property leakage in CFL by subjecting it to both passive and active PIAs. Our empirical analysis shows that the passive PIA performance on CFL is substantially better than that on FL in terms of the attack AUC score. Moreover, we propose an enhanced active PIA method tailored for CFL to improve the attack performance. Our method introduces a scale-up parameter that amplifies the impact of malicious local updates, resulting in better performance than the previous technique. Furthermore, we demonstrate that the vulnerability of CFL can be alleviated by applying differential privacy (DP) mechanisms at the client-level. Unlike previous works, which have shown that applying DP to FL can induce a high utility loss, our empirical results indicate that DP can be used as a defense mechanism in CFL, leading to a better trade-off between privacy and utility. Yungi Cho, Younghan Lee 0001, Ho Bae, Yunheung Paek |
RAID | 5 |
| 2023 | TRust: A Compilation Framework for In-process Isolation to Protect Safe Rust against Untrusted Code
Inyoung Bang, Martin Kayondo, Hyungon Moon, Yunheung Paek |
USENIX Security Symposium | 4 |
| 2023 | ZOMETAG: Zone-Based Memory Tagging for Fast, Deterministic Detection of Spatial Memory Violations on ARMabstractAgainst spatial memory violations threatening a vast amount of legacy software, various safety solutions have been suggested for decades. However, their practical uses have been impeded by diverse reasons, such as significant overheads and mandatory modifications of existing architectures. Accordingly, there has been a clear need for a practical safety solution that is fast enough and yet runs on commodity systems for its wide applicability in the field. As an effort to meet this need, a major processor vendor, ARM, recently announced a hardware extension, called Memory Tagging Extension (MTE), that helps engineers to implement efficient safety solutions. However, due to lack of hardware tags to isolate all data objects, MTE either resorts to a probabilistic memory safety guarantee, which is susceptible to a security loophole, or suffers from severe performance degradation to guarantee deterministic security. The aim of our work is to develop a MTE-based deterministic spatial safety solution, called ZOMETAG, with high efficiency by capitalizing on salient architectural features. Our key idea for fast, deterministic safety is to somehow assign permanently all objects unique tags throughout program execution. For this, ZOMETAG first divides the data memory into a number of small regions, called zones, and distributes data objects over the zones subject to certain constraints (to be discussed later). Then, we extend the notion of a tag in a way that each object stored with MTE tag$t$in zone$z$is uniquely assigned the zone-tag pair$z$,$t$> as a new tag. To work with this new tag assignment, we devise a novel mechanism, called two-layer isolation, that is basically a combination of MTE-based tagging (for one-layer of isolation) with zone-based tagging (for the other) both of which collaborate together to ensure spatial safety for all objects by preventing a pointer currently assigned one zone-tag pair from erroneously referring to objects assigned different pairs. Our experimental results are quite encouraging. ZOMETAG enforces deterministic spatial safety with overheads of 35% in SPEC CPU2006 and merely of 6% in real world applications like nginx. Junseung You, Donghyun Kwon, Yeongpil Cho, Yunheung Paek |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Exploring effective uses of the tagged memory for reducing bounds checking overheads
Inyoung Bang, Yungi Cho, Jangseop Shin, Dongil Hwang, Donghyun Kwon, Yeongpil Cho, Yunheung Paek |
J. Supercomput. | 8 |
| 2023 | Ambassy: A Runtime Framework to Delegate Trusted Applications in an ARM/FPGA Hybrid SystemabstractMany mobile systems run on ARM-based devices today. People use these for increasingly diverse yet security-sensitive applications. ARM has adopted a security model to tackle this threat, where they manage private information in an isolatedtrusted execution environment(TEE) provided byTrustZone. This TrustZone-based model has been proven effective, but due to security concerns, it is available solely for the vendor's applications, thereby hindering the broad use of TrustZone. Consequently, we propose a runtime framework backed by TrustZone to construct a secondary TEE.Ambassyhas its residence built on an on-chip field-programmable gate array (FPGA), which is a standard component in anARM/FPGA hybridsystem readily available on the market today. This study, to the best of our knowledge, is the first attempt to broaden the use of TrustZone by using an FPGA to build a secondary TEE for arbitrary third-parties, which otherwise should be expelled to the Normal World. This paper describes many design challenges that we have overcome to fully implementAmbassyon an FPGA. Our experiments demonstrate the practicality ofAmbassyby presenting the security analysis and performance results of third-party application samples. The samples all run safely onAmbassy, with shorter execution times than regular TEE applications in TrustZone (by a factor of 5.5–52). Dongil Hwang, Sanzhar Yeleuov, Minu Chung, Hyungon Moon, Yunheung Paek |
IEEE Trans. Mob. Comput. | 6 |
| 2022 | Practical Binary Code Similarity Detection with BERT-based Transferable Similarity LearningabstractBinary code similarity detection (BCSD) serves as a basis for a wide spectrum of applications, including software plagiarism, malware classification, and known vulnerability discovery. However, the inference of contextual meanings of a binary is challenging due to the absence of semantic information available in source codes. Recent advances leverage the benefits of a deep learning architecture into a better understanding of underlying code semantics and the advantages of the Siamese architecture into better BCSD. Sunwoo Ahn, Seonggwan Ahn, Hyungjoon Koo, Yunheung Paek |
ACSAC | 4 |
| 2022 | XTENSTORE: Fast Shielded In-memory Key-Value Store on a Hybrid x86-FPGA SystemabstractWe propose XtenStore, a system that extends the existing SGX-based secure in-memory key-value store with an external hardware accelerator in order to ensure comparable security guarantees with lower performance degradation. The accelerator is implemented on a commodity FPGA card that is readily connected with the x86 CPU via PCIe interconnect to form a hybrid x86-FPGA system. In comparison to the prior SGX-based work, XtenStore improves the throughput by 4–33x, and exhibits considerably shorter tail latency (>23x, 99th-percentile). Hyunyoung Oh, Dongil Hwang, Maja Malenko, Myunghyun Cho, Hyungon Moon, Marcel Baunach, Yunheung Paek |
DATE | 7 |
| 2022 | Precise Extraction of Deep Learning Models via Side-Channel Attacks on Edge/Endpoint Devices
Younghan Lee 0001, Sohee Jun, Yungi Cho, Woorim Han, Hyungon Moon, Yunheung Paek |
ESORICS (3) | 6 |
| 2022 | Accelerating N-Bit Operations over TFHE on Commodity CPU-FPGAabstractTFHE is a fully homomorphic encryption (FHE) scheme that evaluates Boolean gates, which we will hereafter call Tgates, over encrypted data. TFHE is considered to have higher expressive power than many existing schemes in that it is able to compute not only N-bit Arithmetic operations but also Logical/Relational ones as arbitrary ALR operations can be represented by Tgate circuits. Despite such strength, TFHE has a weakness that like all other schemes, it suffers from colossal computational overhead. Incessant efforts to reduce the overhead have been made by exploiting the inherent parallelism of FHE operations on ciphertexts. Unlike other FHE schemes, the parallelism of TFHE can be decomposed into multilayers: one inside each FHE operation (equivalent to a single Tgate) and the other between Tgates. Unfortunately, previous works focused only on exploiting the parallelism inside Tgate. However, as each N-bit operation over TFHE corresponds to a Tgate circuit constructed from multiple Tgates, it is also necessary to utilize the parallelism between Tgates for optimizing an entire operation. This paper proposes an acceleration technique to maximize performance of a TFHE N-bit operation by simultaneously utilizing both parallelism comprising the operation. To fully profit from both layers of parallelism, we have implemented our technique on a commodity CPU-FPGA hybrid machine with parallel execution capabilities in hardware. Our implementation outperforms prior ones by 2.43× in throughput and 12.19× in throughput per watt when performing N-bit operations under the 128-bit quantum security parameters. Kevin Nam, Hyunyoung Oh, Hyungon Moon, Yunheung Paek |
ICCAD | 4 |
| 2021 | A metadata-driven approach to efficiently detect code-reuse attacks on ARM multiprocessors
Hyunyoung Oh, Yeongpil Cho, Yunheung Paek |
J. Supercomput. | 3 |
| 2020 | TRUSTORE: Side-Channel Resistant Storage for SGX using Intel Hybrid CPU-FPGAabstractIntel SGX is a security solution promising strong and practical security guarantees for trusted computing. However, recent reports demonstrated that such security guarantees of SGX are broken due to access pattern based side-channel attacks, including page fault, cache, branch prediction, and speculative execution. In order to stop these side-channel attackers, Oblivious RAM (ORAM) has gained strong attention from the security community as it provides cryptographically proven protection against access pattern based side-channels. While several proposed systems have successfully applied ORAM to thwart side-channels, those are severely limited in performance and its scalability due to notorious performance issues of ORAM. This paper presents TrustOre, addressing these issues that arise when using ORAM with Intel SGX. TrustOre leverages an external device, FPGA, to implement a trusted storage service within a completed isolated environment secure from side-channel attacks. TrustOre tackles several challenges in achieving such a goal: extending trust from SGX to FPGA without imposing architectural changes, providing a verifiably-secure connection between SGX applications and FPGA, and seamlessly supporting various access operations from SGX applications to FPGA.We implemented TrustOre on the commodity Intel Hybrid CPU-FPGA architecture. Then we evaluated with three state-of-the-art ORAM-based SGX applications, ZeroTrace, Obliviate, and Obfuscuro, as well as an end-to-end key-value store application. According to our evaluation, TrustOre-based applications outperforms ORAM-based original applications ranging from 10x to 43x, while also showing far better scalability than ORAM-based ones. We emphasize that since TrustOre can be deployed as a simple plug-in to SGX machine's PCIe slot, it is readily used to thwart side-channel attacks in SGX, arguably one of the most cryptic and critical security holes today. Hyunyoung Oh, Adil Ahmad, Seonghyun Park 0001, Byoungyoung Lee, Yunheung Paek |
CCS | 5 |
| 2020 | Hawkware: Network Intrusion Detection based on Behavior Analysis with ANNs on an IoT DeviceabstractThe network-based Intrusion detection system (NIDS) plays a key role in Internet of Things (IoT) as most IoT services are network-driven. However, the existing NIDSes for IoT systems are either too costly to scale or vulnerable against advanced attacks such as traffic mimicry. In this paper, we propose a novel IDS named Hawkware, a lightweight ANN-based distributed NIDS that runs on an IoT device and analyzes the device's runtime behavior in tandem with its network traffic. By analyzing device behavior, Hawkware is able to replace expensive, deep data analysis that has traditionally been used to detect advanced attacks. Our evaluations show that Hawkware is lightweight enough to be distributed and deployed on a Raspberry PI, and yet capable of detecting such attacks at a satisfactory level. Sunwoo Ahn, Hayoon Yi, Younghan Lee 0001, Whoi Ree Ha, Giyeol Kim, Yunheung Paek |
DAC | 6 |
| 2020 | PrOS: Light-Weight Privatized Se cure OSes in ARM TrustZoneabstractTrustZone is a hardware security technique in ARM mobile devices. Using TrustZone, software components running within the secure world can be completely isolated from the normal world, which ensures hardware-enforced security access control over the underlying computing resources. In order to support multiple trusted applications, TrustZone runs its own operating system, called the secure OS, within the secure world. Unfortunately, attackers have been exploiting privilege escalation vulnerabilities in a secure OS, as reported in most of major secure OSes from product vendors including Samsung, Huawei, and Qualcomm. More critically, as all trusted applications are running on the same secure OS instance, compromising the secure OS leads to compromising all trusted applications, rendering the secure OS as a single point of failure endangering the entire TrustZone's security. This paper presents PrOS, our mechanism to privatize secure OSes through direct virtualization of TrustZone. PrOS allows each trusted application to run with its own secure OS such that the secure OS is no longer a single point of security failure. One particular challenge for PrOS lies in how efficiently to implement software-only virtualization for TrustZone for a practical deployment in real systems despite the condition that the current ARM architectures do not support hardware-assisted virtualization for TrustZone. As opposed to the common belief that software-only virtualization is inefficient and sluggish, we have found several common design features inherent in the secure OS to leverage for optimally tailoring the TrustZone virtualization scheme. We implemented PrOS on a 64-bit ARM development board. According to our evaluation, PrOS incurs 0.02 and 1.18 percent performance overheads on average in the normal and secure worlds, respectively, demonstrating its effectiveness in the field. Donghyun Kwon, Yeongpil Cho, Byoungyoung Lee, Yunheung Paek |
IEEE Trans. Mob. Comput. | 5 |
| 2019 | RiskiM: Toward Complete Kernel Protection with Hardware SupportabstractThe OS kernel is typically the assumed trusted computing base in a system. Consequently, when they try to protect the kernel, developers often build their solutions in a separate secure execution environment externally located and protected by special hardware. Due to limited visibility into the host system, the external solutions basically all entail the semantic gap problem which can be easily exploited by an adversary to circumvent them. Thus, for complete kernel protection against such adversarial exploits, previous solutions resorted to aggressive techniques that usually come with various adverse side effects, such as high performance overhead, kernel code modifications and/or excessively complicated hardware designs. In this paper, we introduce RiskiM, our new hardware-based monitoring platform to ensure kernel integrity from outside the host system. To overcome the semantic gap problem, we have devised a hardware interface architecture, called PEMI, by which RiskiM is supplied with all internal states of the host system essential for fulfilling its monitoring task to protect the kernel even in the presence of attacks exploiting the semantic gap between the host and RiskiM. To empirically validate the security strength and performance of our monitoring platform in existing systems, we have fully implemented RiskiM in a RISC-V system. Our experiments show that RiskiM succeeds in the host kernel protection by detecting even the advanced attacks which could circumvent previous solutions, yet suffering from virtually no aforementioned side effects. Dongil Hwang, Myonghoon Yang, Seongil Jeon, Younghan Lee 0001, Donghyun Kwon, Yunheung Paek |
DATE | 6 |
| 2019 | Real-Time Anomalous Branch Behavior Inference with a GPU-inspired Engine for Machine Learning ModelsabstractAttacks on embedded devices are likely to occur any time in unexpected manners. Thus, the defense systems based on fixed sets of rules will easily be subverted by such unexpected, unknown attacks. Learning-based anomaly detection may potentially prevent new unknown zero-day attacks by leveraging the capability of machine learning (ML) to learn the intricate true nature of software hidden within raw information. This paper introduces our work to develop an MPSoC, called RTAD, which can efficiently support in hardware various ML models that run to detect anomalous behaviors on embedded devices in a real-time fashion, and thus enable the devices to counteract the anomalies in the field. In the IoT era, the importance of security for embedded devices cannot be exaggerated because they will become an enticing target for adversaries as they are being integrated into everyday life to provide users with various services. The above-mentioned potential of learning-based detection is believed to benefit those deployed devices under attacks occurring any time during their field operations in unexpected manners. We hereby assume that ML models are trained with runtime branch information as their data features since a sequence of branches serves as a record of control flow transfers during program execution. In fact, there have been numerous ML studies that examine various types of branches in order to infer (or detect) anomaly in branch behaviors that may be induced by diverse attacks that can cause deviant control flow in software. Our goal of real-time anomalous branch behavior inference poses two challenges to our development of RTAD. Firstly, RTAD must collect and transfer in a timely fashion a sequence of branches as the input to the ML model. Secondly, RTAD must be able to promptly process the delivered branch data with the ML model. To tackle these challenges, we have implemented in RTAD two core components: an input generation module and a GPU-inspired ML processing engine. According to our experiments, RTAD enables various ML models to infer anomaly instantly after the victim program behaves aberrantly as the result of attacks being injected into the system. Hyunyoung Oh, Hayoon Yi, Hyeokjun Choe, Yeongpil Cho, Sungroh Yoon, Yunheung Paek |
DATE | 6 |
| 2019 | CRCount: Pointer Invalidation with Reference Counting to Mitigate Use-after-free in Legacy C/C++
Jangseop Shin, Donghyun Kwon, Yeongpil Cho, Yunheung Paek |
NDSS | 5 |
| 2019 | uXOM: Efficient eXecute-Only Memory on ARM Cortex-M
Donghyun Kwon, Jangseop Shin, Giyeol Kim, Byoungyoung Lee, Yeongpil Cho, Yunheung Paek |
USENIX Security Symposium | 6 |
| 2019 | KI-Mon ARM: A Hardware-Assisted Event-triggered Monitoring Platform for Mutable Kernel ObjectabstractExternal hardware-based kernel integrity monitors have been proposed to mitigate kernel-level malwares. However, the existing external approaches have been limited to monitoring the static regions of kernel while the latest rootkits manipulate the dynamic kernel objects. To address the issue, we present KI-Mon, a hardware-based platform that introduces event-triggered monitoring techniques for kernel dynamic objects. KI-Mon advances the bus traffic snooping technique to not only detect memory write traffic on the host bus but also filter out all but meaningful traffic to generate events. We show how kernel invariant verification software can be developed around these events, and also provide a set of APIs for additional invariant verification development. We also report our findings and considerations on the unique challenges for external monitors – such as cache coherency, dynamic object tracing. We introduce host-side kernel changes that alleviate these issues that involve changes in kernel's object allocation and cache policy control. We have built a prototype of KI-Mon on the ARM architecture to demonstrate the efficacy of KI-Mon's event-triggered mechanism in terms of performance overhead for the monitored host system and the processor usage of the KI-Mon processor. Hojoon Lee 0001, Hyungon Moon, Ingoo Heo, Daehee Jang, Jin Soo Jang, Yunheung Paek, Brent ByungHoon Kang |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2019 | Safe and Efficient Implementation of a Security System on ARM using Intra-level Privilege SeparationabstractSecurity monitoring has long been considered as a fundamental mechanism to mitigate the damage of a security attack. Recently, intra-level security systems have been proposed that can efficiently and securely monitor system software without any involvement of more privileged entity. Unfortunately, there exists no full intra-level security system that can universally operate at any privilege level on ARM. However, as malware and attacks increase against virtually every level of privileged software including an OS, a hypervisor, and even the highest privileged software armored by TrustZone, we have been motivated to develop an intra-level security system, named Hilps . Hilps realizes true intra-level scheme in all these levels of privileged software on ARM by elaborately exploiting a new hardware feature of ARM’s latest 64-bit architecture, called TxSZ, that enables elastic adjustment of the accessible virtual address range. Furthermore, Hilps newly supports the sandbox mechanism that provides security tools with individually isolated execution environments, thereby minimizing security threats from untrusted security tools. We have implemented a prototype of Hilps on a real machine. The experimental results demonstrate that Hilps is quite promising for practical use in real deployments. Donghyun Kwon, Hayoon Yi, Yeongpil Cho, Yunheung Paek |
ACM Trans. Priv. Secur. | 4 |
| 2019 | DADE: a fast data anomaly detection engine for kernel integrity monitoring
Hayoon Yi, Yeongpil Cho, Yunheung Paek, Kwangman Ko |
J. Supercomput. | 3 |
| 2018 | Hypernel: a hardware-assisted framework for kernel protection without nested pagingabstractLarge OS kernels always suffer from attacks due to their numerous inherent vulnerabilities. To protect the kernel, hypervisors have been employed by many security solutions. However, relying on a hypervisor has a detrimental impact on the system performance due mainly to nested paging. In this paper, we present Hypernel, a security framework combining hardware and software components to address this problem. Hypersec, the software component, provides an isolated execution environment for security solutions, and the hardware monitor component enables a word-granularity monitoring capability on the kernel memory. Our evaluation shows that Hypernel efficiently fulfills the role of a security framework, while imposing mere 3.1% of runtime overhead on the system. Donghyun Kwon, Kuenwhee Oh, Junmo Park, Seungyong Yang, Yeongpil Cho, Brent ByungHoon Kang, Yunheung Paek |
DAC | 7 |
| 2018 | VM-CFI: Control-Flow Integrity for Virtual Machine Kernel Using Intel PT
Donghyun Kwon, Sehyun Baek, Giyeol Kim, Sunwoo Ahn, Yunheung Paek |
ICCSA (5) | 6 |
| 2018 | Hardware Assisted Randomization of Data
Brian Belleville, Hyungon Moon, Jangseop Shin, Dongil Hwang, Joseph Nash, Seonhwa Jung, Yeoul Na, Stijn Volckaert, Per Larsen, Yunheung Paek, Michael Franz |
RAID | 10 |
| 2018 | A dynamic per-context verification of kernel address integrity from external monitors
Hojoon Lee 0001, Yunheung Paek, Brent ByungHoon Kang |
Comput. Secur. | 3 |
| 2018 | Developing a custom DSP for vision based human computer interaction applications
Jangseop Shin, MoonKwon Kim, Yunheung Paek, Kwangman Ko |
Multim. Tools Appl. | 3 |
| 2017 | Instruction-Level Data Isolation for the Kernel on ARMabstractAs more sophisticated services are increasingly offered by the OS kernel on mobile devices, the security and sensitivity of kernel data that they depend on are becoming a critical issue. Data isolation has emerged as a key technique that can address the issue by providing strong protection for sensitive kernel data. However, existing data isolation mechanisms for mobile devices all incur non-negligible performance overhead. We deem that such computational burden would be a serious problem for mobile devices which already suffer from resource poverty. To alleviate this problem, we have developed a new mechanism that enforces data isolation very efficiently on ARM-based machines backed by unique hardware instructions. For evaluation, this instruction-level data isolation mechanism has been implemented in the Android/Linux kernel running on ARM. According to the experiment, it provides a lightweight data isolation capability for security services installed in the kernel. Yeongpil Cho, Donghyun Kwon, Yunheung Paek |
DAC | 3 |
| 2017 | Dynamic Virtual Address Range Adjustment for Intra-Level Privilege Separation on ARM
Yeongpil Cho, Donghyun Kwon, Hayoon Yi, Yunheung Paek |
NDSS | 4 |
| 2017 | Optimization techniques to enable execution offloading for 3D video games
Donghyun Kwon, Seungjun Yang, Yunheung Paek, Kwangman Ko |
Multim. Tools Appl. | 3 |
| 2017 | Detecting and Preventing Kernel Rootkit Attacks with Bus SnoopingabstractTo protect the integrity of operating system kernels, we presentVigilare system, a kernel integrity monitor that is architected to snoop the bus traffic of the host system from a separate independent hardware. Thissnoop-based monitoringenabled by the Vigilare system, overcomes the limitations of thesnapshot-based monitoringemployed in previous kernel integrity monitoring solutions. Being based on inspecting snapshots collected over a certain interval, the previous hardware-based monitoring solutions cannot detecttransient attacksthat can occur in between snapshots, and cannot protect the kernel against permanent damage. We implemented three prototypes of the Vigilare system by addingSnooperhardware connections module to the host system for bus snooping, and a snapshot-based monitor to be comared with, in order to evaluate the benefit of snoop-based monitoring. The prototypes of Vigilare system detected all the transient attacks and the second one protected the kernel with negligible performance degradation while the snapshot-based monitor could not detect all the attacks and induced considerable performance degradation as much as 10 percent in our tuned STREAM benchmark test. Hyungon Moon, Hojoon Lee 0001, Ingoo Heo, Yunheung Paek, Brent ByungHoon Kang |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2017 | Using CoreSight PTM to Integrate CRA Monitoring IPs in an ARM-Based SoCabstractThe ARM CoreSight Program Trace Macrocell (PTM) has been widely deployed in recent ARM processors for real-time debugging and tracing of software. Using PTM, the external debugger can extract execution behaviors of applications running on an ARM processor. Recently, some researchers have been using this feature for other purposes, such as fault-tolerant computation and security monitoring. This motivated us to develop an external security monitor that can detect control hijacking attacks, of which the goal is to maliciously manipulate the control flow of victim applications at an attacker’s disposal. This article focuses on detecting a special type of attack called code reuse attacks (CRA), which use a recently introduced technique that allows attackers to perform arbitrary computation without injecting their code by reusing only existing code fragments. Our external monitor is attached to the outside of the host system via the system bus and ARM CoreSight PTM, and is fed with execution traces of a victim application running on the host. As a majority of CRAs violates the normal execution behaviors of a program, our monitor constantly watches and analyzes the execution traces of the victim application and detects a symptom of attacks when the execution behaviors violate certain rules that normal applications are known to adhere. We present two different implementations for this purpose: a hardware-based solution in which all CRA detection components are implemented in hardware, and a hardware/software mixed solution that can be employed in a more resource-constrained environment where the deployment of full hardware-level CRA detection is burdensome. Yongje Lee, Jinyong Lee, Ingoo Heo, Dongil Hwang, Yunheung Paek |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2017 | Architectural Supports to Protect OS Kernels from Code-Injection Attacks and Their ApplicationsabstractThe kernel code injection is a common behavior of kernel-compromising attacks where the attackers aim to gain their goals by manipulating an OS kernel. Several security mechanisms have been proposed to mitigate such threats, but they all suffer from non-negligible performance overhead. This article introduces a hardware reference monitor, called Kargos, which can detect the kernel code injection attacks with nearly zero performance cost. Kargos monitors the behaviors of an OS kernel from outside the CPU through the standard bus interconnect and debug interface available with most major microprocessors. By watching the execution traces and memory access events in the monitored target system, Kargos uncovers attempts to execute malicious code with the kernel privilege. On top of this, we also applied the architectural supports for Kargos to the detection of ROP attacks. KS-Stack is the hardware component that builds and maintains the shadow stacks using the existing supports to detect this ROP attacks. According to our experiments, Kargos detected all the kernel code injection attacks that we tested, yet just increasing the computational loads on the target CPU by less than 1% on average. The performance overhead of the KS-Stack was also less than 1%. Hyungon Moon, Jinyong Lee, Dongil Hwang, Seonhwa Jung, Yunheung Paek |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2016 | Integration of ROP/JOP monitoring IPs in an ARM-based SoC
Yongje Lee, Jinyong Lee, Ingoo Heo, Dongil Hwang, Yunheung Paek |
DATE | 5 |
| 2016 | A hardware-based technique for efficient implicit information flow trackingabstractTo access sensitive information, some recent advanced attacks have been successful in exploiting implicit flows in a program in which sensitive data affects the control path and in turn affects other data. To track the sensitive data through implicit flows, several software and hardware based approaches have been proposed, but they suffer from the non-negligible performance overhead. In this paper, we propose a hardware tracking engine for implicit flow, called the implicit flow tracking unit (IFTU). By adopting the tracking scheme for implicit flow and mapping it to the specialized hardware, our solution can efficiently perform the implicit flow tracking with reasonable area costs. Jangseop Shin, Hongce Zhang, Jinyong Lee, Ingoo Heo, Yu-Yuan Chen, Ruby B. Lee, Yunheung Paek |
ICCAD | 7 |
| 2016 | HDFI: Hardware-Assisted Data-Flow IsolationabstractMemory corruption vulnerabilities are the root cause of many modern attacks. Existing defense mechanisms are inadequate; in general, the software-based approaches are not efficient and the hardware-based approaches are not flexible. In this paper, we present hardware-assisted data-flow isolation, or, HDFI, a new fine-grained data isolation mechanism that is broadly applicable and very efficient. HDFI enforces isolation at the machine word granularity by virtually extending each memory unit with an additional tag that is defined by dataflow. This capability allows HDFI to enforce a variety of security models such as the Biba Integrity Model and the Bell -- LaPadula Model. We implemented HDFI by extending the RISC-V instruction set architecture (ISA) and instantiating it on the Xilinx Zynq ZC706 evaluation board. We ran several benchmarks including the SPEC CINT 2000 benchmark suite. Evaluation results show that the performance overhead caused by our modification to the hardware is low (<; 2%). We also developed or ported several security mechanisms to leverage HDFI, including stack protection, standard library enhancement, virtual function table protection, code pointer protection, kernel data protection, and information leak prevention. Our results show that HDFI is easy to use, imposes low performance overhead, and allows us to create more elegant and more secure solutions. Chengyu Song, Hyungon Moon, Monjur Alam, Insu Yun, Byoungyoung Lee, Taesoo Kim, Wenke Lee, Yunheung Paek |
IEEE Symposium on Security and Privacy | 8 |
| 2016 | Hardware-Assisted On-Demand Hypervisor Activation for Efficient Security Critical Code Execution on Mobile Devices
Yeongpil Cho, Jun-Bum Shin, Donghyun Kwon, MyungJoo Ham, Yuna Kim, Yunheung Paek |
USENIX ATC | 6 |
| 2016 | Precise execution offloading for applications with dynamic behavior in mobile cloud computing
Yongin Kwon, Hayoon Yi, Donghyun Kwon, Seungjun Yang, Yeongpil Cho, Yunheung Paek |
Pervasive Mob. Comput. | 6 |
| 2016 | Software-Based Selective Validation Techniques for Robust CGRAs Against Soft ErrorsabstractCoarse-Grained Reconfigurable Architectures (CGRAs) are drawing significant attention since they promise both performances with parallelism and flexibility with reconfiguration. Soft errors (or transient faults) are becoming a serious design concern in embedded systems including CGRAs since the soft error rate is increasing exponentially as technology is scaling. A recently proposed software-based technique with TMR (Triple Modular Redundancy) implemented on CGRAs incurs extreme overheads in terms of runtime and energy consumption mainly due to expensive voting mechanisms for the outputs from the triplication of every operation. In this article, we propose selective validation mechanisms for efficient modular redundancy techniques in the datapaths on CGRAs. Our techniques selectively validate the results at synchronous operations rather than every operation in order to reduce the expensive performance overhead from the validation mechanism. We also present an optimization technique to further improve the runtime and the energy consumption by minimizing synchronous operations where a validating mechanism needs to be applied. Our experimental results demonstrate that our selective validation-based TMR technique with our optimization on CGRAs can improve the runtime by 41.0% and the energy consumption by 26.2% on average over benchmarks as compared to the recently proposed software-based TMR technique with the full validation. Yohan Ko, Jihoon Kang, Joonhyun Kim, Hwisoo So, Kyoungwoo Lee, Yunheung Paek |
ACM Trans. Embed. Comput. Syst. | 8 |
| 2016 | Efficient Security Monitoring with the Core Debug Interface in an Embedded ProcessorabstractFor decades, various concepts in security monitoring have been proposed. In principle, they all in common in regard to the monitoring of the execution behavior of a program (e.g., control-flow or dataflow) running on the machine to find symptoms of attacks. Among the proposed monitoring schemes, software-based ones are known for their adaptability on the commercial products, but there have been concerns that they may suffer from nonnegligible runtime overhead. On the other hand, hardware-based solutions are recognized for their high performance. However, most of them have an inherent problem in that they usually mandate drastic changes to the internal processor architecture. More recent ones have strived to minimize such modifications by employing external hardware security monitors in the system. However, these approaches intrinsically suffer from the overhead caused by communication between the host and the external monitor. Our solution also relies on external hardware for security monitoring, but unlike the others, ours tackles the communication overhead by using the core debug interface (CDI), which is readily available in most commercial processors for debugging. We build our system simply by plugging our monitoring hardware into the processor via CDI, precluding the need for altering the processor internals. To validate the effectiveness of our approach, we implement two well-known monitoring techniques on our proposed framework: dynamic information flow tracking and branch regulation. The experimental results on our FPGA prototype show that our external hardware monitors efficiently perform monitoring tasks with negligible performance overhead, mainly with thanks to the support of CDI, which helps us reduce communication costs substantially. Jinyong Lee, Ingoo Heo, Yongje Lee, Yunheung Paek |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2015 | Accelerating bootstrapping in FHEW using GPUsabstractRecently, the usage of GPU is not limited to the jobs associated with graphics and a wide variety of applications take advantage of the flexibility of GPUs to accelerate the computing performance. Among them, one of the most emerging applications is the fully homomorphic encryption (FHE) scheme, which enables arbitrary computations on encrypted data. Despite much research effort, it cannot be considered as practical due to the enormous amount of computations, especially in the bootstrapping procedure. In this paper, we accelerate the performance of the recently suggested fast bootstrapping method in FHEW scheme using GPUs, as a case study of a FHE scheme. In order to optimize, we explored the reference code and carried out profiling to find out candidates for performance acceleration. Based on the profiling results, combined with more flexible tradeoff method, we optimized the bootstrapping algorithm in FHEW using GPU and CUDA's programming model. The empirical result shows that the bootstrapping of FHEW ciphertext can be done in less than 0.11 second after optimization. Moon Sung Lee, Yongje Lee, Jung Hee Cheon, Yunheung Paek |
ASAP | 4 |
| 2015 | Efficient dynamic information flow tracking on a processor with core debug interfaceabstractDynamic information flow tracking (DIFT) is a promising solution to prevent various attacks on software running on a processor. Previous hardware solutions usually mandate drastic change to internal processor architecture. More recent ones to minimize the change have proposed external devices for DIFT. However, these approaches intrinsically suffer from the high overhead to communicate with their external devices. Consequently, they either significantly lose performance, or inevitably make invasive modifications to the processor inside. Our solution also rely on external hardware for DIFT, but unlike theirs, ours exploits the core debug interface (CDI) to tackle the communication issue. CDI is provided in most commercial processors for debugging so that we were able to build our system simply by plugging our hardware to the processor via CDI, precluding the need for altering the processor itself. Experiments show that our hardware efficiently performs DIFT mainly thanks to the support of CDI that helps us cut substantially down the communication costs. Jinyong Lee, Ingoo Heo, Yongje Lee, Yunheung Paek |
DAC | 4 |
| 2015 | Extrax: security extension to extract cache resident information for snoop-based external monitors
Jinyong Lee, Yongje Lee, Hyungon Moon, Ingoo Heo, Yunheung Paek |
DATE | 5 |
| 2015 | Mantis: Efficient Predictions of Execution Time, Energy Usage, Memory Usage and Network Usage on Smart Mobile DevicesabstractWe present Mantis, a framework for predicting the computational resource consumption (CRC) of Android applications on given inputs accurately, and efficiently. A key insight underlying Mantis is that program codes often contain features that correlate with performance and these features can be automatically computed efficiently. Mantis synergistically combines techniques from program analysis and machine learning. It constructs concise CRC models by choosing from many program execution features only a handful that are most correlated with the program's CRC metric yet can be evaluated efficiently from the program's input. We apply program slicing to reduce evaluation time of a feature and automatically generate executable code snippets for efficiently evaluating features. Our evaluation shows that Mantis predicts four CRC metrics of seven Android apps with estimation error in the range of 0-11.1 percent by executing predictor code spending at most 1.3 percent of their execution time on Galaxy Nexus. Yongin Kwon, Hayoon Yi, Donghyun Kwon, Seungjun Yang, Byung-Gon Chun, Ling Huang 0001, Petros Maniatis, Mayur Naik, Yunheung Paek |
IEEE Trans. Mob. Comput. | 10 |
| 2015 | Implementing an Application-Specific Instruction-Set Processor for System-Level Dynamic Program Analysis EnginesabstractIn recent years, dynamic program analysis (DPA) has been widely used in various fields such as profiling, finding bugs, and security. However, existing solutions have their own weaknesses. Software solutions provide flexibility in DPA but they suffer from tremendous performance overhead. In contrast, core-level hardware engines rely on specialized integrated logics and attain extremely fast computation, but they have a limited functional extensibility because the logics are tightly coupled with the host processor. To mend this, a prior system-level approach utilizes an existing channel to integrate their hardware without necessitating the host architecture modification and introduced great potential in performance. Nevertheless, the prior work does not address the detailed design and implementation of the engine, which is quite essential to leverage the deployment on real systems. To address this, in this article, we propose an implementation of programmable DPA hardware engine, called program analysis unit (PAU). PAU is an application-specific instruction-set processor (ASIP) whose instruction set is customized to reflect common features of various DPA methods. With the specialized architecture and programmability of software, our PAU aims at fast computation and sufficient flexibility. In our case studies on several DPA techniques, we show that our ASIP approach can be successfully applicable to complex DPA schemes while providing hardware-backed power in performance and software-based flexibility in analysis. Recent experiments on our FPGA prototype revealed that the performance of PAU is 4.7-13.6 times faster than pure software DPA, and the power/area consumption is also acceptably small compared to today's mobile processors. Ingoo Heo, Yongje Lee, Changho Choi, Jinyong Lee, Brent ByungHoon Kang, Yunheung Paek |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2014 | CMcloud: Cloud Platform for Cost-Effective Offloading of Mobile ApplicationsabstractRecent efforts towards mobile cloud propose to offload mobile applications to cloud servers for the improved performance and battery life of mobile devices. However, existing schemes completely ignore the costs of cloud resources by assuming that idle servers are always available for free of charge. These unrealistic assumptions make each server run only a small load to achieve the guaranteed high offload performance. Therefore, these schemes cannot be applied to real-world commercial clouds which aim to minimize the operation costs by maximizing the server throughput, and then charge users for their resource usage. In this paper, we propose CMcloud, a novel cost-effective mobile-to-cloud offloading platform, which works nicely under the real-world cloud environments. CMcloud minimizes both the server costs and the user service fee by offloading as many mobile applications to a single server as possible, while satisfying the target performance of all applications. To achieve such goals, CMcloud exploits novel architecture performance modeling and server migration techniques. Our implementation shows that CMcloud can improve the data enter throughput by 84% over a conventional static light-load scheme (or a 2.7x higher per-socket throughput.) Alternatively, CMcloud reduces the number of service failures by 83% over a static high-load scheme, while even improving the throughput by 31%. Dongju Chae, Jihun Kim 0002, Jangwoo Kim, Jong Kim 0001, Seungjun Yang, Yeongpil Cho, Yongin Kwon, Yunheung Paek |
CCGRID | 8 |
| 2014 | Improving performance of loops on DIAM-based VLIW architecturesabstractRecent studies show that very long instruction word (VLIW) architectures, which inherently have wide datapath (e.g. 128 or 256 bits for one VLIW instruction word), can benefit from dynamic implied addressing mode (DIAM) and can achieve lower power consumption and smaller code size with a small performance overhead. Such overhead, which is claimed to be small, is mainly caused by the execution of additionally generated special instructions for conveying information that cannot be encoded in reduced instruction bit-width. In this paper, however, we show that the performance impact of applying DIAM on VLIW architecture cannot be overlooked expecially when applications possess high level of instruction level parallelism (ILP), which is mostly the case for loops because of the result of aggressive code scheduling. We also propose a way to relieve the performance degradation especially focusing on loops since loops spend almost 90% of total execution time in programs and tend to have high ILP. We first implement the original DIAM compilation technique in a compiler, and augment it with the proposed loop optimization scheme to show that ours can clearly alleviate the performance loss caused by the excessive number of additional instructions, with the help of slightly modified hardware. Moreover, the well-known loop unrolling scheme, which would produce denser code in loops at the cost of substantial code size bloating, is integrated into our compiler. The experiment result shows that the loop unrolling technique, combined with our augmented DIAM scheme, produces far better code in terms of performance with quite an acceptable amount of code increase. Jinyong Lee, Jongeun Lee, Yunheung Paek |
LCTES | 4 |
| 2014 | Techniques to Minimize State Transfer Costs for Dynamic Execution Offloading in Mobile Cloud ComputingabstractIn order to meet the increasing demand for high performance in smartphones, recent studies suggested mobile cloud computing techniques that aim to connect the phones to adjacent powerful cloud servers to throw their computational burden to the servers. These techniques often employ execution offloading schemes that migrate a process between machines during its execution. In execution offloading, code regions to be executed on the server are decided statically or dynamically based on the complex analysis on execution time and process state transfer costs of every region. Expectedly, the transfer cost is a deciding factor for the success of execution offloading. According to our analysis, it is dominated by the total size of heap objects transferred over the network. But previous work did not try hard to minimize this size. Thus in this paper, we introduce novel techniques based on compiler code analysis that effectively reduce the transferred data size by transferring only the essential heap objects and the stack frames actually referenced in the server. The experiments exhibit that the reduced size positively influences not only the transfer time itself but also the overall effectiveness of execution offloading, and ultimately, improves the performance of our mobile cloud computing significantly in terms of execution time and energy consumption. Seungjun Yang, Donghyun Kwon, Hayoon Yi, Yeongpil Cho, Yongin Kwon, Yunheung Paek |
IEEE Trans. Mob. Comput. | 6 |
| 2013 | Selective validations for efficient protections on Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-Grained Reconfigurable Architectures or CGRAs are drawing significant attention since they promise both performance with parallelism and flexibility with reconfiguration. Soft errors or transient faults are becoming a serious design concern in embedded systems including CGRAs since soft error rate is increasing exponentially as technology scaling. A recently proposed software-based technique with TMR (Triple Modular Redundancy) implemented on CGRAs incurs extreme performance overhead mainly due to expensive voting mechanisms for outputs from triplication of every operation. In this paper, we propose selective validation mechanisms for efficient modular redundancy techniques in the datapaths on CGRAs. Our techniques selectively validate results at synchronous operations rather than every operation in order to reduce the expensive performance overhead from the validation mechanism. Our experimental results demonstrate that our selective validation based TMR technique can improve the performance by 38.3% on average over benchmarks as compared to the recently proposed software-based TMR technique with the full validation. Jihoon Kang, Yohan Ko, Hwisoo So, Kyoungwoo Lee, Yunheung Paek |
ASAP | 7 |
| 2013 | Fast dynamic execution offloading for efficient mobile cloud computingabstractIn order to meet the increasing demand for high performance in smartphones, recent studies suggested mobile cloud computing techniques that aim to connect the phones to adjacent powerful cloud servers to throw their computational burden to the servers. These techniques often employ execution offloading schemes that migrate a process between machines during its execution. In execution offloading, code regions to be executed on the server are decided statically or dynamically based on the complex analysis on execution time and process state transfer time of every region. Expectedly, the transfer time is a deciding factor for the success of execution offloading. According to our analysis, it is dominated by the total size of heap objects transferred over the network. But previous work did not try hard to minimize this size. Thus in this paper, we introduce novel techniques based on compiler code analysis that effectively reduce the transferred data size by transferring only the essential heap objects. The experiments exhibit that the reduced size positively influences not only the transfer time itself but also the overall effectiveness of execution offloading, and ultimately, improves the performance of our mobile cloud computing significantly in terms of execution time and power consumption. Seungjun Yang, Yongin Kwon, Yeongpil Cho, Hayoon Yi, Donghyun Kwon, Jonghee M. Youn, Yunheung Paek |
PerCom | 7 |
| 2013 | Mantis: Automatic Performance Prediction for Smartphone Applications
Yongin Kwon, Hayoon Yi, Donghyun Kwon, Seungjun Yang, Byung-Gon Chun, Ling Huang 0001, Petros Maniatis, Mayur Naik, Yunheung Paek |
USENIX ATC | 10 |
| 2013 | KI-Mon: A Hardware-assisted Event-triggered Monitoring Platform for Mutable Kernel Object
Hojoon Lee 0001, Hyungon Moon, Daehee Jang, Yunheung Paek, Brent ByungHoon Kang |
USENIX Security Symposium | 6 |
| 2013 | Dynamic code duplication with vulnerability awareness for soft error detection on VLIW architecturesabstractSoft errors are becoming a critical concern in embedded system designs. Code duplication techniques have been proposed to increase the reliability in multi-issue embedded systems such as VLIW by exploiting empty slots for duplicated instructions. However, they increase code size, another important concern, and ignore vulnerability differences in instructions, causing unnecessary or inefficient protection when selecting instructions to be duplicated under constraints. In this article, we propose a compiler-assisted dynamic code duplication method to minimize the code size overhead, and present vulnerability-aware duplication algorithms to maximize the effectiveness of instruction duplication with least overheads for VLIW architecture. Our experimental results with SoarGen and Synopsys simulation environments demonstrate that our proposals can reduce the code size by up to 40% and detect more soft errors by up to 82% via fault injection experiments over benchmarks from DSPstone and Livermore Loops as compared to the previously proposed instruction duplication technique. Yohan Ko, Kyoungwoo Lee, Jonghee M. Youn, Yunheung Paek |
ACM Trans. Archit. Code Optim. | 5 |
| 2013 | Reducing instruction bit-width for low-power VLIW architecturesabstractVLIW (very long instruction word) architectures have proven to be useful for embedded applications with abundant instruction level parallelism. But due to the long instruction bus width it often consumes more power and memory space than necessary. One way to lessen this problem is to adopt a reduced bit-width instruction set architecture (ISA) that has a narrower instruction word length. This facilitates a more efficient hardware implementation in terms of area and power by decreasing bus-bandwidth requirements and the power dissipation associated with instruction fetches. In practice, however, it is impossible to convert a given ISA fully into an equivalent reduced bit-width one because the narrow instruction word, due to bit-width restrictions, can encode only a small subset of normal instructions in the original ISA. Consequently, existing processors provide narrow instructions in very limited cases along with severe restrictions on register accessibility. The objective of this work is to explore the possibility of complete conversion, as a case study, of an existing 32-bit VLIW ISA into a 16-bit one that supports effectively all 32-bit instructions. To this objective, we attempt to circumvent the bit-width restrictions by dynamically extending the effective instruction word length of the converted 16-bit operations. Further, we will show that our proposed ISA conversion can create a synergy effect with a VLES (variable length execution set) architecture that is adopted in most recent VLIW processors. According to our experiment, the code size becomes significantly smaller after the conversion to 16-bit VLIW code. Also at a slight run time cost, the machine with the 16-bit ISA consumes much less energy than the original machine. Jonghee M. Youn, Doosan Cho, Yunheung Paek |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2013 | Architecture customization of on-chip reconfigurable acceleratorsabstractIntegrating coarse-grained reconfigurable architectures (CGRAs) into a System-on-a-Chip (SoC) presents many benefits as well as important challenges. One of the challenges is how to customize the architecture for the target applications efficiently and effectively without performing explicit design space exploration. In this article we present a novel methodology for incremental interconnect customization of CGRAs that can suggest a new interconnection architecture which is able to maximize the performance for a given set of application kernels while minimizing the hardware cost. In our methodology, we translate the problem of interconnect customization into that of inexact graph matching, and we devised a heuristic for A* search algorithm to efficiently solve the inexact graph matching problem. Our experimental results demonstrate that our customization method can quickly find application-optimized interconnections that exhibit 80% higher performance on average compared to the base architecture which has mesh interconnections, with little energy and hardware increase in interconnections and muxes. Jonghee W. Yoon, Jongeun Lee, Jinyong Lee, Yunheung Paek, Doosan Cho |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2012 | Vigilare: toward snoop-based kernel integrity monitorabstractIn this paper, we present Vigilare system, a kernel integrity monitor that is architected to snoop the bus traffic of the host system from a separate independent hardware. This snoop-based monitoring enabled by the Vigilare system, overcomes the limitations of the snapshot-based monitoring employed in previous kernel integrity monitoring solutions. Being based on inspecting snapshots collected over a certain interval, the previous hardware-based monitoring solutions cannot detect transient attacks that can occur in between snapshots. We implemented a prototype of the Vigilare system on Gaisler's grlib-based system-on-a-chip (SoC) by adding Snooper hardware connections module to the host system for bus snooping. To evaluate the benefit of snoop-based monitoring, we also implemented similar SoC with a snapshot-based monitor to be compared with. The Vigilare system detected all the transient attacks without performance degradation while the snapshot-based monitor could not detect all the attacks and induced considerable performance degradation as much as 10% in our tuned STREAM benchmark test. Hyungon Moon, Hojoon Lee 0001, Yunheung Paek, Brent ByungHoon Kang |
CCS | 5 |
| 2012 | Dynamic Operands Insertion for VLIW Architecture with a Reduced Bit-width Instruction SetabstractPerformance, code size and power consumption are all primary concern in embedded systems. To this effect, VLIW architecture has proven to be useful for embedded applications with abundant instruction level parallelism. But due to the long instruction bus width it often consumes more power and memory space than necessary. One way to lessen this problem is to adopt a reduced bit-width instruction set architecture (ISA) that has a narrower instruction word length. This facilitates a more efficient hardware implementation in terms of area and power by decreasing bus-bandwidth requirements and the power dissipation associated with instruction fetches. Also earlier studies reported that it helps to reduce the code size considerably. In practice, however, it is impossible to convert a given ISA fully into an equivalent reduced bit-width one because the narrow instruction word, due to bit-width restrictions, can encode only a small subset of normal instructions in the original ISA. Consequently, existing processors provide narrow instructions in very limited cases along with severe restrictions on register accessibility. The objective of this work is to explore the possibility of complete conversion, as a case study, of an existing 32-bit VLIW ISA into a 16-bit one that supports effectively all 32-bit instructions. To this objective, we attempt to circumvent the bit-width restrictions by dynamically extending the effective instruction word length of the converted 16-bit operations. At compile time when a 32-bit operation is converted to a 16-bit word format, we compute how many bits are additionally needed to represent the whole 32-bit operation and store the bits separately in the VLIW code. Then at run time, these bits are retrieved on demand and inserted to a proper 16-bit operation to reconstruct the original 32-bit representation. According to our experiment, the code size becomes significantly smaller after the conversion to 16-bit VLIW code. Also at a slight run time cost, the machine with the 16-bit ISA consumes much less energy than the original machine. Jonghee M. Youn, Minwook Ahn, Yunheung Paek |
IPDPS | 5 |
| 2012 | Compiler and microarchitectural approaches for register file thermal managementabstractRegister files are responsible for the majority of thermal emergencies in high performance processors. In this paper, we propose compiler-microarchitecture cooperative techniques to reduce the performance penalty incurred by existing thermal management schemes. Our experiments show that an existing dynamic thermal management (DTM) scheme may result in 47% performance degradation at 45 nm technology node as compared to the ideal case where the processor can sustain any thermal emergency and no DTM is applied. We also evaluate performance degradation of a couple of other DTMs and demonstrate that our cooperative policy can bring down this performance penalty. Our approach shows the consistent effectiveness even with the technology scaling. Ingoo Heo, Yunheung Paek |
ISCAS | 3 |
| 2012 | Improving performance of nested loops on reconfigurable array processorsabstractPipelining algorithms are typically concerned with improving only the steady-state performance, or the kernel time. The pipeline setup time happens only once and therefore can be negligible compared to the kernel time. However, for Coarse-Grained Reconfigurable Architectures (CGRAs) used as a coprocessor to a main processor, pipeline setup can take much longer due to the communication delay between the two processors, and can become significant if it is repeated in an outer loop of a loop nest. In this paper we evaluate the overhead of such non-kernel execution times when mapping nested loops for CGRAs, and propose a novel architecture-compiler cooperative scheme to reduce the overhead, while also minimizing the number of extra configurations required. Our experimental results using loops from multimedia and scientific domains demonstrate that our proposed techniques can greatly increase the performance of nested loops by up to 2.87 times compared to the conventional approach of accelerating only the innermost loops. Moreover, the mappings generated by our techniques require only a modest number of configurations that can fit in recent reconfigurable architectures. Jongeun Lee, Toan X. Mai, Yunheung Paek |
ACM Trans. Archit. Code Optim. | 4 |
| 2011 | I2CRF: Incremental interconnect customization for embedded reconfigurable fabricsabstractIntegrating coarse-grained reconfigurable architectures (CGRAs) into a System-on-a-Chip (SoC) presents many benefits as well as important challenges. One of the challenges is how to customize the architecture for the target applications efficiently and effectively without explicit design space exploration. In this paper we present a novel methodology for incremental interconnect customization of CGRAs that can suggest a new interconnection architecture that can maximize the performance for a given set of application kernels while minimizing the hardware cost. Applying the inexact graph matching analogy, we translate our problem into graph matching taking into account the cost of various graph edit operations, which we solve using the A* search algorithm with a heuristic tailored to our problem. Our experimental results demonstrate that our customization method can quickly find application-optimized interconnections that exhibit 70% higher performance on average compared to the base architecture, with relatively little hardware increase in interconnections and muxes. Jonghee W. Yoon, Jongeun Lee, Jaewan Jung, Yunheung Paek, Doosan Cho |
DATE | 6 |
| 2011 | Fast graph-based instruction selection for multi-output instructionsabstractAbstract A multi‐output instruction (MOI) is an instruction that produces multiple outputs to its destination locations. Such inherently parallel instructions are becoming more and more popular in embedded processors, due to the advances in application‐specific architectures. In order to provide high‐level programmability and thus guarantee widespread acceptance, sophisticated compiler support for these programmable cores is necessary. However, traditionaltree‐basedapproaches for instruction selection, although very fast, fail to exploit MOIs mainly because of the fundamental limitation of the tree representation. In fact, to generate optimal code with MOIs requires a more generalgraph‐basedformulation of the instruction selection problem, which is at least NP‐complete. In this paper we present a new methodology to automatically generate from simple instruction set descriptions, graph‐based code selectors that can effectively utilize all provided instructions including MOIs. Our experimental results using a set of benchmarks on a target processor with various MOIs of up to two outputs demonstrate that our generated code selectors can quickly and effectively exploit many MOIs at the application level, and therefore are highly desirable both for architecture exploration and as code generators after architecture is fixed. Copyright © 2010 John Wiley & Sons, Ltd. Jonghee M. Youn, Yunheung Paek, Jongeun Lee, Hanno Scharwächter, Rainer Leupers |
Softw. Pract. Exp. | 3 |
| 2011 | High Throughput Data Mapping for Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-grained reconfigurable arrays (CGRAs) are a very promising platform, providing both up to 10–100 MOps/mW of power efficiency and software programmability. However, this promise of CGRAs critically hinges on the effectiveness of application mapping onto CGRA platforms. While previous solutions have greatly improved the computation speed, they have largely ignored the impact of the local memory architecture on the achievable power and performance. This paper motivates the need for memory-aware application mapping for CGRAs, and proposes an effective solution for application mapping that considers the effects of various memory architecture parameters including the number of banks, local memory size, and the communication bandwidth between the local memory and the external main memory. Further we propose efficient methods to handle dependent data on a double-buffering local memory, which is necessary for recurrent loops. Our proposed solution achieves 59% reduction in the energy-delay product, which factors into about 47% and 22% reduction in the energy consumption and runtime, respectively, as compared to memory-unaware mapping for realistic local memory architectures. We also show that our scheme scales across a range of applications and memory parameters, and the runtime overhead of handling recurrent loops by our proposed methods can be less than 1%. Jongeun Lee, Aviral Shrivastava, Jonghee W. Yoon, Doosan Cho, Yunheung Paek |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2011 | Memory access optimization in compilation for coarse-grained reconfigurable architecturesabstractCoarse-grained reconfigurable architectures (CGRAs) promise high performance at high power efficiency. They fulfil this promise by keeping the hardware extremely simple, and moving the complexity to application mapping. One major challenge comes in the form of data mapping. For reasons of power-efficiency and complexity, CGRAs use multibank local memory, and a row of PEs share memory access. In order for each row of the PEs to access any memory bank, there is a hardware arbiter between the memory requests generated by the PEs and the banks of the local memory. However, a fundamental restriction remains in that a bank cannot be accessed by two different PEs at the same time. We propose to meet this challenge by mapping application operations onto PEs and data into memory banks in a way that avoids such conflicts. To further improve performance on multibank memories, we propose a compiler optimization for CGRA mapping to reduce the number of memory operations by exploiting data reuse. Our experimental results on kernels from multimedia benchmarks demonstrate that our local memory-aware compilation approach can generate mappings that are up to 53% better in performance (26% on average) compared to a memory-unaware scheduler. Jongeun Lee, Aviral Shrivastava, Yunheung Paek |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2010 | Implementing dynamic implied addressing mode for multi-output instructionsabstractThe ever-increasing demand for faster execution time, smaller resource usage and lower energy consumption has compelled architects of embedded processors to adopt more specialized hardware features with irregular data paths and heterogeneous registers that are customized to the needs of their target applications. These processors consequently provide a rich set of specialized instructions in order to enable programmers to access these features. Such an instruction is typically a multi-output instruction (MOI), which outputs multiple results parallely in order to exploit inherent underlying hardware parallelism. Earlier study has exhibited that MOIs help to enhance performance in aspect of instruction counts and code size. However, as MOIs require more operands, they tend to increase not only the size of the instruction set but also the size of individual instructions. This can be a serious setback for embedded processors, which are mostly subject to strong resource limitations (particularly in this case, limited instruction encoding space). For this reason, these processors are often allowed to include only a very small subset of the total desired MOIs in their instruction sets, despite there can be sufficient silicon real estate to accommodate these specialized MOIs. To attack this problem, we introduce a novel instruction encoding scheme based on the dynamic implied addressing mode (DIAM). In this paper, we will discuss how we have overcome the encoding space problem for our target embedded processor whose instruction set has been augmented with a variety of MOIs. Our DIAM-based encoding scheme employs a small on-chip buffer to supplement extra encoding information for MOIs at run time. The empirical results are promising: the scheme allows us to encode many more MOIs for our processor; thereby helping us to achieve considerable reduction of code size as well as running time after the DIAM is additively implemented in the original architecture. Jonghee M. Youn, Yunheung Paek, Jongwung Kim, Jeonghun Cho 0001 |
CASES | 3 |
| 2010 | Memory-Aware Application Mapping on Coarse-Grained Reconfigurable Arrays
Jongeun Lee, Aviral Shrivastava, Jonghee W. Yoon, Yunheung Paek |
HiPEAC | 5 |
| 2010 | Operation and data mapping for CGRAs with multi-bank memoryabstractCoarse Grain Reconfigurable Architectures (CGRAs) promise high performance at high power efficiency. They fulfil this promise by keeping the hardware extremely simple, and moving the complexity to application mapping. One major challenge comes in the form of data mapping. For reasons of power-efficiency and complexity, CGRAs use multi-bank local memory, and a row of PEs share memory access. In order for each row of the PEs to access any memory bank, there is a hardware arbiter between the memory requests generated by the PEs and the banks of the local memory. However, a fundamental restriction remains that a bank cannot be accessed by two different PEs at the same time. We propose to meet this challenge by mapping application operations onto PEs and data into memory banks in a way that avoids such conflicts. Our experimental results on kernels from multimedia benchmarks demonstrate that our local memory-aware compilation approach can generate mappings that are up to 40% better in performance (17.3% on average) compared to a memory-unaware scheduler. Jongeun Lee, Aviral Shrivastava, Yunheung Paek |
LCTES | 4 |
| 2010 | Two versions of architectures for dynamic implied addressing mode
Jonghee M. Youn, Minwook Ahn, Yunheung Paek, Jongwung Kim, Jeonghun Cho 0001 |
J. Syst. Archit. | 3 |
| 2009 | Iterative Algorithm for Compound Instruction Selection with Register CoalescingabstractA compound instruction, encoding several ALU or memory operations within an instruction word, has been regarded as an efficient way of improving performance. In the compiler for embedded processors, the code generation algorithm for compound instructions has been built by dealing mainly with instruction selection which is a crucial phase of code generation. In this paper, we propose an iterative code generation algorithm for minimizing the detrimental impact of register coalescing that is applied to the code with compound instructions generated earlier from the instruction selection phase. Minwook Ahn, Jonghee M. Youn, Youngkyu Choi, Doosan Cho, Yunheung Paek |
DSD | 5 |
| 2009 | Orthogonal Instruction Encoding for a 16-bit Embedded Processor with Dynamic Implied Addressing ModeabstractAlthough 32-bit architectures are becoming the norm for modern microprocessors, 16-bit ones are still employed by many low-end processors, for which small size and low power consumption are of high priority. However, 16-bit architectures have a critical disadvantage for embedded processors that they do not provide enough encoding space to add special instructions coined for certain applications. To overcome this, many existing architectures adopt non-orthogonal, irregular instruction sets to accommodate a variety of unusual addressing modes thru which more opcodes and operands are densely encoded within the narrow instruction word. In general, these non-orthogonal architectures are regarded compiler-unfriendly as they tend to requires extremely sophisticated compiler techniques for optimal code generation. To address this issue, we propose a compiler-friendly processor with a new addressing mode, called the dynamic implied addressing mode (DIAM). In this paper, we will demonstrate that the DIAM provides more encoding space for our 16-bit processor so that we are able to support more instructions specially customized for our applications. And yet, the processor maintains a RISC-style orthogonal architecture, thereby allowing us to use traditional code generation algorithms. In our experiment, the architecture augmented with DIAMs shows 6.2% code size reduction and 3.5% performance increase on average, as compared to the basic architecture without DIAMs. Jonghee M. Youn, Daeho Kim, Minwook Ahn, Yunheung Paek |
HPCC | 5 |
| 2009 | Adaptive Scratch Pad Memory Management for Dynamic Behavior of Multimedia ApplicationsabstractExploiting runtime memory access traces can be a complementary approach to compiler optimizations for the energy reduction in memory hierarchy. This is particularly important for emerging multimedia applications since they usually have input-sensitive runtime behavior which results in dynamic and/or irregular memory access patterns. These types of applications are normally hard to optimize by static compiler optimizations. The reason is that their behavior stays unknown until runtime and may even change during computation. To tackle this problem, we propose an integrated approach of software [compiler and operating system (OS)] and hardware (data access record table) techniques to exploit data reusability of multimedia applications in Multiprocessor Systems on Chip. Guided by compiler analysis for generating scratch pad data layouts and hardware components for tracking dynamic memory accesses, the scratch pad data layout adapts to an input data pattern with the help of a runtime scratch pad memory manager incorporated in the OS. The runtime data placement strategy presented in this paper provides efficient scratch pad utilization for the dynamic applications. The goal is to minimize the amount of accesses to the main memory over the entire runtime of the system, which leads to a reduction in the energy consumption of the system. Our experimental results show that our approach is able to significantly improve the energy consumption of multimedia applications with dynamic memory access behavior over an existing compiler technique and an alternative hardware technique. Doosan Cho, Sudeep Pasricha, Ilya Issenin, Nikil Dutt, Minwook Ahn, Yunheung Paek |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2009 | Compiler-in-the-Loop Design Space Exploration Framework for Energy Reduction in Horizontally Partitioned Cache ArchitecturesabstractHorizontally partitioned caches (HPCs) are a power-efficient architectural feature in which the processor maintains two or more data caches at the same level of hierarchy. HPCs help reduce cache pollution and thereby improve performance. Consequently, most previous research has focused on exploiting HPCs to improve performance and achieve energy reduction only as a byproduct of performance improvement. However, with energy consumption becoming the first class design constraint, there is an increasing need for compilation techniques aimed at energy reduction itself. This paper proposes and explores several low-complexity algorithms aimed at reducing the energy consumption. Acknowledging that the compiler has a significant impact on the energy consumption of the HPCs, Compiler-in-the-Loop Design Space Exploration methodologies are also presented to carefully choose the HPC parameters that result in minimum energy consumption for the application. Aviral Shrivastava, Ilya Issenin, Nikil Dutt, Yunheung Paek |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2009 | Register coalescing techniques for heterogeneous register architecture with copy siftingabstractOptimistic coalescing has been proven as an elegant and effective technique that provides better chances of safely coloring more registers in register allocation than other coalescing techniques. Its algorithm originally assumes homogeneous registers, which are all gathered in the same register file. Although this register architecture is still common in most general-purpose processors, embedded processors often contain heterogeneous registers, which are scattered in physically different register files dedicated for each dissimilar purpose and use. In this work, we show that optimistic coalescing is also useful for an embedded processor to better handle such heterogeneity of the register architecture, and developed a modified algorithm for optimal coalescing that helps a register allocator. In the experiment, an existing register allocator was able to achieve up to 13.0% reduction in code size through our coalescing, and avoid many spills that would have been generated without our scheme. Minwook Ahn, Yunheung Paek |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2009 | A Graph Drawing Based Spatial Mapping Algorithm for Coarse-Grained Reconfigurable ArchitecturesabstractRecently coarse-grained reconfigurable architectures (CGRAs) have drawn increasing attention due to their efficiency and flexibility. While many CGRAs have demonstrated impressive performance improvements, the effectiveness of CGRA platforms ultimately hinges on the compiler. Existing CGRA compilers do not model the details of the CGRA, and thus they are i) unable to map applications, even though a mapping exists, and ii) using too many processing elements (PEs) to map an application. In this paper, we model several CGRA details, e.g., irregular CGRA topologies, shared resources and routing PEs in our compiler and develop a graph drawing based approach, split-push kernel mapping (SPKM), for mapping applications onto CGRAs. On randomly generated graphs our technique can map on average 4.5times more applications than the previous approach, while generating mappings which have better qualities in terms of utilized CGRA resources. Utilizing fewer resources is directly translated into increased opportunities for novel power and performance optimization techniques. Our technique shows less power consumption in 71 cases and shorter execution cycles in 66 cases out of 100 synthetic applications, with minimum mapping time overhead. We observe similar results on a suite of benchmarks collected from Livermore loops, Mediabench, Multimedia, Wavelet and DSPStone benchmarks. SPKM is not a customized algorithm only for a specific CGRA template, and it is demonstrated by exploring various PE interconnection topologies and shared resource configurations with SPKM. Jonghee W. Yoon, Aviral Shrivastava, Minwook Ahn, Yunheung Paek |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2008 | SPKM : A novel graph drawing based algorithm for application mapping onto coarse-grained reconfigurable architecturesabstractRecently coarse-grained reconfigurable architectures (CGRAs) have drawn increasing attention due to their efficiency and flexibility. While many CGRAs have demonstrated impressive performance improvements, the effectiveness of CGRA platforms ultimately hinges on the compiler. Existing CGRA compilers do not model the details of the CGRA architecture, due to which they are, i) unable to map applications, even though a mapping exists, and ii) use too many PEs to map an application. In this paper, we model several CGRA details in our compiler and develop a graph mapping based approach (SPKM) for mapping applications onto CGRAs. On randomly generated graphs our technique can map on average 4.5X more applications than the previous approaches, while using fewer CGRA rows 62% times, without any penalty in mapping time. We observe similar results on a suite of benchmarks collected from Livermore Loops, Multimedia and DSPStone benchmarks. Jonghee W. Yoon, Aviral Shrivastava, Minwook Ahn, Reiley Jeyapaul, Yunheung Paek |
ASP-DAC | 6 |
| 2008 | Hiding Cache Miss Penalty Using Priority-based Execution for Embedded ProcessorsabstractThe contribution of memory latency to execution time continues to increase, and latency hiding mechanisms become ever more important for efficient processor design. While high-end processors can use elaborate techniques like multiple issue, out-of-order execution, speculative execution, value prediction etc. to tolerate high memory latencies, they are often not viable solutions for embedded processors, due to significant area, power and chip complexity overheads. This paper proposes a hardware-software cooperative approach, called priority-based execution to hide cache miss penalty for embedded processors. The compiler classifies the instructions into low-priority and high-priority instructions. The processor executes the high-priority instructions, but delays the execution of low priority instructions. They are executed on a cache miss to hide the cache miss penalty. We empirically evaluate our proposal on the Intel XScale compiler and microarchitecture. Experimental results on benchmarks from Multimedia, MediaBench, MiBench, and SPEC2000 demonstrate an average 17% performance improvements, hiding 75% cache miss penalty. Aviral Shrivastava, Yunheung Paek |
DATE | 3 |
| 2008 | Compiler driven data layout optimization for regular/irregular array access patternsabstractEmbedded multimedia applications consist of regular and irregular memory access patterns. Particularly, irregular pattern are not amenable to static analysis for extraction of access patterns, and thus prevent efficient use of a Scratch Pad Memory (SPM) hierarchy for performance and energy improvements. To resolve this, we present a compiler strategy to optimize data layout in regular/irregular multimedia applications running on embedded multiprocessor environments. The goal is to maximize the amount of accesses to the SPM over the entire system which leads to a reduction in the energy consumption of the system. This is achieved by optimizing data placement of application-wide reused data so that it resides in the SPMs of processing elements. Specifically, our scheme is based on a profiling that generates a memory access footprint. The memory access footprint is used to identify data elements with fine granularity that can profitably be placed in the SPMs to maximize performance and energy gains. We present a heuristic approach that efficiently exploits the SPMs using memory access footprint. Our experimental results show that our approach is able to reduce energy consumption by 30% and improve performance by 18% over cache based memory subsystems for various multimedia applications. Doosan Cho, Sudeep Pasricha, Ilya Issenin, Nikil Dutt, Yunheung Paek, SunJun Ko |
LCTES | 5 |
| 2008 | Register File Power Reduction Using Bypass Sensitive CompilerabstractThis paper explores, develops, and investigates several bypass-sensitive compilation techniques to reduce the register file power by reducing the access frequency to the register file. We study the effectiveness of our techniques on the Intel XScale processor, which is based on the previously proposed ldquoon-demand register fetch readrdquo architectural feature. Furthermore, we show that our bypass-sensitive compilation technique is effective on various partial bypass configurations. Aviral Shrivastava, Nikil Dutt, Alexandru Nicolau, Yunheung Paek, Eugene Earlie |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2008 | A retargetable parallel-programming framework for MPSoCabstractAs more processing elements are integrated in a single chip, embedded software design becomes more challenging: It becomes a parallel programming for nontrivial heterogeneous multiprocessors with diverse communication architectures, and design constraints such as hardware cost, power, and timeliness. In the current practice of parallel programming with MPI or OpenMP, the programmer should manually optimize the parallel code for each target architecture and for the design constraints. Thus, the design-space exploration of MPSoC (multiprocessor systems-on-chip) costs become prohibitively large as software development overhead increases drastically. To solve this problem, we develop a parallel-programming framework based on a novel programming model called common intermediate code (CIC). In a CIC, functional parallelism and data parallelism of application tasks are specified independently of the target architecture and design constraints. Then, the CIC translator translates the CIC into the final parallel code, considering the target architecture and design constraints to make the CIC retargetable. Experiments with preliminary examples, including the H.263 decoder, show that the proposed parallel-programming framework increases the design productivity of MPSoC software significantly. Seongnam Kwon, Woo-Chul Jeun, Soonhoi Ha, Yunheung Paek |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2007 | Software controlled memory layout reorganization for irregular array access patternsabstractMany embedded array-intensive applications have irregular access patterns that are not amenable to static analysis for extraction of access patterns, and thus prevent efficient use of a Scratch Pad Memory (SPM) hierarchy for performance and power improvement. We present a profiling based strategy that generates a memory access trace which can be used to identify data elements with fine granularity that can profitably be placed in the SPMs to maximize performance and energy gains. We developed an entire toolchain that allows incorporation of the code required to profitably move data to SPMs; visualization of the extracted access pattern after profiling; and evaluation/exploration of the generated application code to steer mapping of data to the SPM to yield performance and energy benefits.We present a heuristic approach that efficiently exploits the SPM using the profiler-driven access pattern behaviors. Experimental results on EEMBC and other industrial codes obtained with our framework show that we are able to achieve 36% energy reduction and reduce execution time by up to 22% compared to a cache based system. Doosan Cho, Ilya Issenin, Nikil Dutt, Jonghee W. Yoon, Yunheung Paek |
CASES | 5 |
| 2007 | Preprocessing Strategy for Effective Modulo Scheduling on Multi-issue Digital Signal Processors
Doosan Cho, Ravi Ayyagari, Gang-Ryung Uh, Yunheung Paek |
CC | 4 |
| 2007 | Optimistic coalescing for heterogeneous register architecturesabstractIn this paper, Optimistic coalescing has been proven as an elegant and effective technique that provides better chances of safely coloring more registers in register allocation than other coalescing techniques. Its algorithm originally assumes homogeneous registers which are all gathered in the same register file. Although this register architecture is still common in most general-purpose processors, embedded processors often contain heterogeneous registers which are scattered in physically different register files dedicated for each dissimilar purpose and use. In this work, we developed a modified algorithm for optimal coalescing that helps a register allocator for an embedded processor to better handle such heterogeneity of the register architecture. In the experiment, an existing register allocator was able to achieve up to 10% reduction in code size through our coalescing, and avoid many spills that would have been generated without our scheme. Minwook Ahn, Jooyeon Lee, Yunheung Paek |
LCTES | 3 |
| 2007 | Efficient embedded code generation with multiple load/store instructionsabstractAbstract In a recent study, we discovered that many single load/store operations in embedded applications can be parallelized and thus encoded simultaneously in a single‐instruction multiple‐data instruction, called the multiple load/store (MLS) instruction. In this work, we investigate the problem of utilizing MLS instructions to produce optimized machine code, and propose an effective approach to the problem. Specifically, we formalize the MLS problem, that is, the problem of maximizing the use of MLS instructions with an unlimited register file size. Based on this analysis, we show that we can solve the problem efficiently by translating it into a variant of the problem finding a maximum weighted path cover in a dynamic weighted graph. To handle a more realistic case of the finite size of the register file, our solution is then extended to take into account the constraints of register sequencing in MLS instructions and the limited register resource available in the target processor. We demonstrate the effectiveness of our approach experimentally by using a set of benchmark programs. In summary, our approach can reduce the number of loads/stores by 13.3% on average, compared with the code generated from existing compilers. The total code size reduction is 3.6%. This code size reduction comes at almost no cost because the overall increase in compilation time as a result of our technique remains quite minimal. Copyright © 2007 John Wiley & Sons, Ltd. Yunheung Paek, Minwook Ahn, Doosan Cho, Taehwan Kim 0007 |
Softw. Pract. Exp. | 1 |
| 2007 | Automatic Design Space Exploration of Register Bypasses in Embedded ProcessorsabstractRegister bypassing is a popular and powerful architectural feature to improve processor performance in pipelined processors by eliminating certain data hazards. However, extensive bypassing comes with a significant impact on cycle time, area, and power consumption of the processor. Recent research therefore advocates the use of partial bypassing in a processor. However, accurate performance evaluation of partially bypassed processors is still a challenge, primarily due to the lack of bypass-sensitive retargetable compilation techniques. No existing partial bypass exploration framework estimates the power and area overhead of partial bypassing. As a result, the designers end up making suboptimal design decisions during the exploration of partial bypass design space. This paper presents PBExplore - an automatic design-space-exploration framework for register bypasses. PBExplore accurately evaluates the performance of a partially bypassed processor using a bypass-sensitive compilation technique. It synthesizes the bypass control logic and estimates the area and energy overhead of each bypass configuration. PBExplore is thus able to effectively perform multidimensional exploration of the partial bypass design space. We present experimental results of benchmarks from the MiBench suite on the Intel XScale architecture on and demonstrate the need, utility, and exploration capabilities of PBExplore. Aviral Shrivastava, Eugene Earlie, Nikil Dutt, Alexandru Nicolau, Yunheung Paek |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2007 | Introduction to the special LCTES'05 issueabstractNo abstract available. Rajiv Gupta 0001, Yunheung Paek |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2006 | A spatial mapping algorithm for heterogeneous coarse-grained reconfigurable architecturesabstractIn this work, we investigate the problem of automatically mapping applications onto a coarse-grained reconfigurable architecture and propose an efficient algorithm to solve the problem. We formalize the mapping problem and show that it is NP-complete. To solve the problem within a reasonable amount of time, we divide it into three subproblems: covering, partitioning and layout. Our empirical results demonstrate that our technique produces nearly as good performance as hand-optimized outputs for many kernels. Minwook Ahn, Jonghee W. Yoon, Yunheung Paek, Yoonjin Kim, Mary Kiemb, Kiyoung Choi |
DATE | 3 |
| 2006 | Automatic generation of operation tables for fast exploration of bypasses in embedded processorsabstractCustomizing the bypasses in an embedded processor uncovers valuable trade-offs between the power, performance and the cost of the processor. Meaningful exploration of bypasses requires bypass-sensitive compiler. Operation tables (OTs) have been proposed to perform bypass-sensitive compilation. However, due to lack of automated methods to generate OTs, OTs are currently manually specified by the designer. Manual specification of OTs is not only an extremely time consuming task, but is also highly error-prone. In this paper, we present AutoOT, an algorithm to automatically generate OTs from a high-level processor description. Our experiments on the Intel XScale processor model running MiBench benchmarks demonstrate that AutoOT greatly reduces the time and effort of specification. Automatic generation of OTs makes it feasible to perform full bypass exploration on the Intel XScale and thus discover interesting alternate bypass configurations in a reasonable time. To further reduce the compile-time overhead of OT generation, we propose another novel algorithm, AutoOTDB. AutoOTDB is able to cut the compile-time overhead of OT generation by half Eugene Earlie, Aviral Shrivastava, Alexandru Nicolau, Nikil Dutt, Yunheung Paek |
DATE | 6 |
| 2006 | Power-conscious configuration cache structure and code mapping for coarse-grained reconfigurable architectureabstractCoarse-grained reconfigurable architecture aims to achieve both performance and flexibility. However, power consumption is no less important for the reconfigurable architecture to be used as a competitive processing core in embedded systems. In this paper, we show how power is consumed in a typical coarse-grained reconfigurable architecture. Based on the power breakdown data, we suggest a power-conscious configuration cache structure and code mapping technique, which reduce power consumption without performance degradation. Experimental results show that the proposed approach saves much power even with reduced configuration cache size. Yoonjin Kim, Ilhyun Park, Kiyoung Choi, Yunheung Paek |
ISLPED | 4 |
| 2006 | Bypass aware instruction scheduling for register file power reductionabstractSince register files suffer from some of the highest power densities within processors, designers have investigated several architectural strategies for register file power reduction, including "On Demand RF Read" where the register file is read only if the operand value is not available from the bypasses. However, we show in this paper that significant additional reductions in the register file power consumption can be obtained by scheduling instructions so that they transfer the operands via bypasses, rather than reading from the register file. Such instruction scheduling requires the compiler to be cognizant of the bypasses in the processor pipeline. In this paper, we develop several bypass aware instruction scheduling heuristics varying in time complexity, and study their effectiveness on the Intel XScale processor pipeline running MiBench benchmarks. Our experimental results show additional power consumption reductions of up to 26% and on average 12% over and above the register file power reduction achieved through existing techniques. Aviral Shrivastava, Nikil Dutt, Alexandru Nicolau, Yunheung Paek, Eugene Earlie |
LCTES | 5 |
| 2006 | VISTA: VPO interactive system for tuning applicationsabstractSoftware designers face many challenges when developing applications for embedded systems. One major challenge is meeting the conflicting constraints of speed, code size, and power consumption. Embedded application developers often resort to hand-coded assembly language to meet these constraints since traditional optimizing compiler technology is usually of little help in addressing this challenge. The results are software systems that are not portable, less robust, and more costly to develop and maintain. Another limitation is that compilers traditionally apply the optimizations to a program in a fixed order. However, it has long been known that a single ordering of optimization phases will not produce the best code for every application. In fact, the smallest unit of compilation in most compilers is typically a function and the programmer has no control over the code improvement process other than setting flags to enable or disable certain optimization phases. This paper describes a new code improvement paradigm implemented in a system called VISTA that can help achieve the cost/performance trade-offs that embedded applications demand. The VISTA system opens the code improvement process and gives the application programmer, when necessary, the ability to finely control it. VISTA also provides support for finding effective sequences of optimization phases. This support includes the ability to interactively get static and dynamic performance information, which can be used by the developer to steer the code improvement process. This performance information is also internally used by VISTA for automatically selecting the best optimization sequence from several attempted. One such feature is the use of a genetic algorithm to search for the most efficient sequence based on specified fitness criteria. We include a number of experimental results that evaluate the effectiveness of using a genetic algorithm in VISTA to find effective optimization phase sequences. Prasad A. Kulkarni, Wankang Zhao, Stephen Roderick Hines, David B. Whalley, Xin Yuan 0001, Robert A. van Engelen, Kyle A. Gallivan, Jason Hiser, Jack W. Davidson, Baosheng Cai, Mark W. Bailey, Hwashin Moon, Kyunghwan Cho, Yunheung Paek |
ACM Trans. Embed. Comput. Syst. | 14 |
| 2005 | Compiler transformations for effectively exploiting a zero overhead loop bufferabstractAbstract A Zero Overhead Loop Buffer (ZOLB) is an architectural feature that is commonly found in DSP processors. This buffer can be viewed as a compiler managed cache that contains a sequence of instructions that will be executed a specified number of times without incurring any loop overhead. Unlike loop unrolling, a loop buffer can be used to minimize loop overhead without the penalty of increasing code size. In addition, a ZOLB requires relatively little space and power, which are both important considerations for most DSP applications. This paper describes strategies for generating code to effectively use a ZOLB. We have found that many common code improving transformations used by optimizing compilers on conventional architectures can be easily used to (1) allow more loops to be placed in a ZOLB, (2) further reduce loop overhead of the loops placed in a ZOLB, and (3) avoid redundant loading of ZOLB loops. The results given in this paper demonstrate that this architectural feature can often be exploited with substantial improvements in execution time and slight reductions in code size for various signal processing applications. Copyright © 2004 John Wiley & Sons, Ltd. Gang-Ryung Uh, David B. Whalley, Sanjay Jinturkar, Yunheung Paek, Vincent Cao |
Softw. Pract. Exp. | 5 |
| 2004 | Code optimizations for a VLIW-style network processing unitabstractAbstract The explosive growth in network bandwidth and Internet services such as QoS (quality of service) and SLA (service level agreement) monitoring have created the need for new networking hardware called aNetwork Processing Unit (NPU). In order to rapidly reconfigure the NPU for frequently varying Internet services and technologies, a high‐performance C compiler is urgently needed. Several code generation techniques, which are intended to meet the high code quality demands of other types ofapplication specific instruction‐set processors(ASIPs) likedigital signal processors(DSPs), have already been developed. However, these techniques are insufficient for NPUs due to striking architectural differences such as asymmetric data paths. The main purpose of this paper is to discuss our recent experience with the development of a commercial compiler for a new NPU called thePaion PPII, which is basically apacket enginefor NPU to meet the growing need for new high‐bandwidth communication equipment targeted for Internet routers and ethernet adapters. For this purpose, we will first show the architectural challenges posed by the target NPU. Then, we will describe several compiler techniques that we found to be effective for the target NPU with various unorthogonal architectural features. The current implementations of the PPII use a VLIW (Very Long Instruction Word) architecture. So, we handled this VLIW‐style architecture by employing a simplecode compactionscheme which packs multiple parallel instructions into one long instruction word. The experimental results show that our techniques are effective for significantly reducing the dynamic instruction count. Copyright © 2004 John Wiley & Sons, Ltd. Jinhwan Kim, Yunheung Paek, Gang-Ryung Uh |
Softw. Pract. Exp. | 2 |
| 2004 | Fast memory bank assignment for fixed-point digital signal processorsabstractMost vendors of digital signal processors (DSPs) support a Harvard architecture, which has two or more memory buses, one for program and one or more for data and allow the processor to access multiple words of data from memory in a single instruction cycle. Also, many existing fixed-point DSPs are known to have an irregular architecture with heterogeneous registers, which contains multiple register files that are distributed and dedicated to different sets of instructions. Although there have been several studies conducted to efficiently assign data to multimemory banks, most of them assumed processors with relatively simple, homogeneous general-purpose registers. Thus, several vendor-provided compilers for DSPs that we examined were unable to efficiently assign data to multiple data memory banks, thereby often failing to generate highly optimized code for their machines. As a consequence, programmers for these DSPs often manually assign program variables to memories so as to fully utilize multimemory banks in their code. This paper reports on our recent attempt to address this problem by presenting an algorithm that helps the compiler to efficiently assign data to multimemory banks. Our algorithm differs from previous work in that it assigns variables to memory banks in separate, decoupled code generation phases, instead of a single, tightly coupled phase. The experimental results have revealed that our decoupled algorithm greatly simplifies our code generation process; thus our compiler runs extremely fast, yet generates target code that is comparable in quality to the code generated by a coupled approach. Jeonghun Cho 0001, Yunheung Paek, David B. Whalley |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2003 | A Quantitative Comparison of Two Retargetable Compilation ApproachesabstractIn the design of an embedded processor, the compiler design is tightly coupled with the underlying processor architecture, and thus it is crucial to rapidly retarget a compiler along with the change in the architecture in order to expedite the processor design. However, among many compiler writers, there is a controversial issue that has long been argued - whether a compiler can be easily retargetable while it performs sophisticated machine-specific optimizations for a new architecture configuration. We examine this issue by finding some possible cases where optimizations may be impeded on a pathway to building a retargetable compiler. For this, we developed two types of compilation frameworks called user-retargetable and developer-retargetable, and compared their performance. Sejong Oh, Yunheung Paek |
ICPP | 2 |
| 2003 | Finding effective optimization phase sequencesabstractIt has long been known that a single ordering of optimization phases will not produce the best code for every application. This phase ordering problem can be more severe when generating code for embedded systems due to the need to meet conflicting constraints on time, code size, and power consumption. Given that many embedded application developers are willing to spend time tuning an application, we believe a viable approach is to allow the developer to steer the process of optimizing a function. In this paper, we describe support in VISTA, an interactive compilation system, for finding effective sequences of optimization phases. VISTA provides the user with dynamic and static performance information that can be used during an interactive compilation session to gauge the progress of improving the code. In addition, VISTA provides support for automatically using performance information to select the best optimization sequence among several attempted. One such feature is the use of a genetic algorithm to search for the most efficient sequence based on specified fitness criteria. We have included a number of experimental results that evaluate the effectiveness of using a genetic algorithm in VISTA to find effective optimization phase sequences. Prasad A. Kulkarni, Wankang Zhao, Hwashin Moon, Kyunghwan Cho, David B. Whalley, Jack W. Davidson, Mark W. Bailey, Yunheung Paek, Kyle A. Gallivan |
LCTES | 8 |
| 2003 | Case Studies on Automatic Extraction of Target-Specific Architectural Parameters in Complex Code Generation
Yunheung Paek, Minwook Ahn, Soonho Lee |
SCOPES | 1 |
| 2002 | Experience with a retargetable compiler for a commercial network processorabstractThe Paion PPII network processor is designed to meet the growing need for new high bandwidth network equipment. In order to rapidly reconfigure the processor for frequently varying internet services and technologies, a high performance compiler is urgently needed. Albeit various code generation techniques have been proposed for DSPs or ASIPs, we experienced these techniques are not easily tailored towards the target Paion PPII processor due to striking architectural differences. First, we will show the architectural challenges posed by the target processor. Second, novel compiler techniques will be described that effectively exploit unorthogonal architectural features. The techniques include virtual data path, compiler intrinsics, and interprocedural register allocation. Third, intermediate benchmark results will be presented to demonstrate the effectiveness of our techniques. Jinhwan Kim, Sungjoon Jung, Yunheung Paek, Gang-Ryung Uh |
CASES | 3 |
| 2002 | A proof method for the correctness of modularized 0CFA
Oukseh Lee, Kwangkeun Yi, Yunheung Paek |
Inf. Process. Lett. | 3 |
| 2002 | Efficient and precise array access analysisabstractA number of existing compiler techniques hinge on the analysis of array accesses in a program. The most important task in array access analysis is to collect the information about array accesses of interest and summarize it in some standard form. Traditional forms used in array access analysis are sensitive to the complexity of array subscripts; that is, they are usually quite accurate and efficient for simple array subscripting expressions, but lose accuracy or require potentially expensive algorithms for complex subscripts. Our study has revealed that in many programs, particularly numerical applications, many access patterns are simple in nature even when the subscripting expressions are complex. Based on this analysis, we have developed a new, general array region representational form, called the linear memory access descriptor (LMAD). The key idea of the LMAD is to relate all memory accesses to the linear machine memory rather than to the shape of the logical data structures of a programming language. This form helps us expose the simplicity of the actual patterns of array accesses in memory, which may be hidden by complex array subscript expressions. Our recent experimental studies show that our new representation simplifies array access analysis and, thus, enables efficient and accurate compiler analysis. Yunheung Paek, Jay P. Hoeflinger, David A. Padua |
ACM Trans. Program. Lang. Syst. | 1 |
| 2002 | An Advanced Compiler Framework for Non-Cache-Coherent MultiprocessorsabstractThe Cray T3D and T3E are non-cache-coherent (NCC) computers with a NUMA structure. They have been shown to exhibit a very stable and scalable performance for a variety of application programs. Considerable evidence suggests that they are more stable and scalable than many other shared-memory multiprocessors. However, the principal drawback of these machines is a lack of programmability, caused by the absence of the global cache coherence that is necessary to provide a convenient shared view of memory in hardware. This forces the programmer to keep careful track of where each piece of data is stored, a complication that is unnecessary when a pure shared-memory view is presented to the user. We believe that a remedy for this problem is advanced compiler technology. In this paper, we present our experience with a compiler framework for automatic parallelization and communication generation that has the potential to reduce the time-consuming hand-tuning that would otherwise be necessary to achieve good performance with this type of machine. From our experiments, we learned that our compiler performs well for a variety of applications on the T3D and T3E and we found a few sophisticated techniques that could improve performance even more once they are fully implemented in the compiler. Yunheung Paek, Angeles G. Navarro, Emilio L. Zapata, Jay P. Hoeflinger, David A. Padua |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2001 | The very portable optimizer for digital signal processorsabstractAlthough retargetability has been a major design concern for many compilers, retargetability is a vitally important issue for Digital Signal Processors(DSPs) because the architectural variations of DSPs are much wider than those of General-Purpose Processors (GPPs). This paper describes our preliminary study on a retargetable code generator, called the Very Portable Optimizer (VPO), that has been recently engineered to target DSPs. The compiler was originally implemented to target GPPs, but thanks to its novel intermediate representation designed to support retargetability, it has been successfully retargeted to a commercial DSP within a relatively short period of time. This retargetable compiler study is at its early stage, so the code quality is still amenable to further improvement. However, our recent enhancements to VPO helped us to achieve quite encouraging results when being compared with a production-quality, highly-customized, compiler. Sungjoon Jung, Yunheung Paek |
CASES | 2 |
| 2001 | A Parallel Programming Environment for a V-Busbased PC-clusteabstractNowadays PC-cluster architectures are widely accepted for parallel computing. In a PC-cluster system, memories are physically distributed. To harness the computational power of a distributed-memory PC-cluster, a user must write efficient software for the machine by hand. The absence of global address space makes the process laborious because the user must manually assign computations to processors, distribute data across processors and explicitly manage communication among processors. This paper introduces our recent three related studies on parallel computing. The first is a V-Bus based PC-cluster in which all PCs are interconnected through V-Bus network cards. The second is the implementation of an one-sided communication library on our PC-cluster system, which provides the user with a view of global address space on top of the distributed-memory PC-cluster. This encapsulation of global address allows the user to write shared memory code on our cluster system, which simplifies programming our cluster because the user does not need to explicitly cope with distributed memories. The third is a parallelizing compiler that automatically translates sequential code to shared-memory code for the PC-cluster. This compiler not only enables legacy code to be compiled for our cluster, but also further simplifies programming our cluster by allowing the user to continue using conventional sequential languages such as Fortran 77. In this work, the compiler was optimized particularly for our V-Bus based PC-cluster. The paper also reports our experimental results with the compiler on our PC-cluster system. Sang Seok Lim, Yunheung Paek, Jay P. Hoeflinger |
CLUSTER | 2 |
| 1999 | Communication Studies of Single-Threaded and Multithreaded Distributed-Memory MultiprocessorsabstractThis report explicates the communication overlapping capabilities of three distributed-memory machines, SGI/Cray T3E, IBM SP-2 with wide nodes, and the ETL EM-X. Bitonic sorting and Fast Fourier Transform are selected for experiments. Various message sizes are used to determine when, where, how much and why the overlapping takes place. Experimental results with up to 64 processors indicated that the communication performance of EM-X is insensitive to various message sizes while SP-2 is the most sensitive. T3E stayed in between. The EM-X gave the highest communication overlapping capability while T3E did the lowest. The experimental results are compared with the analytical results based on LogP and LogGP communication models. Andrew Sohn, Yunheung Paek, Jui-Yuan Ku, Yuetsu Kodama, Yoshinori Yamaguchi |
HPCA | 2 |
| 1998 | Simplification of Array Access Patterns for Compiler OptimizationsabstractExisting array region representation techniques are sensitive to the complexity of array subscripts. In general, these techniques are very accurate and efficient for simple subscript expressions, but lose accuracy or require potentially expensive algorithms for complex subscripts. We found that in scientific applications, many access patterns are simple even when the subscript expressions are complex. In this work, we present a new, general array access representation and define operations for it. This allows us to aggregate and simplify the representation enough that precise region operations may be applied to enable compiler optimizations. Our experiments show that these techniques hold promise for speeding up applications. Yunheung Paek, Jay P. Hoeflinger, David A. Padua |
PLDI | 1 |
| 1997 | Compiler Techniques for Effective Communication on Distributed-Memory MultiprocessorsabstractThe Polaris restructurer transforms conventional Fortran programs into parallel form for various types of multiprocessor systems. This paper presents the results of a study on strategies to improve the effectiveness of Polaris' techniques for distributed-memory multiprocessors. Our study, which is based on the hand analysis of MDG and TRFD from the Perfect Benchmarks and TOYCATV and SWIM from SPEC benchmarks, identified three techniques that are important for improving communication optimization. Their application produces almost perfect speedups for the four programs on the Cray T3D. Angeles G. Navarro, Emilio L. Zapata, Yunheung Paek, David A. Padua |
ICPP | 3 |