Daehoon Kim 0001

dblp:97/7677-1 · DBLP profile ↗
← Back
26ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0003-0837-0877ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 5 first-author · 14 since 2021Software engineering, systems software and programming languages · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ARIADNE: Adaptive UVM Management for Efficient GPU Memory Oversubscription
abstract
Unified Virtual Memory (UVM) simplifies GPU programming and supports memory oversubscription, but suffers from severe performance degradation under high memory pressure due to page fault overhead and thrashing. Existing approaches such as prefetching, access counter-based migration, and dynamic Zero-copy offer limited benefits and often require hardware or compiler modifications, undermining UVM's portability and ease of deployment. We present ARIADNE, a runtime UVM management framework that preserves UVM's GPU memory abstraction while ensuring high and robust performance under memory oversubscription. ARIADNE is guided by three principles: (1) pipelined fault handling to hide migration latency, (2) Sharing Degree, a runtime metric that captures thread-level access locality without requiring hardware or compiler changes, to inform placement decisions, and (3) dynamic placement of memory regions between GPU memory and Zero-copy based on real-time access patterns. Implemented entirely within NVIDIA's UVM driver, ARIADNE requires no recompilation or hardware modifications and applies transparently to any executable or closed-source GPU UVM applications. Our experimental results show that ARIADNE delivers average speedups of$1.9 \times, 5.0 \times$, and$4.8 \times$over a state-of-the-art method at$\text{1 3 0 \%, ~} \text{1 7 5 \%}$, and 300 % oversubscription, respectively, while effectively preventing thrashing and maintaining near-linear performance scaling.
Hyunkyun Shin, Seongtae Bang, Hyungwon Park 0001, Daehoon Kim 0001
HPCA4
2026 DarkStream: Exploiting Internal Throughput Contention in Data Streaming Accelerator for Timing Attacks
Hyosang Kim, Ki-Dong Kang, Gyeongseo Park, Sungju Kim, Daehoon Kim 0001
ISCA5
2026 SMOOTH: Hardware-Assisted Fine-Grained On-Chip Memory Management for Efficient On-Device LLM Inference
Seulki Kim, Bokyeong Kim, Kyeonghyeon Ryu, Yeji Jung, Hwanjun Lee, Sungju Kim, Yunhyeong Jeon, Daehoon Kim 0001
ISCA8
2025 Co-UP: Comprehensive Core and Uncore Power Management for Latency-Critical Workloads
abstract
Improving energy efficiency to reduce costs in server environments has attracted considerable attention. Considering that processors account for a significant portion of energy consumption in servers, Dynamic Voltage and Frequency Scaling (DVFS) enhances their energy efficiency by adjusting the operational speed and power consumption of processors. Additionally, modern high-end processors extend DVFS functionality not only to core components but also to uncore parts. This is because the increasing complexity and integration of System on Chips (SoCs) have highlighted the substantial energy consumption. However, existing uncore voltage/frequency scaling fails to effectively consider Latency-Critical (LC) applications, leading to sub-optimal energy efficiency or degraded performance. In this paper, we introduce Co-UP, power management that simultaneously scales core and uncore frequencies for latency-critical applications, designed to improve energy efficiency without violating Service Level Objectives (SLOs). To this end, Co-UP incorporates a prediction model that estimates outcomes of energy consumption and performance as uncore and core frequency changes. Based on the estimated gains, Co-UP adjusts to uncore and/or core frequencies to further enhance energy efficiency or performance. This predictive model can rapidly adapt to new and unlearned loads, enabling Co-UP to operate online without any prior profiling. Our experiments show that Co-UP can reduce energy consumption by up to 28.2% compared to existing Intel's policy and up to 17.6% compared to state-of-the-art power management studies, without SLO violations.
Ki-Dong Kang, Gyeongseo Park, Daehoon Kim 0001
DATE3
2025 BrokenSleep: Remote Power Timing Attack Exploiting Processor Idle States
abstract
Power and energy consumption emerge as critical aspects in computing systems, spanning from mobile devices to data-center servers. Modern processors typically support idle states (i.e., C-states), which deactivate specific hardware components, in addition to offering multiple voltage and frequency states (i.e., P-states). While C-states can significantly reduce static power when processor cores are idle, a notable security vulnerability arises due to differences in wake-up latency among various C-states when the processor cores become active again. This paper proposes a security vulnerability arising from processor idle state management, called BrokenSleep, which exploits the aforementioned wake-up latency differences to create covert and side-channel between computing nodes connected via an external network. This study presents the first remote timing attack based on power management, overcoming the limitations of previous research that required the co-location of attacker and victim applications on the same local machine. This advancement significantly extends the range of existing remote timing attacks by integrating power-related factors. Regardless of the computing system types, our experiments demonstrate that an attacker can transfer data to remote machines without direct network access and deduce the keystroke timing. This vulnerability is not confined to a single processor architecture; it affects processors designed by both Intel and ARM, indicating a widespread potential risk across different hardware platforms.
Hyosang Kim, Ki-Dong Kang, Gyeongseo Park, Seungkyu Lee 0002, Daehoon Kim 0001
HPCA5
2025 Jack Unit: An Area- and Energy-Efficient Multiply-Accumulate (MAC) Unit Supporting Diverse Data Formats
abstract
In this work, we introduce an area- and energy-efficient multiply-accumulate (MAC) unit, named Jack Unit, that is a jack-of-all-trades, supporting various data formats such as integer (INT), floating point (FP), and microscaling data format (MX). It provides bit-level flexibility and enhances hardware efficiency by i) replacing the carry-save multiplier (CSM) in the FP multiplier with a precision-scalable CSM, ii) performing the adjustment of significands based on the exponent differences within the CSM, and iii) utilizing 2D sub-word parallelism. To assess effectiveness, we implemented the layout of the Jack unit and three baseline MAC units. Additionally, we designed an AI accelerator equipped with our Jack units to compare with a state-of-the-art AI accelerator supporting various data formats. The proposed MAC unit achieves an area reduction of 14.53∼50.25% and a power reduction of 4.76 ∼ 45.65% compared to the baseline MAC units. On five AI benchmarks, the accelerator de-signed with our Jack units improves energy efficiency by 1.32 ∼ 5.41× over the baseline across various data formats.
Seock-Hwan Noh, Sungju Kim, Daehoon Kim 0001, Jaeha Kung 0001, Yeseong Kim
ISLPED4
2025 Beyond Page Migration: Enhancing Tiered Memory Performance via Integrated Last-Level Cache Management and Page Migration
Hwanjun Lee, Yeji Jung, Seonmu Oh, Ki-Dong Kang, Seunghak Lee, Daehoon Kim 0001
MICRO7
2025 EcoCore: Dynamic Core Management for Improving Energy Efficiency in Latency-Critical Applications
Gyeongseo Park, Ki-Dong Kang, Yunhyeong Jeon, Seulki Kim, Daehoon Kim 0001
MICRO6
2025 MTAT: Adaptive Fast Memory Management for Co-located Latency-Critical Workloads in Tiered Memory System
abstract
Modern data centers increasingly employ multi-tenant deployment models in which multiple applications or virtual machines share a single physical server. However, existing tiered memory management schemes classify pages solely by access frequency to govern promotions and demotions across memory tiers without accounting for the distinct access patterns of latency-critical (LC) and best-effort (BE) workloads. LC workloads demand low-latency service yet lack sustained high-frequency access; consequently, frequency-based tiering demotes LC data to slower memory (SMem), degrading responsiveness and violating service-level objectives (SLOs).
Seonggyu Han, Gyeongseo Park, Daehoon Kim 0001
Middleware4
2025 Para-ksm: Parallelized Memory Deduplication with Data Streaming Accelerator
Houxiang Ji, Seonmu Oh, Daehoon Kim 0001, Nam Sung Kim
USENIX ATC4
2024 vSPACE: Supporting Parallel Network Packet Processing in Virtualized Environments through Dynamic Core Management
abstract
Data centers face significant performance challenges with parallel processing for network I/O in virtualized environments, particularly for latency-critical (LC) workloads that must satisfy strict Service Level Objectives (SLOs). While previous studies have addressed performance challenges in network I/O virtualization, they overlook the impact of excessive parallelism on the performance of Virtual Machines (VMs). We observe that excessive parallelization for VMs and network I/O processing can lead to core oversubscription, resulting in significant resource contention, frequent preemptions, and task migrations. Based on these observations, we propose vSPACE, dynamic core management specifically designed to support parallel network I/O processing in virtualized environments efficiently. To reduce scheduling contention, vSPACE creates distinct core allocation groups for VM and network I/O and assigns dedicated cores to each. Then, it dynamically adjusts the number of allocated cores to enforce appropriate parallelism for VMs and network I/O processing based on varying demands. vSPACE employs continuous monitoring and a heuristic algorithm to periodically determine appropriate core allocation, addressing excessive contention and improving energy and resource efficiency. vSPACE operates in three modes: performance improvement, energy efficiency, and resource efficiency. Our evaluations demonstrate that vSPACE significantly enhances throughput by up to 4.2 × compared to existing core allocation approaches and improves energy and resource efficiency by up to 16.5% and 30.5%, respectively.
Gyeongseo Park, Ki-Dong Kang, Yunhyeong Jeon, Sungju Kim, Hyosang Kim, Daehoon Kim 0001
PACT7
2022 IDIO: Network-Driven, Inbound Network Data Orchestration on Server Processors
abstract
High-bandwidth network interface cards (NICs), each capable of transferring 100s of Gigabits per second, are making inroads into the servers of next-generation datacenters. Such unprecedented data delivery rates impose immense pressure, especially on the server’s memory subsystem, as NICs first transfer network data to DRAM before processing. To alleviate the pressure, the cache hierarchy has evolved, supporting a direct data I/O (DDIO) technology to directly place network data in the last-level cache (LLC). Subsequently, various policies have been explored to manage such LLC and have proven to effectively reduce service latency and memory bandwidth consumption of network applications. However, the more recent evolution of the cache hierarchy decreased the size of LLC per core but significantly increased that of midlevel cache (MLC) with a non-inclusive policy. This calls for a re-examination of the aforementioned DDIO technology and management policies. In this paper, first, we identify three shortcomings of the current static data placement policy placing network data to LLC first and the non-inclusive policy with a commercial server system: (1) ineffectively using large MLC, (2) suffering from high rates of writebacks from MLC to LLC, and (3) breaking the isolation between application and network data enforced by limiting cache ways for DDIO. Second, to tackle the three shortcomings, we propose an intelligent direct I/O (IDIO) technology that extends DDIO to MLC and provides three synergistic mechanisms: (1) self-invalidating I/O buffer, (2) network-driven MLC prefetching, and (3) selective direct DRAM access. Our detailed experiments using a full-system simulator — capable of running modern DPDK userspace network functions while sustaining 100Gbps + network bandwidth — show that IDIO significantly reduces data movement (up to 84% MLC and LLC writeback reduction), provides LLC isolation (up to 22% performance improvement), and improves tail latency (up to 38% reduction in 99thlatency) for receive-intensive network applications.
Mohammad Alian, Siddharth Agarwal, Daehoon Kim 0001, Ren Wang 0001, Nam Sung Kim
MICRO6
2021 InnerSP: A Memory Efficient Sparse Matrix Multiplication Accelerator with Locality-Aware Inner Product Processing
abstract
Sparse matrix multiplication is one of the key computational kernels in large-scale data analytics. However, a naive implementation suffers from the overheads of irregular memory accesses due to the representation of sparsity. To mitigate the memory access overheads, recent accelerator designs advocated the outer product processing which minimizes input accesses but generates intermediate products to be merged to the final output matrix. Using real-world sparse matrices, this study first identifies the memory bloating problem of the outer product designs due to the unpredictable intermediate products. Such an unpredictable increase in memory requirement during computation can limit the applicability of accelerators. To address the memory bloating problem, this study revisits an alternative inner product approach, and proposes a new accelerator design called InnerSP. This study shows that nonzero element distributions in real-world sparse matrices have a certain level of locality. Using a smart caching scheme designed for inner product, the locality is effectively exploited with a modest on-chip cache. However, the row-wise inner product relies on on-chip aggregation of intermediate products. Due to uneven sparsity per row, overflows or underflows of the on-chip storage for aggregation can occur. To maximize the parallelism while avoiding costly overflows, the proposed accelerator uses pre-scanning for row splitting and merging. The simulation results show that the performance of InnerSP can exceed or be similar to those of the prior outer product approaches without any memory bloating problem.
Daehyeon Baek, Soojin Hwang, Taekyung Heo, Daehoon Kim 0001, Jaehyuk Huh 0001
PACT4
2021 NMAP: Power Management Based on Network Packet Processing Mode Transition for Latency-Critical Workloads
abstract
Processor power management exploiting Dynamic Voltage and Frequency Scaling (DVFS) plays a crucial role in improving the data-center’s energy efficiency. However, we observe that current power management policies in Linux (i.e., governors) often considerably increase tail response time (i.e., violate a given Service Level Objective (SLO)) and energy consumption of latency-critical applications. Furthermore, the previously proposed SLO-aware power management policies oversimplify network request processing and ignore the fact that network requests arrive at the application layer in bursts. Considering the complex interplay between the OS and network devices, we propose a power management framework exploiting network packet processing mode transitions in the OS to quickly react to the processing demands from the received network requests. Our proposed power management framework tracks the transitions between polling and interrupt in the network software stack to detect excessive packet processing on the cores and immediately react to the load changes by updating the voltage and frequency (V/F) states. Our experimental results show that our framework does not violate SLO and reduces energy consumption by up to 35.7% and 14.8% compared to Linux governors and state-of-the-art SLO-aware power management techniques, respectively.
Ki-Dong Kang, Gyeongseo Park, Hyosang Kim, Mohammad Alian, Nam Sung Kim, Daehoon Kim 0001
MICRO6
2021 GreenDIMM: OS-assisted DRAM Power Management for DRAM with a Sub-array Granularity Power-Down State
abstract
Power and energy consumed by DRAM comprising main memory of data-center servers have increased substantially as the capacity and bandwidth of memory increase. Especially, the fraction of DRAM background power in DRAM total power is already high, and it will continue to increase with the decelerating DRAM technology scaling as we will have to plug more DRAM modules in servers or stack more DRAM dies in a DRAM package to provide necessary DRAM capacity in the future. To reduce the background power, we may exploit low average utilization of the DRAM capacity in data-center servers (i.e., 40–60%) for DRAM power management. Nonetheless, the current DRAM power management supports low-power states only at the rank granularity, which becomes ineffective with memory interleaving techniques devised to disperse memory requests across ranks. That is, ranks need to be frequently woken up from low-power states with aggressive power management, which can significantly degrade system performance, or they do not get a chance to enter low-power states with conservative power management.
Seunghak Lee, Ki-Dong Kang, Hwanjun Lee, Hyungwon Park 0001, Young Hoon Son, Nam Sung Kim, Daehoon Kim 0001
MICRO7
2020 Improving the Efficiency of Power Management via Dynamic Interrupt Management
abstract
In this paper, we first analyze the effects of interrupt management on response latency of the latency-critical application and the efficiency of current dynamic power management governor which determines Voltage and Frequency States (V/F states) based on CPU utilization. We also demonstrate that interrupt management provides a governor an opportunity to decrease the V/F state without performance degradation. Next, we propose I-state that adjusts the interrupt rate based on the V/F state determined by the governor. When a core is highly utilized with a high V/F state, I-state improves response latency by decreasing interrupt rate, moderating the load on the CPU. I-state also improves energy-efficiency by making the governor decrease the V/F state more often. When a core is not highly utilized while operating at low V/F states, I-state improves response latency by increasing interrupt rate, which can notify the processor of packet arrivals faster so that the CPU processes the packets quickly. Our experimental results show that I-state improves 95thpercentile latency by up to 15.9x while reducing energy consumption by up to 17.6%.
Ki-Dong Kang, Hyungwon Park 0001, Gyeongseo Park, Daehoon Kim 0001
ICCD4
2018 VIP: Virtual Performance-State for Efficient Power Management of Virtual Machines
abstract
A power management policy aims to improve energy efficiency by choosing an appropriate performance (voltage/frequency) state for a given core. In current virtualized environments, multiple virtual machines (VMs) running on the same core must follow a single power management policy governed by the hypervisor. However, we observe that such a per-core power management policy has two limitations. First, it cannot offer the flexibility of choosing a desirable power management policy for each VM (or client). Second, it often hurts the power efficiency of some or even all VMs especially when the VMs desire conflicting power management policies. To tackle these limitations, we propose a per-VM power management mechanism, VIP supporting Virtual Performance-state for each VM. Specifically, for VMs sharing a core, VIP allows each VM's guest OS to deploy its own desired power management policy while preventing such VMs from interfering/influencing each other's power management policy. That is, VIP can also facilitate a pricing model based on the choice of a power management policy. Second, identifying some inefficiency in strictly enforcing per-VM power management policies, we propose hypervisor-assisted techniques to further improve power and energy efficiency without compromising the key benefits of per-VM power management. To demonstrate the efficacy of VIP, we take a case that some VMs run CPU-intensive applications and other VMs run latency-sensitive applications sharing the same cores. Our evaluation shows that VIP reduces the overall energy consumption and improves the execution time of CPU-intensive applications compared with the default ondemand governor of Xen hypervisor up to 27% and 32%, respectively, without violating service level agreement (SLA) of latency-sensitive applications.
Ki-Dong Kang, Mohammad Alian, Daehoon Kim 0001, Jaehyuk Huh 0001, Nam Sung Kim
SoCC3
2018 Application-Transparent Near-Memory Processing Architecture with Memory Channel Network
abstract
The physical memory capacity of servers is expected to increase drastically with deployment of the forthcoming non-volatile memory technologies. This is a welcomed improvement for emerging data-intensive applications. For such servers to be cost-effective, nonetheless, we must cost-effectively increase compute throughput and memory bandwidth commensurate with the increase in memory capacity without compromising application readiness. Tackling this challenge, we present Memory Channel Network (MCN) architecture in this paper. Specifically, first, we propose an MCN DIMM, an extension of a buffered DIMM where a small but capable processor called MCN processor is integrated with a buffer device on the DIMM for near-memory processing. Second, we implement device drivers to give the host and MCN processors in a server an illusion that they are independent heterogeneous nodes connected through an Ethernet link. These allow the host and MCN processors in a server to run a given data-intensive application together based on popular distributed computing frameworks such as MPI and Spark without any change in the host processor hardware and its application software, while offering the benefits of high-bandwidth and low-latency communications between the host and the MCN processors over memory channels. As such, MCN can serve as an application-transparent framework which can seamlessly unify near-memory processing within a server and distributed computing across such servers for data-intensive applications. Our simulation running the full software stack shows that a server with 8 MCN DIMMs offers 4.56X higher throughput and consume 47.5% less energy than a cluster with 9 conventional nodes connected through Ethernet links, as it facilitates up to 8.17X higher aggregate DRAM bandwidth utilization. Lastly, we demonstrate the feasibility of MCN with an IBM POWER8 system and an experimental buffered DIMM.
Mohammad Alian, Seungwon Min, Hadi Asghari Moghaddam, Ashutosh Dhar, Dong Kai Wang, Adam J. McPadden, Oliver O'Halloran, Deming Chen, Jinjun Xiong, Daehoon Kim 0001, Wen-Mei W. Hwu, Nam Sung Kim
MICRO11
2017 Janus: supporting heterogeneous power management in virtualized environments
abstract
The cloud servers have routinely adopted machine virtualization for high energy efficiency. Such virtualization notably improves energy efficiency not only through consolidation, but also through Dynamic Voltage/Frequency Scaling (DVFS). Thus, current hypervisors such as Xen and KVM support power management (PM) policies statically or dynamically setting a Voltage/Frequency (V/F) level, similar to ones deployed by the Linux. However, the current hypervisors can promote only a single PM policy (i.e., host governor) per physical core. This poses a unique challenge for VMs sharing a physical core and running applications with opposite runtime characteristics in a time-shared manner (i.e., heterogeneous VMs); note that the consolidation policy often encourages heterogeneous VMs to share a physical core, since such VMs use different resources in the system [2].
Daehoon Kim 0001, Mohammad Alian, Jaehyuk Huh 0001, Nam Sung Kim
SoCC1
2017 NCAP: Network-Driven, Packet Context-Aware Power Management for Client-Server Architecture
abstract
The rate of network packets encapsulating requests from clients can significantly affect the utilization, and thus performance and sleep states of processors in servers deploying a power management policy. To improve energy efficiency, servers may adopt an aggressive power management policy that frequently transitions a processor to a low-performance or sleep state at a low utilization. However, such servers may not respond to a sudden increase in the rate of requests from clients early enough due to a considerable performance penalty of transitioning a processor from a sleep or low-performance state to a high-performance state. This in turn entails violations of a service level agreement (SLA), discourages server operators from deploying an aggressive power management policy, and thus wastes energy during low-utilization periods. For both fast response time and high energy-efficiency, we propose NCAP, Network-driven, packet Context-Aware Power management for client-server architecture. NCAP enhances a network interface card (NIC) and its driver such that it can examine received and transmitted network packets, determine the rate of network packets containing latency-critical requests, and proactively transition a processor to an appropriate performance or sleep state. To demonstrate the efficacy, we evaluate on-line data-intensive (OLDI) applications and show that a server deploying NCAP consumes 37~61% lower processor energy than a baseline server while satisfying a given SLA at various load levels.
Mohammad Alian, Ahmed H. M. O. Abulila, Lokesh Jindal, Daehoon Kim 0001, Nam Sung Kim
HPCA4
2017 dist-gem5: Distributed simulation of computer clusters
abstract
When analyzing a distributed computer system, we often observe that the complex interplay among processor, node, and network sub-systems can profoundly affect the performance and power efficiency of the distributed computer system. Therefore, to effectively cross-optimize hardware and software components of a distributed computer system, we need a full-system simulation infrastructure that can precisely capture the complex interplay. Responding to the aforementioned need, we present dist-gem5, a flexible, detailed, and open-source full-system simulation infrastructure that can model and simulate a distributed computer system using multiple simulation hosts. Then we validate dist-gem5 against a physical cluster and show that the latency and bandwidth of the simulated network sub-system are within 18% of the physical one. Compared with the single threaded and parallel versions of gem5, dist-gem5 speeds up the simulation of a 63-node computer cluster by 83.1x and 12.8x, respectively.
Mohammad Alian, Umur Darbaz, Gábor Dózsa, Stephan Diestelhorst, Daehoon Kim 0001, Nam Sung Kim
ISPASS5
2016 Virtual Snooping Coherence for Multi-Core Virtualized Systems
abstract
Proliferation of virtualized systems opens a new opportunity to improve the scalability of multi-core architectures. Among the scalability bottlenecks in multi-cores, cache coherence has been one of the most critical problems. Although snoop-based protocols have been dominating commercial multi-core designs, it has been difficult to scale them for more cores, as snooping protocols require high network bandwidth and power consumption for snooping all the caches.In this paper, we propose a novel snoop-based cache coherence protocol, called virtual snooping, for virtualized multi-core architectures. Virtual snooping exploits memory isolation across virtual machines and prevents unnecessary snoop requests from crossing the virtual machine boundaries. Each virtual machine becomes a virtual snoop domain, consisting of a subset of the cores in a system. Although the majority of virtual machine memory is isolated, sharing of cachelines across VMs still occur. To address such data sharing, this paper investigates three factors, data sharing through the hypervisor, virtual machine relocation, and content-based sharing. In this paper, we explore the design space of virtual snooping with experiments on emulated and real virtualized systems including the mechanisms and overheads of the hypervisor. In addition, the paper discusses the scheduling impact on the effectiveness of virtual snooping.
Daehoon Kim 0001, Chang Hyun Park 0001, Hwanju Kim, Jaehyuk Huh 0001
IEEE Trans. Parallel Distributed Syst.1
2015 vCache: architectural support for transparent and isolated virtual LLCs in virtualized environments
abstract
A key role of virtualization is to give an illusion that a consolidated workload runs on a dedicated machine although the underlying resources are actively shared by multiple workloads. Technical advances have enabled a virtual machine (VM) to exercise many shared resources of a machine in a transparent and isolated manner. However, such an illusion of resource dedication has not been supported for the last-level cache (LLC), although the LLC is the largest on-chip shared resource with a significant performance impact. In this paper, we propose vCache--architectural support to provide a transparent and isolated virtual LLC (vLLC) for each VM and interfaces to manage the vLLC. More specifically, this study first proposes architectural support for the guest OS of a VM to index the LLC with its guest physical address instead of a host physical address. This in turn allows that the guest OS transparently view its vLLC and preserve the effectiveness of its page placement policy. Second, this study extends the architectural support for each VM to keep its vLLC strongly isolated from other VMs. Such resource dedication is critical to offer performance isolation and preserve vLLC transparency for each VM in a highly consolidated machine. With little hardware overhead, vCache can facilitate various unchartered vLLC capacity-based services for the public clouds while providing up to 17% higher performance than a traditional virtualized system.
Daehoon Kim 0001, Hwanju Kim, Nam Sung Kim, Jaehyuk Huh 0001
MICRO1
2012 Subspace Snooping: Exploiting Temporal Sharing Stability for Snoop Reduction
abstract
Although snoop-based coherence protocols provide fast cache-to-cache transfers with a simple and robust coherence mechanism, scaling the protocols has been difficult due to the overheads of broadcast snooping. In this paper, we propose a coherence filtering technique called subspace snooping, which stores the potential sharers of each memory page in the page table entry. By using the sharer information in the page table entry, coherence transactions for a page generate snoop requests only to the subset of nodes in the system. However, the coherence subspace of a page may evolve, as the phases of applications may change or the operating system may migrate threads to different nodes. To adjust subspaces dynamically, subspace snooping supports two different shrinking mechanisms, which remove obsolete nodes from subspaces. Among the two shrinking mechanisms, subspace snooping with safe shrinking can be integrated to any type of coherence protocols and network topologies, as it guarantees that a subspace always contains the precise sharers of a page. Speculative shrinking breaks the subspace superset property, but achieves better snoop reductions than safe shrinking. We evaluate subspace snooping with Token Coherence on unordered mesh networks. Subspace snooping reduces 58 percent of snoops on average for a set of parallel scientific and server workloads, and 87 percent for our multiprogrammed workloads.
Jeongseob Ahn, Daehoon Kim 0001, Jaehong Kim 0006, Jaehyuk Huh 0001
IEEE Trans. Computers2
2010 Subspace snooping: filtering snoops with operating system support
abstract
Although snoop-based coherence protocols provide fast cache-to-cache transfers with a simple and robust coherence mechanism, scaling the protocols has been difficult due to the overheads of broadcast snooping. In this paper, we propose a coherence filtering technique called subspace snooping, which stores the potential sharers of each memory page in the page table entry. By using the sharer information in the page table entry, coherence transactions for a page generate snoop requests only to the subset of nodes in the system (subspace). However, the coherence subspace of a page may evolve, as the phases of applications may change or the operating system may migrate threads to different nodes. To adjust subspaces dynamically, subspace snooping supports a shrinking mechanism, which removes obsolete nodes from subspaces.
Daehoon Kim 0001, Jeongseob Ahn, Jaehong Kim 0006, Jaehyuk Huh 0001
PACT1
2010 Virtual Snooping: Filtering Snoops in Virtualized Multi-cores
abstract
Virtualization has been rapidly expanding its applications in numerous server and desktop environments to improve the utilization and manageability of physical systems. Such proliferation of virtualized systems opens a new opportunity to improve the scalability of future multi-core architectures. Among the scalability bottlenecks in multi-cores, cache coherence has been a critical problem. Although snoop-based protocols have been dominating commercial multi-core designs, it has been difficult to scale them for more cores, as snooping protocols require high network bandwidth and power consumption for snooping all the caches. In this paper, we propose a novel snoop-based cache coherence protocol, called virtual snooping, for virtualized multi-core architectures. Virtual snooping exploits memory isolation across virtual machines and prevents unnecessary snoop requests from crossing the virtual machine boundaries. Each virtual machine becomes a virtual snoop domain, consisting of a subset of the cores in a system. However, in real virtualized systems, virtual machines cannot partition the cores perfectly without any data sharing across the snoop partitions. This paper investigates three factors, which break the memory isolation among virtual machines: data sharing with a hyper visor, virtual machine relocation, and content-based data sharing. In this paper, we explore the design space of virtual snooping with experiments on real virtualized systems and approximate simulations. The results show that virtual snooping can reduce snoops significantly even if virtual machines migrate frequently. We also propose mechanisms to address content-based data sharing by exploiting its read-only property.
Daehoon Kim 0001, Hwanju Kim, Jaehyuk Huh 0001
MICRO1