Jeeho Ryoo

dblp:121/2281 · also Jee Ho Ryoo · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
6since 2021 · last 2026
0009-0003-0401-3685ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 6 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Memory systems · 45% Processor architecture and microarchitecture · 10% Parallel and multicore computing · 8%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 22 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory management › virtual memory
address translation
0.622017
CSALT: context switch aware large TLB · MICRO 2017
Rethinking TLB Designs in Virtualized Environments: A Very Large Part-of-Memory TLB · ISCA 2017
Storage systems
data migration
0.312017
SILC-FM: Subblocked InterLeaved Cache-Like Flat Memory Organization · HPCA 2017
Memory systems
hybrid memory
0.312017
SILC-FM: Subblocked InterLeaved Cache-Like Flat Memory Organization · HPCA 2017
Memory systems › memory management › virtual memory
page table walk
0.312017
CSALT: context switch aware large TLB · MICRO 2017
Memory systems › memory management › virtual memory › address translation
TLB
0.312017
CSALT: context switch aware large TLB · MICRO 2017
Memory systems › memory management › virtual memory › address translation › TLB
TLB design
0.312017
Rethinking TLB Designs in Virtualized Environments: A Very Large Part-of-Memory TLB · ISCA 2017
Memory systems › memory management
virtual memory
0.312017
CSALT: context switch aware large TLB · MICRO 2017
Memory systems
cache
0.212016
Dynamic Core Allocation and Packet Scheduling in Multicore Network Processors · IEEE Trans. Computers 2016
Processor architecture and microarchitecture
chip multiprocessor
0.212016
Dynamic Core Allocation and Packet Scheduling in Multicore Network Processors · IEEE Trans. Computers 2016
Processor architecture and microarchitecture › special-purpose processor
network processor
0.212016
Dynamic Core Allocation and Packet Scheduling in Multicore Network Processors · IEEE Trans. Computers 2016
Interconnection networks and networks-on-chip › network scheduling
packet scheduling
0.212016
Dynamic Core Allocation and Packet Scheduling in Multicore Network Processors · IEEE Trans. Computers 2016
Electronic design automation › high-level synthesis
scheduling
0.212016
Dynamic Core Allocation and Packet Scheduling in Multicore Network Processors · IEEE Trans. Computers 2016
Parallel and multicore computing
data distribution
0.212015
Data partitioning strategies for graph workloads on heterogeneous clusters · SC 2015
Parallel and multicore computing
graph processing
0.212015
Data partitioning strategies for graph workloads on heterogeneous clusters · SC 2015
High-performance computing › cluster computing
heterogeneous clusters
0.212015
Data partitioning strategies for graph workloads on heterogeneous clusters · SC 2015
Cloud and datacenter computing
virtualization
0.222017
CSALT: context switch aware large TLB · MICRO 2017
Rethinking TLB Designs in Virtualized Environments: A Very Large Part-of-Memory TLB · ISCA 2017
Distributed systems › fault tolerance
checkpointing
0.112012
Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems · SC 2012
Interconnection networks and networks-on-chip › error control
error detection and recovery
0.112012
Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems · SC 2012
High-performance computing › supercomputing
exascale computing
0.112012
Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems · SC 2012
Distributed systems
fault tolerance
0.112012
Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems · SC 2012
Distributed systems › fault tolerance
resilience
0.112012
Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems · SC 2012
Memory systems › virtual memory management
address translation overhead
0.112017
Rethinking TLB Designs in Virtualized Environments: A Very Large Part-of-Memory TLB · ISCA 2017

Methods — techniques the papers use, named apart from their topics

heterogeneity-aware data ingress · 0.4graph cutting · 0.4radix-4 page tables · 0.3performance simulation · 0.3page walk cache · 0.3energy analysis · 0.3context switch aware TLB design · 0.3hardware implementation · 0.2trace-driven simulation · 0.1analytical modeling · 0.1
YearPublicationVenuePosition
2026 POSTER: Hardware Acceleration for Graph Neural Networks
Cory Davis, Patrick M. Stockton, Eugene John, Jeeho Ryoo, Ebod Shojaei
CF4
2026 Performance Analysis and Optimization of 3D Generative Diffusion Models across GPU Architectures
abstract
Diffusion models have become essential for high-fidelity 3D MRI synthesis, yet their deployment remains constrained by substantial GPU resource demands arising from hundreds of U-Net evaluations per sample and a highly heterogeneous kernel behavior. This paper performs a comprehensive performance analysis of the state-of-the-art medical diffusion model, Med-DDPM, across three generations of NVIDIA architectures to study kernel-level runtime breakdowns, instruction-mix characteristics, memory system utilization, warp-level activities, and profiler priority-score estimates. We show that training is overwhelmingly dominated by cuDNN convolution and implicit-GEMM kernels, with inefficiencies arising from memory-access patterns, tensor-layout conversions, and limited Tensor Core utilization. Guided by these insights, we evaluate two architecture-aware optimizations TF32 Tensor Core activation and a 3D channels-last layout and demonstrate that they reduce SM cycles by up to 100x, cut dynamic instructions by 100x, raise Tensor Core utilization from 1.45 to 9.98x, and increase IPC by 7% on A100, all without degrading synthesis quality.
Jeeho Ryoo, Yongchan Jung, Muhammad Ali Khaliq, Jiatong Han, Byeong Kil Lee
ICPE1
2025 Oneiros: KV Cache Optimization through Parameter Remapping for Multi-tenant LLM Serving
abstract
KV cache accelerates LLM inference by avoiding redundant computation, at the expense of memory. To support larger KV caches, prior work extends GPU memory with CPU memory via CPU-offloading. This involves swapping KV cache between GPU and CPU memory. However, because the cache updates dynamically, such swapping incurs high CPU memory traffic. We make a key observation that model parameters remain constant during runtime, unlike the dynamically updated KV cache. Building on this, we introduce Oneiros, which avoids KV cache swapping by remapping, and thereby repurposing, the memory allocated to model parameters for KV cache. This parameter remapping is especially beneficial in multi-tenant environments, where the memory used for the parameters of the inactive models can be more aggressively reclaimed. Exploiting the high CPU-GPU bandwidth offered by the modern hardware, such as the NVIDIA Grace Hopper Superchip, we show that Oneiros significantly outperforms state-of-the-art solutions, achieving a reduction of 44.8%-82.5% in tail time-between-token latency, 20.7%-99.3% in tail time-to-first-token latency, and 6.6%-86.7% higher throughput compared to vLLM. Source code of Oneiros is available at https://github.com/UT-SysML/Oneiros/.
Ruihao Li 0002, Shagnik Pal, Vineeth Narayan Pullu, Prasoon Sinha, Jeeho Ryoo, Lizy Kurian John, Neeraja J. Yadwadkar
SoCC5
2025 Evaluating the Impact of Assistive AI Tools on Learning Outcomes and Ethical Considerations in Programming Education
abstract
This study critically evaluates the efficacy of GitHub Copilot in low-level programming education, specifically within C programming tasks involving complex concepts like memory management and pointer manipulation. While AI tools have shown promise in supporting high-level programming, its impact on skill-intensive, low-level contexts remains underexplored. We conducted a within-subject experimental study with 34 graduate computer science students, assessing performance on AI -assisted and independent tasks. Statistical analyses revealed that Copilot, one of the AI programming tools, enhances productivity in routine coding activities; however, it is insufficient for tasks requiring deep problem-solving skills. Notably, a significant performance decline in AI-free tasks suggests a dependency on Copilot that may hinder the development of essential independent problem-solving abilities. Survey feedback underscores ethical concerns, with 40.6 % of students expressing uncertainty about responsible AI usage and potential over-reliance. These findings highlight the ne-cessity for structured instructional practices, including AI-free assessments and clear ethical guidelines, to promote balanced technology integration in programming education. This study contributes to educational theory by illuminating the limitations of generative AI within constructivist and self-regulated learning frameworks. Future research should explore the long-term effects of AI dependency on technical skill development and investigate AI advancements tailored for low-level programming to better support foundational skills.
Seong Min Park, Marco Ho, Michael Pin-Chuan Lin, Jeeho Ryoo
EDUCON4
2025 Mapping AI Tools in Education: A Topic Modeling Analysis of Cognitive, Metacognitive, and Affective Insights
Michael Pin-Chuan Lin, Arita Li Liu, Saeed Saffari, Daniel Chang, Jeeho Ryoo
ITS (1)5
2023 Evaluation of Pruning Techniques
abstract
CNNs are widely used in a variety of computer vision tasks such as image processing, image classification, etc. The state-of-the-art neural networks are bigger and have greater number of parameters which translates to greater computations and memory footprint. This makes it difficult to deploy the real-time applications using these models on resource constrained edge devices. To address this issue, pruning techniques have been explored which reduce the number of computations and memory requirements of modern CNNs with negligible loss in accuracy. In this paper, we analyze the performance of three different pruning techniques - L1-norm based filter pruning, channel pruning and weight pruning and compare the performance in terms of inference time and accuracy. We also compared the performance of the pruned networks on two different GPU architectures and found that the inference time of pruned networks is more boosted in NVIDIA V100 compared to NVIDIA GTX1080 Ti due to its superior architecture features. We also perform inference speedup comparison between the different pruning techniques and analyze performance benefits between the different pruning techniques.
Shvetha S. Kumar, Reshma R. Nayak, Jismi S. Kannampuzha, Jeeho Ryoo, Sahil Rai, Lizy Kurian John
IPCCC4
2018 Puzzle Memory: Multifractional Partitioned Heterogeneous Memory Scheme
abstract
As current main memory technology scaling is coming close to an end due to its physical limitations, many emerging memory technologies are coming to the market to fill the scaling gap. Future memory systems will require a heterogeneous memory architecture where one technology acts as a low latency memory whereas the other acts as a high capacity memory. This will allow the future main memory system to continue to scale in terms of capacity, yet have similar or slightly better latency than today's DRAM technology. Prior work on data management in heterogeneous memory has optimized one or a maximum of two components in the computing stack. However, different components are good at different tasks in data management, so in the era of heterogeneous memory, it is inevitable that cooperative multi-component data management will be adopted in future systems. We propose a heterogeneous memory layout where two memories are laid out asymmetrically. The operating system is aware of this layout and places pages with different locality characteristics in different regions of memory. Finally, a custom hardware performs the data remapping to optimize the data placement at finer granularity than what is visible to the operating system. In the end, we show that our multi-component cooperative data management scheme can improve the overall system performance by up to 40%.
Jeeho Ryoo, Shuang Song 0007, Lizy Kurian John
ICCD1
2018 A Case for Granularity Aware Page Migration
abstract
Memory is becoming increasingly heterogeneous with the emergence of disparate memory technologies ranging from non-volatile memories like PCM, STT-RAM, and memristors to 3D-stacked memories like HBM. In such systems, data is of ten migrated across memory regions backed by different technologies for better overall performance. An effective migration mechanism is a prerequisite in such systems.
Jeeho Ryoo, Lizy Kurian John, Arkaprava Basu
ICS1
2017 SILC-FM: Subblocked InterLeaved Cache-Like Flat Memory Organization
abstract
With current DRAM technology reaching its limit, emerging heterogeneous memory systems have become attractive to continue scaling memory performance. This paper argues for using a small, fast memory closer to the processor as part of a flat address space where the memory system is composed of two or more memory types. OS-transparent management of such memory has been proposed in prior works such as CAMEO and Part of Memory (PoM). Data migration is typically handled either at coarse granularity with high bandwidth overheads (as in PoM) or at fine granularity with low hit rate (as in CAMEO). Prior work uses restricted address mapping from only congruence groups in order to simplify the mapping. At any time, only one page (block) from a congruence group is resident in the fast memory. In this paper, we present a flat address space organization called SILC-FM that uses large granularity but allows subblocks from two pages to coexist in an interleaved fashion in fast memory. Data movement is done at subblocked granularity, avoiding fetching of useless subblocks and consuming less bandwidth compared to migrating the entire large block. SILC-FM can achieve more spatial locality hits than CAMEO and PoM due to page-level operation and interleaving blocks respectively. The interleaved subblock placement improves performance by 55% on average over a static placement scheme without data migration. We also selectively lock hot blocks to prevent them from being involved in hardware swapping operations. Additional features such as locking, associativity and bandwidth balancing improve performance by 11%, 8%, and 8% respectively, resulting in a total of 82% performance improvement over a no migration static placement scheme. Compared to the best state-of-the-art scheme, SILC-FM gets performance improvement of 36% with 13% energy savings.
Jeeho Ryoo, Mitesh R. Meswani, Andreas Prodromou, Lizy Kurian John
HPCA1
2017 Rethinking TLB Designs in Virtualized Environments: A Very Large Part-of-Memory TLB
abstract
With increasing deployment of virtual machines for cloud services and server applications, memory address translation overheads in virtualized environments have received great attention. In the radix-4 type of page tables used in x86 architectures, a TLB-miss necessitates up to 24 memory references for one guest to host translation. While dedicated page walk caches and such recent enhancements eliminate many of these memory references, our measurements on the Intel Skylake processors indicate that many programs in virtualized mode of execution still spend hundreds of cycles for translations that do not hit in the TLBs.
Jeeho Ryoo, Nagendra Dwarakanath Gulur, Shuang Song 0007, Lizy Kurian John
ISCA1
2017 CSALT: context switch aware large TLB
abstract
Computing in virtualized environments has become a common practice for many businesses. Typically, hosting companies aim for lower operational costs by targeting high utilization of host machines maintaining just enough machines to meet the demand. In this scenario, frequent virtual machine context switches are common, resulting in increased TLB miss rates (often, by over 5X when contexts are doubled) and subsequent expensive page walks. Since each TLB miss in a virtual environment initiates a 2D page walk, the data caches get filled with a large fraction of page table entries (often, in excess of 50%) thereby evicting potentially more useful data contents.
Yashwant Marathe, Nagendra Dwarakanath Gulur, Jeeho Ryoo, Shuang Song 0007, Lizy Kurian John
MICRO3
2016 POSTER: SILC-FM: Subblocked InterLeaved Cache-Like Flat Memory Organization
abstract
In this paper, we present a flat address space organization called SILC-FM that allows subblocks from two pages to co-exist in an interleaved fashion in die-stacked DRAM. Data movement at subblocked granularity consumes less bandwidth compared to migrating the entire large block and prevents fetching useless subblocks that may never get accessed. SILC-FM can get more spatial locality hits than CAMEO and PoM due to page-level operation and interleaving blocks respectively. The interleaved subblock placement improves performance by 55% on average over a static placement scheme without data migration. We also selectively lock hot blocks to prevent them from being involved in the hardware swapping operations. Additional features such as locking, associativity and bandwidth balancing improve performance by 11%, 8%, and 8% respectively, resulting in a total of 82% performance improvement over no migration static placement scheme. Compared to the best state-of-the-art scheme, SILC-FM gets performance improvement of 36%.
Jeeho Ryoo, Mitesh R. Meswani, Reena Panda, Lizy Kurian John
PACT1
2016 Proxy-Guided Load Balancing of Graph Processing Workloads on Heterogeneous Clusters
abstract
Big data decision-making techniques take advantage of large-scale data to extract important insights from them. One of the most important classes of such techniques falls in the domain of graph applications, where data segments and their inherent relationships are represented as vertices and edges. Efficiently processing large-scale graphs involves many subtle tradeoffs and is still regarded as an open-ended problem. Furthermore, as modern data centers move towards increased heterogeneity, the traditional assumption of homogeneous environments in current graph processing frameworks is no longer valid. Prior work estimates the graph processing power of heterogeneous machines by simply reading hardware configurations, which leads to suboptimal load balancing. In this paper, we propose a profiling methodology leveraging synthetic graphs for capturing a node's computational capability and guiding graph partitioning in heterogeneous environments with minimal overheads. We show that by sampling the execution of applications on synthetic graphs following a power-law distribution, the computing capabilities of heterogeneous clusters can be captured accurately (<;10% error). Our proxy-guided graph processing system results in a maximum speedup of 1.84x and 1.45x over a default system and prior work, respectively. On average, it achieves 17.9% performance improvement and 14.6% energy reduction as compared to prior heterogeneity-aware work.
Shuang Song 0007, Xinnian Zheng, Michael LeBeane, Jeeho Ryoo, Reena Panda, Andreas Gerstlauer, Lizy Kurian John
ICPP5
2016 Dynamic Core Allocation and Packet Scheduling in Multicore Network Processors
abstract
With ever increasing network traffic rates, multicore architectures for network processors have successfully provided performance improvements through high parallelism. However, naively allocating the network traffic to multiple cores without considering diversified applications and flow locality results in issues such as packet reordering, load imbalance and inefficient cache usage. Consequently, these issues degrade the performance of latency sensitive network processors by dropping packets or delivering packets out of order. In this paper, we propose a packet scheduling scheme that considers the multiple dimensions of locality to improve the throughput of a network processor while minimizing out of order packets. Our scheduling policy tries to maintain packet order my maintaining the flow locality, minimizes the migration of flows from one core to another by identifying the aggressive flows, and partitions the cores among multiple services to gain instruction cache locality. Our light weight hardware implementation shows improvement of 60 percent in the number of packets dropped and 80 percent in the number of out-of-order packet deliveries over previously proposed techniques.
Muhammad Faisal Iqbal, Jim Holt, Jeeho Ryoo, Gustavo de Veciana, Lizy Kurian John
IEEE Trans. Computers3
2015 GPGPU Benchmark Suites: How Well Do They Sample the Performance Spectrum?
abstract
Recently, GPGPUs have positioned themselves in the mainstream processor arena with their potential to perform a massive number of jobs in parallel. At the same time, many GPGPU benchmark suites have been proposed to evaluate the performance of GPGPUs. Both academia and industry have been introducing new sets of benchmarks each year while some already published benchmarks have been updated periodically. However, some benchmark suites contain benchmarks that are duplicates of each other or use the same underlying algorithm. This results in an excess of workloads in the same performance spectrum. In this paper, we provide a methodology to obtain a set of new GPGPU benchmarks that are located in the unexplored region of the performance spectrum. Our proposal uses statistical methods to understand the performance spectrum coverage and uniqueness of existing benchmark suites. Later we show techniques to identify areas that are not explored by existing benchmarks by visually showing the performance spectrum coverage. Finding unique key metrics for future benchmarks to broaden its performance spectrum coverage is also explored using hierarchical clustering and ranking by Hotel ling's T2 method. Finally, key metrics are categorized into GPGPU performance related components to show how future benchmarks can stress each of the categorized metrics to distinguish themselves in the performance spectrum. Our methodology can serve as a performance spectrum oriented guidebook for designing future GPGPU benchmarks.
Jeeho Ryoo, Saddam Quirem, Michael LeBeane, Reena Panda, Shuang Song 0007, Lizy Kurian John
ICPP1
2015 PowerTrain: A learning-based calibration of McPAT power models
abstract
As research on improving energy efficiency becomes prevalent, the necessity of a tool to accurately estimate power is increasing. Among various tools proposed, McPAT has gained some popularity due to its easy-to-use analytical power models. However, McPAT's prediction has several limitations. Although under- or over-estimated power from unmodeled and mis-modeled parts offset each other, it still incorporates errors in each block. Moreover, the lack of awareness to the implementation details exacerbates the prediction inaccuracies. To alleviate this problem, we propose a new methodology to train McPAT towards precise processor power prediction using power measurements from real hardware. This calibration enables McPAT's power to fit to the target processor power. Once we adjusted the power consumption of each block to best match those in the target processor, our trained McPAT delivered more precise power estimation. We calibrated the outputs of McPAT against a Cortex-A15 within a Samsung Exynos 5422 SoC. We observe that our methodology successfully reduces the errors, particularly for workloads with fluctuating power behaviors. The results show that the mean percentage error and the mean percentage absolute error of the calibrated power against real hardware are 2.04 percent and 4.37 percent, respectively.
Wooseok Lee, Youngchun Kim, Jeeho Ryoo, Dam Sunwoo, Andreas Gerstlauer, Lizy Kurian John
ISLPED3
2015 Watt Watcher: Fine-Grained Power Estimation for Emerging Workloads
abstract
Extensive research has focused on estimating power to guide advances in power management schemes, thermal hot spots, and voltage noise. However, simulated power models are slow and struggle with deep software stacks, while direct measurements are typically coarse-grained. This paper introduces Watt Watcher, a multicore power measurement framework that offers fine-grained functional unit breakdowns. Watt Watcher operates by passing event counts and a hardware descriptor file into configurable back-end power models based on McPAT. Researchers and vendors can add other processors to our tool by mapping to the Watt Watcher interface. We show that Watt Watcher, when calibrated, has a MAPE (mean absolute percentage error) of 2.67% aggregated over all benchmarks when compared to measured power consumption on SPEC CPU 2006 and multithreaded PARSEC benchmarks across three different machines of various form factors and manufacturing processes. We present two use cases showing how Watt Watcher can derive insights that are difficult to obtain through other measurement infrastructures. Additionally, we illustrate how Watt Watcher can be used to provide insights into challenging big data and cloud workloads on a server CPU. Through the use of Watt Watcher, it is possible to obtain a detailed power breakdown on real hardware without vendor proprietary models or hardware instrumentation.
Michael LeBeane, Jeeho Ryoo, Reena Panda, Lizy Kurian John
SBAC-PAD2
2015 Performance Characterization of Modern Databases on Out-of-Order CPUs
abstract
Big data revolution has created an unprecedented demand for intelligent data management solutions on a large scale. While data management has traditionally been used as a synonym for relational data processing, in recent years a new group popularly known as NoSQL databases have emerged as a competitive alternative. There is a pressing need to gain greater understanding of the characteristics of modern databases to architect targeted computers. In this paper, we investigate four popular NoSQL/SQL-style databases and evaluate their hardware performance on modern computer systems. Based on data collected from real hardware, we evaluate how efficiently modern databases utilize the underlying systems and make several recommendations to improve their performance efficiency. We observe that performance of modern databases is severely limited by poor cache/memory performance. Nonetheless, we demonstrate that dynamic execution techniques are still effective in hiding a significant fraction of the stalls, thereby improving performance. We further show that NoSQL databases suffer from greater performance inefficiencies than their SQL counterparts. SQL databases outperform NoSQL databases for most operations and are beaten by NoSQL databases only in a few cases. NoSQL databases provide a promising competitive alternative to SQL-style databases, however, they are yet to be optimized to fully reach the performance of contemporary SQL systems. We also show that significant diversity exists among different database implementations and big-data benchmark designers can leverage our analysis to incorporate representative workloads to encapsulate the full spectrum of data-serving applications. In this paper, we also compare data-serving applications with other popular benchmarks such as SPEC CPU2006 and SPECjbb2005.
Reena Panda, Christopher Erb, Michael LeBeane, Jeeho Ryoo, Lizy Kurian John
SBAC-PAD4
2015 i-MIRROR: A Software Managed Die-Stacked DRAM-Based Memory Subsystem
abstract
This paper presents an operating system managed die-stacked DRAM called i-MIRROR that mirrors high locality pages from off-chip DRAM. Optimizing the problems of reducing cache tag area, reducing transfer bandwidth and improving hit latency altogether while using die-stacked DRAM as hardware cache is extremely challenging. In this paper, we show that performance and energy efficiency can be obtained by software management of die-stacked DRAM, which eliminates the need for tags, the source of aforementioned problems. In the proposed scheme, the operating system loads pages from disks to die-stacked DRAM on a page fault at the same time as they are loaded to off-chip DRAM. Our scheme maintains the pages in off-chip and die-stacked DRAM in a synchronized/mirrored state by exploiting the parallel loading capability to die-stacked and off-chip DRAM from the disk. This eliminates the need for physical page movement to the slower off-chip DRAM upon eviction from die-stacked DRAM. Requests for pages that got evicted from die-stacked DRAM are simply serviced by the slower off-chip DRAM to prevent frequent data movements of large pages and thrashing between conflicting pages. The operating system periodically monitors the usage of the pages in off-chip DRAM and promotes high locality pages to die-stacked DRAM. Our evaluations show that the proposed hardware-assisted software-managed i-MIRROR scheme achieves an IPC improvement of 13% while consuming 6% less energy than prior state-of-the-art die-stacked caching schemes and 79% improvement in terms of IPC and 72% in terms of energy savings over systems without die-stacked DRAM support.
Jeeho Ryoo, Karthik Ganesan 0006, Yao-Min Chen, Lizy Kurian John
SBAC-PAD1
2015 Data partitioning strategies for graph workloads on heterogeneous clusters
abstract
Large scale graph analytics are an important class of problem in the modern data center. However, while data centers are trending towards a large number of heterogeneous processing nodes, graph analytics frameworks still operate under the assumption of uniform compute resources. In this paper, we develop heterogeneity-aware data ingress strategies for graph analytics workloads using the popular PowerGraph framework. We illustrate how simple estimates of relative node computational throughput can guide heterogeneity-aware data partitioning algorithms to provide balanced graph cutting decisions. Our work enhances five online data ingress strategies from a variety of sources to optimize application execution for throughput differences in heterogeneous data centers. The proposed partitioning algorithms improve the runtime of several popular machine learning and data mining applications by as much as a 65% and on average by 32% as compared to the default, balanced partitioning approaches.
Michael LeBeane, Shuang Song 0007, Reena Panda, Jeeho Ryoo, Lizy Kurian John
SC4
2013 Flow Migration on Multicore Network Processors: Load Balancing While Minimizing Packet Reordering
abstract
With ever increasing network traffic rates, multicore architectures for network processors have successfully provided performance improvements through high parallelism. However, naively allocating the network traffic to multiple cores without considering diversified applications and flow locality results in issues such as packet reordering, load imbalance and inefficient cache usage. Consequently, these issues degrade the performance of latency sensitive network processors by dropping packets or delivering packets out of order. In this paper, we propose a packet scheduling scheme that considers the multiple dimensions of locality to improve the throughput of a network processor while minimizing out of order packets. Our scheduling policy tries to maintain packet order by maintaining the flow locality, minimizes the migration of flows from one core to another by identifying the aggressive flows, and partitions the cores among multiple services to gain instruction cache locality. The scheduler uses a novel low cost two-level caching scheme to identify top aggressive flows. Our light weight hardware implementation shows improvement of 60% in the number of packets dropped and 80% improvement in the out-of-order packet deliveries over previously proposed techniques.
Muhammad Faisal Iqbal, Jim Holt, Jeeho Ryoo, Lizy Kurian John, Gustavo de Veciance
ICPP3
2012 Containment domains: a scalable, efficient, and flexible resilience scheme for exascale systems
abstract
This paper describes and evaluates a scalable and efficient resilience scheme based on the concept of containment domains. Containment domains are a programming construct that enable applications to express resilience needs and to interact with the system to tune and specialize error detection, state preservation and restoration, and recovery schemes. Containment domains have weak transactional semantics and are nested to take advantage of the machine and application hierarchies and to enable hierarchical state preservation, restoration, and recovery. We evaluate the scalability and efficiency of containment domains using generalized trace-driven simulation and analytical analysis and show that containment domains are superior to both checkpoint restart and redundant execution approaches.
Jinsuk Chung, Ikhwan Lee, Michael B. Sullivan 0001, Jeeho Ryoo, Dong-Wan Kim, Doe Hyun Yoon, Larry Kaplan, Mattan Erez
SC4