EDBT 2026 Demo / reviewers in the wild / expert
Ipoom Jeong
dblp:221/0606
· DBLP profile ↗
19ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0001-7513-2858ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 4 first-author · 15 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Compiler and System Optimizations for Gem5 SimulatorabstractArchitectural simulators are indispensable for modern computer architecture research, but they remain notoriously slow due to their event-driven, cycle-level execution model. In this work, we present a set of software- and system-level optimizations to accelerate large-scale design-space exploration with gem5. First, we reduce per-instance simulation time via compiler-level optimization. We demonstrate that although gem5 suffers severe frontend stalls on modern CPUs stemming from its large instruction footprints, naïve Profile-Guided Optimization (PGO) is impractical in this setting because it requires frequent reprofiling and recompilation. To address this, we challenge the conventional reliance on self-profiling and instead construct a universal, performance-driven profile that generalizes across simulation inputs. Second, we improve aggregate simulation throughput by strengthening performance isolation using Sub-NUMA clustering (SNC). Finally, we show that a simple co-scheduling heuristic has great potential for reducing resource stranding and boosting multi-instance efficiency. Together, these techniques improve single simulation speed by 17 % and aggregate throughput by $27 \%$, making large-scale design-space exploration more practical and efficient. Haneul Park, Siddharth Agarwal, Pradyun Narkadamilli, Kiung Jung, Yongjun Park 0001, Ipoom Jeong, Nam Sung Kim |
ISPASS | 6 |
| 2025 | Warped-Compaction: Maximizing GPU Register File Bandwidth Utilization via Operand Compaction
Eunbi Jeong, Ipoom Jeong, Myung Kuk Yoon, Nam Sung Kim |
HPCA | 2 |
| 2025 | Marching Page Walks: Batching and Concurrent Page Table Walks for Enhancing GPU ThroughputabstractVirtual memory, with the support of address translation hardware, is a key technique in expanding programmability and memory management in GPUs. However, the nature of the GPU execution model heavily pressures its translation hardware, particularly due to a discrepancy in the behavior of page table walkers and thousands of concurrently running threads. In GPU workloads, multiple threads simultaneously access a number of pages necessitating a substantial number of translations whereas each walker handles only a single walk request at a time. Such a limitation significantly increases the queueing latency of walk requests, which we observe as a major bottleneck for servicing page table walks in GPUs. To tackle this challenge, we investigate a design of page walkers that facilitates multiple walk requests to be handled together in batches. Then, we make the following observations: 1) allowing a page walker to issue beyond a single memory request significantly improves the throughput of walkers, and 2) GPU applications tend to concurrently access pages in wide address ranges. By leveraging the above implications, we propose Marching Page Walks (MPW) that effectively mitigate the contention in GPU page table walkers. MPW scans pending walk requests to identify ones that can be grouped together. Then, MPW batches these requests and concurrently handles them by issuing multiple memory instructions. Experiments show that MPW reduces the queueing latency of page walks by 86.7% and improves GPU performance by 55.6% over the baseline design. Jiwon Lee 0001, Gun Ko, Myung Kuk Yoon, Ipoom Jeong, Yunho Oh, Won Woo Ro |
HPCA | 4 |
| 2025 | UPP: Universal Predicate Pushdown to Smart StorageabstractIn large-scale analytics, in-storage processing (ISP) can significantly boost query performance by letting ISP engines (e.g., FPGAs) preselect only the relevant data before sending them to databases.This reduces the amount of not only data transfer between storage and host, but also database computation, facilitating faster query processing.However, existing ISP solutions cannot effectively support a wide range of modern analytical queries because they only support simple combinations of frequently used operators (e.g., =, <), particularly on fixed-length columns.As modern databases allow filter predicates to include numerous operators/functions (e.g., dateadd) compatible with diverse data formats (and their complex combinations), it becomes more challenging for existing approaches to accelerate such queries efficiently.To address the limitations, we propose a new ISP approach, called Universal Predicate Pushdown (UPP), that can accelerate modern analytical databases, leveraging hardware/software co-design for a high level of flexibility.Our core insight is that instead of programming for individual filter operators/functions, we should devise a compact instruction set architecture (ISA) tailored explicitly for predicate pushdown.The software (i.e., database) layer recognizes and compiles various general filters (called a universal predicate) to a set of UPP-compliant instructions, which are then processed efficiently by FPGA using bitwise comparisons, leveraging lightweight metadata.In our experiments with a 100 GB TPC-H dataset, UPP running on SmartSSD could speed up Spark's end-to-end query performance by 1.2×-7.9×without changing input data formats. Ipoom Jeong, Jinghan Huang 0001, Chuxuan Hu, Dohyun Park, Jaeyoung Kang 0004, Nam Sung Kim, Yongjoo Park |
ISCA | 1 |
| 2025 | A4: Microarchitecture-Aware LLC Management for Datacenter Servers with Emerging I/O DevicesabstractIn modern server CPUs, the Last-Level Cache (LLC) serves not only as a victim cache for higher-level private caches but also as a buffer for low-latency DMA transfers between CPU cores and I/O devices through Direct Cache Access (DCA).However, prior work has shown that high-bandwidth network-I/O devices can rapidly flood the LLC with packets, often causing significant contention with co-running workloads.One step further, this work explores hidden microarchitectural properties of the Intel Xeon CPUs, uncovering two previously unrecognized LLC contentions triggered by emerging high-bandwidth I/O devices.Specifically, (C1) DMAwritten cache lines in LLC ways designated for DCA (referred to as DCA ways) are migrated to certain LLC ways (denoted as inclusive ways) when accessed by CPU cores, unexpectedly contending with non-I/O cache lines within the inclusive ways.In addition, (C2) high-bandwidth storage-I/O devices, which are increasingly common in datacenter servers, benefit little from DCA while contending with (latency-sensitive) network-I/O devices within DCA ways.To this end, we present A4, a runtime LLC management framework designed to alleviate both (C1) and (C2) among diverse co-running workloads, using a hidden knob and other hardware features implemented in those CPUs.Additionally, we demonstrate that A4 can also alleviate other previously known network-I/Odriven LLC contentions.Overall, it improves the performance of latency-sensitive, high-priority workloads by 51% without notably compromising that of low-priority workloads. Haneul Park, Jiaqi Lou, Sangjin Lee 0003, KyoungSoo Park, Yongseok Son, Ipoom Jeong, Nam Sung Kim |
ISCA | 7 |
| 2025 | Intel ® in-Memory Analytics Accelerator: Performance Characterization and GuidelinesabstractImprovements in CPU performance have significantly slowed due to the demise of Dennard scaling, making it increasingly challenging to cost-effectively process the exponentially growing volumes of data. As a result, there is a growing trend of offloading frequently used functions to hardware accelerators to enhance application performance while reducing expensive CPU cycle consumption. In line with this trend, Intel introduced the InMemory Analytics Accelerator (IAA) as an on-chip accelerator in its$\mathbf{4}^{\text{th}}$-generation Xeon®Scalable CPUs (Sapphire Rapids). IAA is designed to offload common big data and in-memory analytics functions from CPUs, such as CRC64, expand, extract, scan, and select, in addition to (de)compression. As IAA is integrated with CPUs as an on-chip accelerator, it can directly access the CPU's cache and memory in a cache-coherent manner, offering low latency and power consumption with reduced programming complexity compared to off-chip accelerators. In this paper, we first introduce the hardware and software architectures of IAA, highlighting its latest features. We then evaluate IAA's performance using various microbenchmarks and widely used analytics applications, including Pandas, Citus, and ClickHouse. Finally, based on our findings, we provide guidelines for effectively utilizing and optimizing IAA for analytics applications. Jaeyoung Kang 0004, Qirong Xia, Ipoom Jeong, Yongjoo Park, Nam Sung Kim |
ISPASS | 3 |
| 2025 | Re-architecting End-host Networking with CXL: Coherence, Memory, and Offloading
Houxiang Ji, Yang Zhou 0050, Ipoom Jeong, Ren Wang 0001, Saksham Agarwal, Nam Sung Kim |
MICRO | 4 |
| 2024 | A Quantitative Analysis and Guidelines of Data Streaming Accelerator in Modern Intel Xeon Scalable ProcessorsabstractAs semiconductor power density is no longer constant with the technology process scaling down, we need different solutions if we are to continue scaling application performance. To this end, modern CPUs are integrating capable data accelerators on the chip, aiming to improve performance and efficiency for a wide range of applications and usages. One such accelerator is the Intel® Data Streaming Accelerator (DSA) introduced since Intel® 4th Generation Xeon® Scalable CPUs (Sapphire Rapids). DSA targets data movement operations in memory that are common sources of overhead in datacenter workloads and infrastructure. In addition, it supports a wider range of operations on streaming data, such as CRC32 calculations, computation of deltas between data buffers, and data integrity field (DIF) operations. This paper aims to introduce the latest features supported by DSA, dive deep into its versatility, and analyze its throughput benefits through a comprehensive evaluation with both microbenchmarks and real use cases. Along with the analysis of its characteristics and the rich software ecosystem of DSA, we summarize several insights and guidelines for the programmer to make the most out of DSA, and use an in-depth case study of DPDK Vhost to demonstrate how these guidelines benefit a real application. Reese Kuper, Ipoom Jeong, Ren Wang 0001, Narayan Ranganathan, Philip Lantz, Nam Sung Kim |
ASPLOS (2) | 2 |
| 2024 | TAROT: A CXL SmartNIC-Based Defense Against Multi-bit Errors by Row-Hammer AttacksabstractRow Hammer (RH) has been demonstrated as a security vulnerability in modern systems. Although commodity CPUs can handle RH-induced single-bit errors in DRAM through ECC, RH can still give rise to multi-bit uncorrectable errors (UEs) and crash the systems. Meanwhile, recent work has indicated that the DRAM cells vulnerable to RH are determined by manufacturing imperfections and resulting defects. Taking one step further from the recent work, we first conduct RH experiments on contemporary DRAM modules for 3 weeks. This demonstrates that RH-induced UEs occur only at specific DRAM addresses (RH-UE-vulnerable addresses) and the percentage of such addresses is small in these DRAM modules. Second, to protect the systems from RH-induced UEs, we propose two RH defense solutions: H- and S-TAROT (TArgeted ROw-Hammer Therapy). H-TAROT is a software-based solution running on the host CPU. It obtains RH-UE-vulnerable addresses during the system boot and then periodically accesses such addresses before UEs may occur. Since it accesses only a small percentage of addresses, it does not incur a notable performance penalty for throughput applications (e.g., a 1.5% increase in execution time of the SPECrate 2017 benchmark suite running on a system even with 128GB of DRAM). Yet, it imposes a significant performance penalty on latency-sensitive applications (e.g., a 28.2% increase in tail latency of Redis). To minimize the performance penalty, for a system with a SmartNIC (SNIC), S-TAROT offloads H-TAROT from the host CPU to the SNIC CPU. Our experiment shows that S-TAROT increases the execution time and tail latency of the SPECrate 2017 benchmark suite and Redis by only 0.1% and 1.0%, respectively. Chihun Song, Michael Jaemin Kim, Houxiang Ji, Jinghan Huang 0001, Ipoom Jeong, Jaehyun Park 0006, Hwayong Nam, Minbok Wi, Jung Ho Ahn, Nam Sung Kim |
ASPLOS (3) | 6 |
| 2024 | ScaleCache: A Scalable Page Cache for Multiple Solid-State DrivesabstractThis paper presents a scalable page cache called ScaleCache for improving SSD scalability. Specifically, we first propose a concurrent data structure of page cache based on XArray (ccXArray) to enable access and update the page cache concurrently. Second, we introduce a direct page flush (dflush) which directly flushes pages to storage devices in a parallel and opportunistic manner. We implement ScaleCache with two techniques in the Linux kernel and evaluate it on a 64-core machine with eight NVMe SSDs. Our evaluations show that ScaleCache improves the performance of Linux file systems by up to 6.81× and 4.50× compared with the existing scheme and scalable scheme for multiple SSDs, respectively. Kiet Tuan Pham, Seokjoo Cho, Sangjin Lee 0003, Lan Anh Nguyen, Hyeongi Yeo, Ipoom Jeong, Sungjin Lee 0001, Nam Sung Kim, Yongseok Son |
EuroSys | 6 |
| 2024 | HAL: Hardware-assisted Load Balancing for Energy-efficient SNIC-Host Cooperative ComputingabstractA typical SmartNIC (SNIC) integrates a processor comprising Arm CPU and accelerators with a conventional NIC. The processor is designed to energy-efficiently execute network functions frequently used by datacenter applications. With such a processor, the SNIC has promised to notably improve the system-wide energy efficiency of datacenter servers. Nevertheless, the latest trend of integrating accelerators into server CPUs for these functions sparks a question on the SNIC processor’s superiority over a host processor (i.e., server CPU with accelerators) in system-wide energy efficiency, especially under given tail latency constraints. Answering this question, we first take an Intel Xeon processor, integrated with various accelerators (e.g., QuickAssist Technology), as a host processor, and then compare it to an NVIDIA BlueField-2 SNIC processor. This uncovers that (1) the host accelerator, coupled with a more powerful memory subsystem, can outperform the SNIC accelerator, and (2) the SNIC processor can improve system-wide energy efficiency only at low packet rates for most functions under tail latency constraints. To provide high system-wide energy efficiency without compromising tail latency at any packet rates, we propose HAL, consisting of a hardware-based load balancer and an intelligent load balancing policy implemented inside the SNIC. When HAL determines that the SNIC processor cannot efficiently process a given function beyond a specific packet rate, it limits the rate of packets to the SNIC processor and lets the host processor handle the excess. We implement a HAL-enabled SNIC with a commodity FPGA and a BlueField-2 SNIC, plug it into a commodity server, and run 10 popular network functions. Our evaluation shows that HAL can improve the system-wide energy efficiency and throughput of the server running these functions by 31% and 10%, respectively, without notably increasing the tail latency. Jinghan Huang 0001, Jiaqi Lou, Srikar Vanavasam, Xinhao Kong, Houxiang Ji, Ipoom Jeong, Danyang Zhuo, Nam Sung Kim |
ISCA | 6 |
| 2024 | Intel Accelerators Ecosystem: An SoC-Oriented Perspective : Industry ProductabstractA growing demand for hyperscale services has compelled hyperscalers to deploy more compute resources at an unprecedented pace, further accelerated by the demise of Dennard scaling. Meanwhile, a considerable portion of the compute resources are consumed to execute common functions present across the hyperscale services, i.e., datacenter taxes. These challenges motivated many to explore specialized accelerators for these functions. Leading such a technology trend, Intel has integrated diverse on-chip accelerators into its recent flagship datacenter CPU products. Furthermore, to support the easy and efficient use of these accelerators for successful deployment in production hyperscale services, Intel has developed a hardware/software ecosystem. In this paper, we first focus on Intel’s holistic efforts to build the hardware/software ecosystem, presenting key SoC-level features that facilitate efficient CPU-accelerator interaction, effortless programming and use, and scalable accelerator sharing and virtualization. Next, we delve into the functions, microarchitectures, and software stacks of three new on-chip accelerators: Data Streaming Accelerator (DSA), In-memory Analytics Accelerator (IAA), and Dynamic Load Balancer (DLB). Lastly, we demonstrate that Intel’s on-chip accelerators can not only significantly reduce the datacenter taxes but also accelerate data-intensive applications essential for hyperscale services, with little effort to use the accelerators. Ren Wang 0001, Narayan Ranganathan, Philip Lantz, Vivekananthan Sanjeepan, Jorge Cabrera 0003, Atul Kwatra, Rajesh Sankaran, Ipoom Jeong, Nam Sung Kim |
ISCA | 11 |
| 2024 | Demystifying a CXL Type-2 Device: A Heterogeneous Cooperative Computing PerspectiveabstractCXL is the latest interconnect technology built on PCIe, providing three protocols to facilitate three distinct types of devices, each with unique capabilities. Among these devices, a CXL Type-2 device has become commercially available, followed by CXL Type-3 devices. Therefore, it is timely to understand capabilities and characteristics of the CXL Type-2 device, as well as explore suitable applications. In this work, first, we delve into three key features of a CXL Type-2 device: cache-coherent device accelerator to host memory, device accelerator to device memory, and host CPU to device memory accesses. Second, using microbenchmarks, we comprehensively characterize the latency and bandwidth of these memory accesses with a CXL Type-2 device, and then compare them with those of equivalent memory accesses with comparable devices, such as emulated CXL Type-2, CXL Type-3, and PCIe devices. Lastly, as applications that exploit the unique capabilities of a CXL Type-2 device, we propose two CXL-based Linux memory optimization features: compressed RAM cache for swap (zswap) and memory deduplication (ksm). Our evaluation shows that Redis, when running with traditional CPU-based zswap and ksm, suffers from a tail latency increase of 4.5-10.3× compared to Redis running alone. While PCIe-based zswap and ksm still experience a tail latency increase of up to 8.1×, CXL-based zswap and ksm practically eliminate the tail latency increase with faster and more efficient host-device communication than PCIe-based zswap and ksm. Houxiang Ji, Srikar Vanavasam, Yang Zhou 0050, Qirong Xia, Jinghan Huang 0001, Ren Wang 0001, Pekon Gupta, Bhushan Chitlur, Ipoom Jeong, Nam Sung Kim |
MICRO | 10 |
| 2023 | INTERPRET: Inter-Warp Register Reuse for GPU Tensor CoreabstractTensor cores in the recent NVIDIA GPUs are under the spotlight due to their superior computation throughput for general matrix-matrix multiplication (GEMM) that has been widely used for deep learning applications. For massive-scale GEMMs, the entire matrix is practically divided into sub-matrices and assigned to multiple thread blocks and warps, and then processed by the tensor cores. Meanwhile, the same sub-matrix is regularly reused as an input to different sub-GEMMs, which causes redundant load operations from different warps and waste of register file spaces. To tackle this issue, we propose INTERPRET, a novel tensor core microarchitecture designed to minimize unnecessary accesses to the cache/memory hierarchy by leveraging the inter-warp data reuse characteristics. INTERPRET adopts a register renaming scheme to reduce the redundant load requests as well as the waste of register files, resulting in the reduction of the effective data load latency. INTERPRET further improves performance via non-speculative tensor preloading by leveraging the register file space saved by the register renaming. As INTERPRET is implemented based on the data access patterns of tensor core operations exhibiting a high level of regularity, the synergistic integration of the register renaming and tensor preloading can significantly improve the processing efficiency. Our experiments show that the proposed design achieves an average speedup of 34.1% and reduces energy consumption by 27.9%. Jae Seok Kwak, Myung Kuk Yoon, Ipoom Jeong, Seunghyun Jin 0002, Won Woo Ro |
PACT | 3 |
| 2023 | Demystifying CXL Memory with Genuine CXL-Ready Systems and DevicesabstractThe ever-growing demands for memory with larger capacity and higher bandwidth have driven recent innovations on memory expansion and disaggregation technologies based on Compute eXpress Link (CXL). Especially, CXL-based memory expansion technology has recently gained notable attention for its ability not only to economically expand memory capacity and bandwidth but also to decouple memory technologies from a specific memory interface of the CPU. However, since CXL memory devices have not been widely available, they have been emulated using DDR memory in a remote NUMA node. In this paper, for the first time, we comprehensively evaluate a true CXL-ready system based on the latest 4th-generation Intel Xeon CPU with three CXL memory devices from different manufacturers. Specifically, we run a set of microbenchmarks not only to compare the performance of true CXL memory with that of emulated CXL memory but also to analyze the complex interplay between the CPU and CXL memory in depth. This reveals important differences between emulated CXL memory and true CXL memory, some of which will compel researchers to revisit the analyses and proposals from recent work. Next, we identify opportunities for memory-bandwidth-intensive applications to benefit from the use of CXL memory. Lastly, we propose a CXL-memory-aware dynamic page allocation policy, Caption to more efficiently use CXL memory as a bandwidth expander. We demonstrate that Caption can automatically converge to an empirically favorable percentage of pages allocated to CXL memory, which improves the performance of memory-bandwidth-intensive applications by up to 24% when compared to the default page allocation policy designed for traditional NUMA systems. Zeduo Yu, Reese Kuper, Chihun Song, Jinghan Huang 0001, Houxiang Ji, Siddharth Agarwal, Jiaqi Lou, Ipoom Jeong, Ren Wang 0001, Jung Ho Ahn, Tianyin Xu, Nam Sung Kim |
MICRO | 10 |
| 2023 | A convertible neural processor supporting adaptive quantization for real-time neural networks
Hongju Kal, Hyoseong Choi, Ipoom Jeong, Joon-Sung Yang, Won Woo Ro |
J. Syst. Archit. | 3 |
| 2022 | Reconstructing Out-of-Order Issue QueueabstractOut-of-order cores provide high performance at the cost of energy efficiency. Dynamic scheduling is one of the major contributors to this: generating highly optimized issue schedules considering both data dependences and underlying execution resources, but relying heavily on complex wakeup and select operations of an out-of-order issue queue (IQ). For decades, researchers have proposed several complexity-effective dynamic scheduling schemes by leveraging the energy efficiency of an in-order IQ. However, they are either costly or not capable of delivering sufficient performance to substitute for a conventional wide-issue out-of-order IQ. In this work, we revisit two previous designs: one classical dependence-based design and the other state-of-the-art readiness-based design. We observe that they are complementary to each other, and thus their synergistic integration has the potential to be a good alternative to an out-of-order IQ. We first combine these two designs, and further analyze the main architectural bottlenecks that incur the underutilization of aggregate issue capability, thereby limiting the exploitation of instruction-level and memory-level parallelisms: 1) memory dependences not exposed by the register-based dependence analysis and 2) wide and shallow nature of dynamic dependence chains due to the long-latency memory accesses. To this end, we propose Ballerino, a novel microarchitecture that performs balanced and cache-miss-tolerable dynamic scheduling via a complementary combination of cascaded and clustered in-order IQs. Ballerino is built upon three key functionalities: 1) speculatively filtering out ready-at-dispatch instructions, 2) eliminating wasteful wakeup operations via a simple steering technique leveraging the awareness of memory dependences, and 3) reacting to program phase changes by allowing different load-dependent chains to share a single IQ while guaranteeing their out-of-order issue. The net effect is minimal scheduling energy consumption per instruction while providing comparable scheduling performance to a fully out-of-order IQ. In our analysis, Ballerino achieves comparable performance to an 8-wide out-of-order core by using twelve in-order IQs, improving core-wide energy efficiency by 20%. Ipoom Jeong, Jiwon Lee 0001, Myung Kuk Yoon, Won Woo Ro |
MICRO | 1 |
| 2020 | CASINO Core Microarchitecture: Generating Out-of-Order Schedules Using Cascaded In-Order Scheduling WindowsabstractThe performance gap between in-order (InO) and out-of-order (OoO) cores comes from the ability to dynamically create highly optimized instruction issue schedules. In this work, we observe that a significant amount of performance benefit of OoO scheduling can also be attained by supplementing a traditional InO core with a small and speculative instruction scheduling window, namely SpecInO. SpecInO monitors a small set of instructions ahead of a conventional InO scheduling window, aiming at issuing ready instructions behind long-latency stalls. Simulation results show that SpecInO captures and issues 62% of dynamic instructions out of program order. To this end, we propose a CASINO core microarchitecture that dynamically and speculatively generates OoO schedules with near-InO complexity, using CAScaded IN-Order scheduling windows. A Speculative IQ (S-IQ) issues an instruction if it is ready, or otherwise passes it to the next IQ. At the last IQ, instructions are scheduled in program order along serial dependence chains. The net effect is OoO scheduling via collaboration between cascaded InO IQs. To support speculative execution with minimal cost overhead, we propose a novel register renaming technique that allocates free physical registers only to instructions issued from the S-IQ. The proposed core performs dynamic memory disambiguation via an on-commit value check by extending the store buffer already existing in an InO core. We further optimize energy efficiency by filtering out redundant associative searches performed by speculated loads. In our analysis, CASINO core improves performance by 51% over an InO core (within 10 percentage points of an OoO core), which results in 25% and 42% improvements in energy efficiency over InO and OoO cores, respectively. Ipoom Jeong, Seihoon Park, Changmin Lee 0002, Won Woo Ro |
HPCA | 1 |
| 2019 | OverCome: Coarse-Grained Instruction Commit with Handover Register RenamingabstractCoarse-grained instruction commit mechanisms enabled the effective size of the instruction window to be as large as possible by committing a group of instructions atomically. Within a group, the reorder buffer (ROB) and physical registerfile (PRF) entries are conservatively managed, and thus the instruction window can handle more in-flight instructions beyond the hardware limit. However, previous approaches have suffered from high storage requirements for managing group information and unbalanced lifetime of instruction window resources, i.e., the ROB and PRF. In this paper, we propose an OverCome microarchitecture based on a history-based approach to address these problems. First, OverCome retains the conservative allocation of the ROB regardless of the group size limit, thereby providing high scalability. Second, it handles the information of numerous groups with a low storage cost. These two techniques achieve a significant reduction in the pressure on the ROB; thus, a new bottleneck arises: the pressure on the PRF. To address this issue, we propose a novel register renaming technique to reduce the lifetime of physical registers to a large extent, by tightly coupling the early release and lazy allocation schemes. Thus, the proposed design strikes a balance between the ROB and PRF requirements. Detailed evaluation of the proposed techniques on a state-of-the-art superscalar processor shows that our proposals augment the effective size of the instruction window by more than 4×, with a net overhead of less than 3 percent of the core area. Ipoom Jeong, Changmin Lee 0002, Keunsoo Kim, Won Woo Ro |
IEEE Trans. Computers | 1 |