EDBT 2026 Demo / reviewers in the wild / expert
Eishi Arima
dblp:129/7741
· DBLP profile ↗
7ranked-venue papers
3as first author
3since 2021 · last 2025
0009-0002-7043-4288ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cache Miss Curve Analysis via Cardinality DomainabstractAnalyzing and understanding the memory access behaviors of applications are essential when optimizing computing systems and applications. One prominent example is the cache miss curve estimation, i.e., detecting the cache miss ratio or frequency as a function of the capacity using a memory access sequence. Historically, cache miss curves have been derived from their corresponding stack distance distributions that are obtained by keeping track of the depth on the LRU stack. As this procedure requires significant computational complexity, a variety of approximation techniques have been proposed ever since it was originally proposed by Mattson et al. in 70s. We, however, claim that stack distances are not necessarily required to derive a cache miss curve.This paper proposes an alternative, efficient approach: instead of relying on the stack distances, we approximate a cache miss curve from the cardinality of accesses - total access count as a function of unique access count. By doing so, we can apply wellestablished efficient probabilistic data structures (e.g., Log-Log counting and its variants) widely used in various domains. In our approach, we model LRU stack behavior as a series of Bernoulli trials and derive a macroscopical relationship between a cache miss curve and its corresponding cardinality curve. Driven by this relationship, we offer our miss curve estimation mechanism using cardinality detection data structures and a curve fitting approach. We further offer several advanced techniques by taking advantage of the cardinality domain: (1) a compensation mechanism for sampled memory traces and (2) an algorithm to synthesize a miss curve of multiple access sequences by following the additivity principle in the cardinality domain. We comprehensively validate our approach using SPEC CPU 2017 benchmark suite under a variety of scenarios. Eishi Arima, Martin Schulz 0001 |
PACT | 1 |
| 2024 | Reinforcement Learning-Driven Co-Scheduling and Diverse Resource Assignments on NUMA SystemsabstractAs modern HPC systems are typically composed of fat and rich compute nodes, it is usually difficult to fully utilize all node resources with a single application. Co-scheduling, i.e., co-executing multiple complementary applications (or jobs) on the same node in a space sharing manner, is a promising solution and thus has been widely studied in the past decade. As one major drawback of co-scheduling is that it induces the interference effects among co-located applications due to contention among shared resources, the industry has started to support several resource/traffic partitioning features, e.g., in shared caches or memory controllers, on modern commercial processors. Recent studies proposed effective approaches to make use of these advanced features, however, the interactions between these features and (1) job scheduling decisions as well as (2) NUMA (Non-Uniform Memory Access) effects were generally overlooked. This paper explicitly targets these two missing pieces and comprehensively harmonizes the following decisions using reinforcement learning: (a) job selections for co-execution from a given job queue; and (b) diverse resource assignments to co-executed jobs, leveraging emerging hardware partitioning features, while taking NUMA-awareness into account. Our evaluation result demonstrates that our approach can improve the total system throughput by up to 78.1% over time sharing-based naive scheduling. Urvij Saroliya, Eishi Arima, Dai Liu, Martin Schulz 0001 |
ICCD | 2 |
| 2023 | Hierarchical Resource Partitioning on Modern GPUs: A Reinforcement Learning ApproachabstractGPU-based heterogeneous architectures are now commonly used in HPC clusters. Due to their architectural simplicity specialized for data-level parallelism, GPUs can offer much higher computational throughput and memory bandwidth than CPUs in the same generation do. However, as the available resources in GPUs have increased exponentially over the past decades, it has become increasingly difficult for a single program to fully utilize them. As a consequence, the industry has started supporting several resource partitioning features in order to improve the resource utilization by co-scheduling multiple programs on the same GPU die at the same time.Driven by the technological trend, this paper focuses on hierarchical resource partitioning on modern GPUs, and as an example, we utilize a combination of two different features available on recent NVIDIA GPUs in a hierarchical manner: MPS (Multi-Process Service), a finer-grained logical partitioning; and MIG (Multi-Instance GPU), a coarse-grained physical partitioning. We propose a method for comprehensively co-optimizing the setup of hierarchical partitioning and the selection of co-scheduling groups from a given set of jobs, based on reinforcement learning using their profiles. Our thorough experimental results demonstrate that our approach can successfully set up job concurrency, partitioning, and co-scheduling group selections simultaneously. This results in a maximum throughput improvement by a factor of 1.87 compared to the time-sharing scheduling. Urvij Saroliya, Eishi Arima, Dai Liu, Martin Schulz 0001 |
CLUSTER | 2 |
| 2020 | Evaluation of Power Management Control on the Supercomputer FugakuabstractThe supercomputer “Fugaku”, which recently ranked number one on multiple supercomputing lists, including the Top500 in June 2020, has various power control features, such as (1) an eco mode that utilizes only one of two floating-point pipelines while decreasing the power supply to the chip; (2) a boost mode that increases clock frequency; and (3) a core retention function that turns unused cores into a low-power state. By orchestrating these power-performance features while considering the characteristics of currently running applications, we can potentially gain even better system-level energy efficiency. In this article, we report on the effectiveness of these features using the pre-evaluation environment for Fugaku. As a result, we confirmed several prominent results useful for the operation of the Fugaku system, including: remarkable power reduction and energy-efficiency improvement by coordinating the eco mode and the core retention feature in the memory intensive case; a 10% speed-up with a 17% power consumption increase using the boost mode in the CPU intensive case; and considerable power variations across over 20K nodes. Yuetsu Kodama, Tetsuya Odajima, Eishi Arima, Mitsuhisa Sato |
CLUSTER | 3 |
| 2020 | Classification-Based Unified Cache Replacement via Partitioned Victim Address HistoryabstractIn modern microprocessors, lower level cache memories are usually implemented as unified caches where different classes of cachelines such as data, instructions, and Page Table Entries (PTEs) coexist. Particularly, frequent PTE accesses following after TLB missies can happen on modern systems, which is driven by the increasing demands of applications for larger working set size, and this trend naturally leads to significant conflicts among these different kinds of cachelines.This paper targets the emerging conflict problem and provides a systematic mechanism using a partitioned victim address history. Prior studies have shown the effectiveness of history-based cache managements to predict the reuseness and thus to improve the hit rate. This work augments the following functionalities: (1) partitioning the history into multiple areas to separately keep track of the reuseness for all the different cacheline categories; and (2) setting different allocation priorities to the different cacheline categories when cache replacement. Furthermore, this paper proposes a control system to dynamically optimize the history partitions and the cache allocation priorities at the same time by using the statistics of the history structure. The experimental result indicates that the proposed technique improves performance considerably compared with the conventional LRU-based approach and others. Eishi Arima |
DSD | 1 |
| 2015 | Immediate sleep: Reducing energy impact of peripheral circuits in STT-MRAM cachesabstractImplementing last level caches (LLCs) with STT-MRAM is a promising approach for designing energy efficient microprocessors due to high density and low leakage power of its memory cells. However, peripheral circuits of an STT-MRAM cache still suffer from leakage power because large and leaky transistors are required to drive large write current to STT-MRAM element. To overcome this problem, we propose a new power management scheme called Immediate Sleep (IS). IS immediately turns off a subarray of an STT-MRAM cache if the next access is predicted to be not critical in performance. Thus, IS can effectively reduce leakage energy with little impact on performance. Our experimental results show that our technique can save the leakage energy of an STT-MRAM LLC by 32% compared to an STT-MRAM LLC with the conventional scheme at the same performance. Eishi Arima, Hiroki Noguchi, Takashi Nakada, Shinobu Miwa, Susumu Takeda, Shinobu Fujita, Hiroshi Nakamura |
ICCD | 1 |
| 2013 | D-MRAM cache: enhancing energy efficiency with 3T-1MTJ DRAM/MRAM hybrid memoryabstractThis paper describes a proposal of non-volatile cache architecture utilizing novel DRAM / MRAM cell-level hybrid structured memory (D-MRAM) that enables effective power reduction for high performance mobile SoCs without area overhead. Here, the key point to reduce active power is intermittent refresh process for the DRAM-mode. D-MRAM has advantage to reduce static power consumptions compared to the conventional SRAM, because there are no static leakage paths in the D-MRAM cell and it is not needed to supply voltage to its cells when used as the MRAM-mode. Besides, with advanced perpendicular magnetic tunnel junctions (p-MTJ), which decreases the write energy and latency without shortening its retention time, D-MRAM is capable of power reduction by replacing the traditional SRAM caches. Considering the 65-nm CMOS technology, the access latencies of 1MB memory macro are 2.2 ns / 1.5 ns for read / write in DRAM mode, and 2.2 ns / 4.5 ns in MRAM mode, while those of SRAM are 1.17 ns. The SPEC CPU2006 benchmarks have revealed that the energy per instruction (EPI) of the total cache memory can be dramatically reduced by 71 % on average, and the instruction per cycle (IPC) performance of the D-MRAM cache architecture degraded only by approximately 4 % on average in spite of its latency overhead. Hiroki Noguchi, Kumiko Nomura, Keiko Abe, Shinobu Fujita, Eishi Arima, Kyundong Kim, Takashi Nakada, Shinobu Miwa, Hiroshi Nakamura |
DATE | 5 |