EDBT 2026 Demo / reviewers in the wild / expert
Toru Koizumi 0001
dblp:196/1934-1
· DBLP profile ↗
14ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-0990-1916ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-author · 9 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RUNLTS: Branch Prediction with Register-Value Correlations and Hierarchical Table Orchestration
Toru Koizumi 0001, Toshiki Maekawa, Masanari Mizuno, Maru Kuroki, Tomoaki Tsumura, Ryota Shioya |
ISCA | 1 |
| 2025 | Trailing-Ones Anticipation for Reducing the Latency of the Rounding Incrementer in FP FMA UnitsabstractFloating-point fused multiply-add (FMA) operations are fundamental in many fields such as scientific computing, graphics processing, and machine learning. In conventional floating-point FMA designs, the internal steps of multiplication and addition, which are performed using integer arithmetic, are followed by a post-processing stage. It has been known that the post-processing stage contributes to approximately 60% of the total latency of the double-precision FMA operation. We propose a novel trailing-ones anticipation technique that predicts trailing-ones bits of the mantissa before rounding, in parallel with the post-processing stage. With this technique, the rounding incrementer can be implemented using a single XOR operation, thereby reducing the total latency. We evaluated the latency using Synopsys Design Compiler for synthesis, confirming that the proposed technique reduced total latency by 4 %. Toru Koizumi 0001, Ryota Shioya, Takuya Yamauchi, Tomoya Adachi, Ken Namura, Jun Makino |
ARITH | 1 |
| 2025 | PEZY-SC4s : The Fourth Generation MIMD Many-core Processor with High Energy Efficiency and Flexibility for HPC and AI Applications
Naoya Hatta, Shuntaro Tsunoda, Kouhei Uchida, Taichi Ishitani, Toru Koizumi 0001, Ryota Shioya, Kei Ishii |
HCS | 5 |
| 2025 | Register Bridging: A Lightweight Microarchitectural Approach for Skipping Overhead Instructions in Distance-Based ISA ProcessorsabstractOut-of-order superscalar processors achieve high performance at the cost of control complexity and energy overhead, with register renaming contributing significantly. Distancebased instruction set architectures (ISAs) provide an alternative to avoid register renaming by specifying operands using relative instruction distances. As a representative design, STRAIGHT implements this approach to support out-of-order execution and eliminate false dependencies in a lightweight design. However, to simplify the overall system design and ensure clear instruction semantics, distance-based architectures require additional instructions (e.g., RMOV) to adjust operand distances, which consume execution resources and, more critically, may delay dependent instructions, resulting in performance degradation. In this paper, we propose Register Bridging, a mechanism that redirects semantically equivalent operands to bypass RMOV dependencies, enabling parallel execution of instructions previously constrained by data-flow ordering. Specifically, a circular buffer is introduced to support operand redirecting with low complexity, in contrast to traditional renaming tables. We implemented the proposed method on a cycle-accurate simulator and compiled benchmarks using the optimizing STRAIGHT compiler. Through a series of simulation experiments on both realistic and synthetic benchmarks, we demonstrate that our proposal enables 41.9% of relay instructions to be bypassed on average, as well as improves performance by up to 5.7%, compared to related methods. Toru Koizumi 0001, Shu Sugita, Yuriko Yamauchi, Ryota Shioya, Junichiro Kadomoto, Hidetsugu Irie |
ICCD | 2 |
| 2023 | A Sound and Complete Algorithm for Code Generation in Distance-Based ISAabstractThe single-thread performance of a processor core is essential even in the multicore era. However, increasing the processing width of a core to improve the single-thread performance leads to a super-linear increase in power consumption. To overcome this power consumption issue, an instruction set architecture for general-purpose processors, called STRAIGHT, has been proposed. STRAIGHT adopts a distance-based ISA, in which source operands are specified by the distance between instructions. In STRAIGHT, it is necessary to satisfy constraints on the distance used as operands to generate executable code. However, it is not yet clear how to generate code that satisfies these constraints in the general case. In this paper, we propose three compiling techniques for STRAIGHT code generation and prove that our techniques can reliably generate code that satisfies the distance constraints. We implemented the proposed method on a compiler and evaluated benchmark programs compiled with it through simulation. The evaluation results showed that the proposed method works in all cases, including conditions where the number of registers is small and existing methods fail to generate code. Shu Sugita, Toru Koizumi 0001, Ryota Shioya, Hidetsugu Irie, Shuichi Sakai |
CC | 2 |
| 2023 | TURBULENCE: Complexity-effective Out-of-order Execution on GPU with Distance-based ISAabstractA graphic processing unit (GPU) is a processor that achieves high throughput by exploiting data parallelism. We found that many GPU workloads also contain instruction-level parallelism, which can be extracted through out-of-order execution to provide additional performance improvement opportunities. We propose the TURBULENCE architecture for very low-cost out-of-order execution on GPUs. TURBULENCE consists of 1) a novel ISA that introduces the concept of referencing operands by inter-instruction distance instead of register numbers and 2) a novel microarchitecture that executes the novel ISA. Our proposed ISA and microarchitecture enable cost-effective out-of-order execution on GPUs without introducing expensive hardware. Reoma Matsuo, Toru Koizumi 0001, Hidetsugu Irie, Shuichi Sakai, Ryota Shioya |
DATE | 2 |
| 2023 | An Out-of-Order Superscalar Processor Using STRAIGHT Architecture in 28 nm CMOSabstractThe single-thread performance of a CPU is an essential factor in a computer system. However, increasing the processing width of a CPU to improve performance often results in a super-linear enlargement of the circuit area and, consequently, a massive increase in power consumption. In this paper, we present an out-of-order superscalar processor based on a new architecture, STRAIGHT, which overcomes the circuit area and power consumption problems. We have designed and evaluated the first real processor chip based on the STRAIGHT architecture. The processor chip was fabricated using 28nm CMOS technology, and we confirmed that it could correctly execute real programs. We evaluated its performance, circuit area, and power consumption, and as a result, demonstrated that a large processing width can be achieved in a small area using the new STRAIGHT architecture. Taichi Amano, Junichiro Kadomoto, Satoshi Mitsuno, Toru Koizumi 0001, Ryota Shioya, Hidetsugu Irie, Shuichi Sakai |
ISCAS | 4 |
| 2023 | Clockhands: Rename-free Instruction Set Architecture for Out-of-order ProcessorsabstractOut-of-order superscalar processors are currently the only architecture that speeds up irregular programs, but they suffer from poor power efficiency. To tackle this issue, we focused on how to specify register operands. Specifying operands by register names, as conventional RISC does, requires register renaming, resulting in poor power efficiency and preventing an increase in the front-end width. In contrast, a recently proposed architecture called STRAIGHT specifies operands by inter-instruction distance, thereby eliminating register renaming. However, STRAIGHT has strong constraints on instruction placement, which generally results in a large increase in the number of instructions. Toru Koizumi 0001, Ryota Shioya, Shu Sugita, Taichi Amano, Yuya Degawa, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai |
MICRO | 1 |
| 2022 | T-SKID: Predicting When to Prefetch Separately from Address PredictionabstractPrefetching is an important technique for reducing the number of cache misses and improving processor performance, and thus various prefetchers have been proposed. Many prefetchers are focused on issuing prefetches sufficiently earlier than demand accesses to hide miss latency. In contrast, we propose aT-SKID prefetcher, which focuses on delaying prefetching. If a prefetcher issues prefetches for demand accesses too early, the prefetched line will be evicted before it is referenced. We found that existing prefetchers often issue such too-early prefetches, and this observation offers new opportunities to improve performance. To tackle this issue, T-SKID performs timing prediction indepen-dently of address prediction. In addition to issuing prefetches sufficiently early as existing prefetchers do, T-SKID can delay the issue of prefetches until an appropriate time if necessary. We evaluated T-SKID by simulations using SPEC CPU 2017. The result shows that T-SKID achieves a 5.6 % performance improve-ment for multi-core environment, compared to Instruction Pointer Classifier based Prefetching, which is a state-of-the-art prefetcher. Toru Koizumi 0001, Tomoki Nakamura, Yuya Degawa, Hidetsugu Irie, Shuichi Sakai, Ryota Shioya |
DATE | 1 |
| 2021 | Compiling and Optimizing Real-world Programs for STRAIGHT ISAabstractThe renaming unit of a superscalar processor is a very expensive module. It consumes large amounts of power and limits the front-end bandwidth. To overcome this problem, an instruction set architecture called STRAIGHT has been proposed. Owing to its unique manner of referencing operands, STRAIGHT does not cause false dependencies and allows out-of-order execution without register renaming. However, the compiler optimization techniques for STRAIGHT are still immature, and we found that the naive code generators currently available can generate inefficient code with additional instructions. In this paper, we propose two novel compiler optimization techniques and a novel calling convention for STRAIGHT to reduce the number of instructions. We compiled real-world programs with a compiler that implemented these techniques and measured their performance through simulation. The evaluation results show that the proposed methods reduced the number of executed instructions by 15% and improved the performance by 17%. Toru Koizumi 0001, Shu Sugita, Ryota Shioya, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai |
ICCD | 1 |
| 2021 | Accurate and Fast Performance Modeling of Processors with Decoupled Front-endabstractVarious techniques, such as cache replacement algorithms and prefetching, have been studied to prevent instruction cache misses from becoming a bottleneck in the processor frontend. In such studies, the goal of the design has been to reduce the number of instruction cache misses. However, owing to the increasing complexity of modern processors, the correlation between reducing instruction cache misses and reducing the number of executed cycles has become smaller than in previous cases. In this paper, we propose a new guideline for improving the performance of modern processors. In addition, we propose a method for estimating the approximate performance of a design two orders of magnitude faster than a full simulation each time the designers modify their design. Yuya Degawa, Toru Koizumi 0001, Tomoki Nakamura, Ryota Shioya, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai |
ICCD | 2 |
| 2020 | A High-Performance Out-of-Order Soft Processor Without Register RenamingabstractOwing to the growth of FPGA-based systems and the increasing complexity of applications, the demand for high-performance soft processors in FPGAs has increased. The performance of processors is enhanced through out-of-order (OoO) superscalar execution using a register renaming mechanism. However, the register renaming mechanism has two problems. First, it requires a register mapping table (RMT), which usually comprises a RAM with a large number of ports. A multi-port RAM is not suitable for an FPGA. Second, register renaming complicates recovery mechanisms for exceptions, such as branch mispredictions. These problems increase the usage of resources and hinder the improvement of performance. Recently, the STRAIGHT architecture was proposed to solve these problems. STRAIGHT has a unique instruction format and enables OoO execution without register renaming. This approach eliminates the RMT and makes the recovery operation more efficient. In this study, we demonstrate a high-performance OoO STRAIGHT soft processor by implementing several mechanisms for adopting the STRAIGHT architecture and fabricate the first STRAIGHT processor capable of executing practical complex programs. Compared to a state-of-the-art OoO soft processor, our processor consumes approximately 17% fewer LUTs and 10% fewer FlipFlops and achieves 15% higher performance in CoreMark, which is a standard benchmark. Satoshi Mitsuno, Junichiro Kadomoto, Toru Koizumi 0001, Ryota Shioya, Hidetsugu Irie, Shuichi Sakai |
FPL | 3 |
| 2018 | An Area-Efficient Out-of-Order Soft-Core Processor Without Register RenamingabstractIn this paper, we present an out-of-order soft-core processor adopting STRAIGHT architecture. STRAIGHT has a unique instruction format in which source operands are expressed as distances from producer instructions. This eliminates the need for register renaming and eliminates a register map table (RMT), which usually consists of a large multi-port RAM. That leads to small area, low power consumption, and high scalability of the front-end pipeline width. Moreover, the simplified architecture enables rapid miss-recovery. The prototype is implemented and evaluated on an FPGA. Compared to an out-of-order soft-core processor with a conventional RISC ISA, the proposed soft-core consumes 147-829 fewer LUTs for the front-end pipeline. The evaluation results show that the proposed soft-core is correctly operating on an FPGA, and estimated dynamic power consumption of the soft-core is 0.120 W. Junichiro Kadomoto, Toru Koizumi 0001, Akifumi Fukuda, Reoma Matsuo, Susumu Mashimo, Akifumi Fujita, Ryota Shioya, Hidetsugu Irie, Shuichi Sakai |
FPT | 2 |
| 2018 | STRAIGHT: Hazardless Processor Architecture Without Register RenamingabstractThe single-thread performance of a processor improves the capability of the entire system by reducing the critical path latency of programs. Typically, conventional superscalar processors improve this performance by introducing out-of-order (OoO) execution with register renaming. However, it is also known to increase the complexity and affect the power efficiency. This paper realizes a novel computer architecture called "STRAIGHT" to resolve this dilemma. The key feature is a unique instruction format in which the source operand is given based on the distance from the producer instruction. By leveraging this format, register renaming is completely removed from the pipeline. This paper presents the practical Instruction Set Architecture (ISA) design, the novel efficient OoO microarchitecture, and the compilation algorithm for the STRAIGHT machine code. Because the ISA has sequential execution semantics, as in general CPUs, and is provided with a compiler, programming for the architecture is as easy as that of conventional CPUs. A compiler, an assembler, a linker, and a cycle-accurate simulator are developed to measure the performance. Moreover, an RTL description of STRAIGHT is developed to estimate the power reduction. The evaluation using standard benchmarks shows that the performance of STRAIGHT is 18.8% better than the conventional superscalar processor of the same issue-width and instruction window size. This improvement is achieved by STRAIGHT's rapid miss-recovery. Compilation technology for resolving the possible overhead of the ISA is also revealed. The RTL power analysis shows that the architecture reduces the power consumption by removing the power for renaming. The revealed performance and efficiencies support that STRAIGHT is a novel viable alternative for designing general purpose OoO processors. Hidetsugu Irie, Toru Koizumi 0001, Akifumi Fukuda, Seiya Akaki, Satoshi Nakae, Yutaro Bessho, Ryota Shioya, Takahiro Notsu, Katsuhiro Yoda, Teruo Ishihara, Shuichi Sakai |
MICRO | 2 |