VLDB 2026 Research / reviewers in the wild / expert
Ryota Shioya
dblp:96/4041
· DBLP profile ↗
34ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0002-9309-5875ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 3 first-author · 15 since 2021Software engineering, systems software and programming languages · 13 · 1 first-author · 10 since 2021Security and privacy · 5 · 1 first-author · 2 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RUNLTS: Branch Prediction with Register-Value Correlations and Hierarchical Table Orchestration
Toru Koizumi 0001, Toshiki Maekawa, Masanari Mizuno, Maru Kuroki, Tomoaki Tsumura, Ryota Shioya |
ISCA | 6 |
| 2026 | A comprehensive analysis of the impact of sub 10-nm CNFET technology on 64-bit parallel prefix adders and 32-bit matrix multiply units
Chenlin Shi, Tongxin Yang, Ryota Shioya, Hayato Yamaki, Hiroki Honda, Shinobu Miwa |
Integr. | 3 |
| 2025 | Trailing-Ones Anticipation for Reducing the Latency of the Rounding Incrementer in FP FMA UnitsabstractFloating-point fused multiply-add (FMA) operations are fundamental in many fields such as scientific computing, graphics processing, and machine learning. In conventional floating-point FMA designs, the internal steps of multiplication and addition, which are performed using integer arithmetic, are followed by a post-processing stage. It has been known that the post-processing stage contributes to approximately 60% of the total latency of the double-precision FMA operation. We propose a novel trailing-ones anticipation technique that predicts trailing-ones bits of the mantissa before rounding, in parallel with the post-processing stage. With this technique, the rounding incrementer can be implemented using a single XOR operation, thereby reducing the total latency. We evaluated the latency using Synopsys Design Compiler for synthesis, confirming that the proposed technique reduced total latency by 4 %. Toru Koizumi 0001, Ryota Shioya, Takuya Yamauchi, Tomoya Adachi, Ken Namura, Jun Makino |
ARITH | 2 |
| 2025 | CACTI-CNFET: an Analytical Tool for Timing, Power, and Area of SRAMs with Carbon Nanotube Field Effect TransistorsabstractCarbon nanotube field effect transistors (CNFETs) are expected to replace silicon-based metal oxide semiconductor field effect transistors (MOSFETs) to improve the power efficiency and performance of microprocessors. However, the design of CNFET processors is not as mature as that of silicon-based processors because there is no architecture-level analytical tool for CNFET processors. Since circuit-level analysis such as RTL analysis is a very time-consuming and troublesome task, architecture-level analysis is needed for the rapid design of processors optimized for CNFETs. Shinobu Miwa, Eiichiro Sekikawa, Tongxin Yang, Ryota Shioya, Hayato Yamaki, Hiroki Honda |
ASP-DAC | 4 |
| 2025 | Biotite: A High-Performance Static Binary Translator using Source-Level InformationabstractResearch on novel Instruction Set Architectures (ISAs) is actively pursued; however, it requires extensive efforts to develop and maintain comprehensive compilation toolchains for each new ISA. Binary translation can provide a practical solution for ISA researchers to port target programs to novel ISAs when such a level of toolchain support is not available. However, to ensure the correct handling of indirect jumps, existing binary translators rely on complex runtime systems, whose implementation on primitive research ISAs demands significant efforts. ISA researchers generally have access to additional source-level information, including the symbol table and source code, when using binary translators. The symbol table can provide potential jump targets for optimizing indirect jumps, and ISA-independent functions in source code can be directly compiled without translation. Leveraging source-level information as additional input, in this paper, we propose Biotite, a high-performance static binary translator that correctly handles arbitrary indirect jumps. Currently, Biotite supports the translation of RV64GC Linux binaries to self-contained LLVM IR. Our evaluation shows that Biotite successfully translates all benchmarks in SPEC CPU 2017 and achieves a 2.346× performance improvement over QEMU for the integer benchmark suite. Changbin Chen, Shu Sugita, Yotaro Nada, Hidetsugu Irie, Shuichi Sakai, Ryota Shioya |
CC | 6 |
| 2025 | AceCov: Auxiliary Composite Edge Coverage for FuzzingabstractSecurity flaws in software, including bugs and vulnerabilities, can cause serious issues, and fuzzing has been extensively studied and employed to identify them. A typical fuzzer automatically generates test inputs and feeds back information obtained from executed programs for efficient software verification. As the feedback information, edge coverage, which is based on edge traversals in a control flow graph, is widely used. We focus on branches whose results change depending on the traversal of a particular edge in the control flow graph. In particular, we found that when a branch depends on a combination of multiple edge traversals, it is difficult for existing edge coverage to cover the branch outcome. Based on this observation, we propose a novel coverage metric for such branches: auxiliary composite-edge coverage (AceCov). AceCov works together with existing edge coverage to provide an auxiliary coverage metric only for branches such as the one described above, where the edge coverage does not cover the result. We implemented AceCov in AFL++ and evaluated it using FuzzBench and MAGMA. The evaluation results show a performance improvement of up to 3.36% and an average improvement of 1.19% over AFL++. We also identified four bugs that AFL++ failed to find. Haruki Yoshida, Yuichi Sugiyama, Ryota Shioya |
EuroS&P | 3 |
| 2025 | PEZY-SC4s : The Fourth Generation MIMD Many-core Processor with High Energy Efficiency and Flexibility for HPC and AI Applications
Naoya Hatta, Shuntaro Tsunoda, Kouhei Uchida, Taichi Ishitani, Toru Koizumi 0001, Ryota Shioya, Kei Ishii |
HCS | 6 |
| 2025 | Register Bridging: A Lightweight Microarchitectural Approach for Skipping Overhead Instructions in Distance-Based ISA ProcessorsabstractOut-of-order superscalar processors achieve high performance at the cost of control complexity and energy overhead, with register renaming contributing significantly. Distancebased instruction set architectures (ISAs) provide an alternative to avoid register renaming by specifying operands using relative instruction distances. As a representative design, STRAIGHT implements this approach to support out-of-order execution and eliminate false dependencies in a lightweight design. However, to simplify the overall system design and ensure clear instruction semantics, distance-based architectures require additional instructions (e.g., RMOV) to adjust operand distances, which consume execution resources and, more critically, may delay dependent instructions, resulting in performance degradation. In this paper, we propose Register Bridging, a mechanism that redirects semantically equivalent operands to bypass RMOV dependencies, enabling parallel execution of instructions previously constrained by data-flow ordering. Specifically, a circular buffer is introduced to support operand redirecting with low complexity, in contrast to traditional renaming tables. We implemented the proposed method on a cycle-accurate simulator and compiled benchmarks using the optimizing STRAIGHT compiler. Through a series of simulation experiments on both realistic and synthetic benchmarks, we demonstrate that our proposal enables 41.9% of relay instructions to be bypassed on average, as well as improves performance by up to 5.7%, compared to related methods. Toru Koizumi 0001, Shu Sugita, Yuriko Yamauchi, Ryota Shioya, Junichiro Kadomoto, Hidetsugu Irie |
ICCD | 6 |
| 2024 | ReOVE: Restricted Out-of-Order Execution for Superscalar Processors with Vector ExtensionabstractVector instructions have recently been introduced in general-purpose CPUs. These general-purpose CPUs typically employ out-of-order superscalar processors to execute scalar instructions at high speed. However, the naive implementation of vector instructions in an out-of-order CPU can require considerable circuit area due to the complex register and memory access inherent to vector instructions. While reordering between vector and scalar instructions is crucial for performance in out-of-order CPUs, we found that reordering among vector instructions does not significantly contribute to performance improvement. Based on this observation, we propose ReOVE, a novel architecture designed to implement vector instructions in out-of-order CPUs while maintaining low hardware costs by introducing partial restrictions on vector instruction reordering. We evaluated ReOVE on a cycle-accurate simulator with the RISC-V vector extension. The evaluation results show that ReOVE reduces the energy consumption and circuit area by 21.9% and 21.1%, respectively, with a performance degradation of 2.7% compared to existing out-of-order CPUs. Masayuki Kimura, Ryota Shioya |
ISLPED | 2 |
| 2024 | Dynamic Possible Source Count Analysis for Data Leakage PreventionabstractDynamic Taint Analysis (DTA) is a widely studied technique that can effectively detect various attacks and information leakage. In the context of detecting information leakage, taint is a flag added to data to indicate whether secret data can be inferred from it. DTA tracks the flow of tainted data in a language runtime environment and identifies secret data leakage when tainted data is transmitted externally. We found that existing DTAs can produce false negatives and false positives in complex data flows because of the binary nature of taint. Since taint is binary, meaning either secret data is inferable (=1) or non-inferable (=0), it cannot represent intermediate states that may slightly infer the secret data, and these states are quantized to 0 or 1. As a result of this quantization, existing methods are unable to distinguish between outputs that are practically secure and those that pose a real security threat in complex data flows, resulting in false positives and false negatives. To address this problem, we introduce the concept of Possible Source Count (PSC) and propose Dynamic Possible source Count Analysis (DPCA), which tracks PSC instead of taint. PSC is a metric that indicates how many secrets can be identified by observing the data. DPCA tracks and computes the PSC of each data item using dynamic symbolic execution. By evaluating the PSC of data that reaches the sink point, DPCA can effectively distinguish between data that is practically secure and data that poses a security threat. Eri Ogawa, Tetsuro Yamazaki, Ryota Shioya |
MPLR | 3 |
| 2024 | Dynamic Controllability Analysis for Preventing Injection AttacksabstractInjection attacks are some of the most serious security threats, and various techniques have been studied to prevent such attacks through program analysis. One of the typical dynamic analysis methods is Dynamic Taint Analysis (DTA), which adds a flag called taint to externally input data and detects an injection attack when these data reach a sink point where the system can be manipulated. However, DTA- based attack detection may produce many false positives and false negatives, especially in complex data flows. We consider that the high rate of false positives and negatives arises because the taint in DTA indicates whether data was controlled, not how much data was controlled. We propose Dynamic Controllability Analysis (DCA), an approach that approximates controllability by generalizing binary taint into natural numbers, indicating the extent of data control. We implemented DCA on a JavaScript runtime and evaluated the controllability computed by DCA. The evaluation results show that the controllability computed by DCA is sensitive to the presence or absence of an injection attack, yielding very low values when the system is safe and very high values when an attack is present. Eri Ogawa, Tetsuro Yamazaki, Ryota Shioya |
PRDC | 3 |
| 2023 | CNFET7: An Open Source Cell Library for 7-nm CNFET TechnologyabstractIn this paper, we propose CNFET7, the first open-source cell library for 7-nm carbon nanotube field-effect transistor (CNFET) technology. CNFET7 is based on an open-source CNFET SPICE model called VS-CNFET, and various model parameters such as the channel width and carbon nanotube diameter are carefully tuned to mimic the predictive 7-nm CNFET technology presented in a published paper. Some nondisclosure parameters, such as the cell size and pin layout, are derived from those of the NanGate 15-nm open-source cell library in the same way as for an open-source framework for CNFET circuit design. CNFET7 includes two types of delay model (i.e., the composite current source and nonlinear delay model), each having 56 cells, such as INV_X1 and BUF_X1. CNFET7 supports both logic synthesis and timing-driven place and route in the Cadence design flow. Our experimental results for several synthesized circuits show that CNFET7 has reductions of up to 96%, 62% and 82% in dynamic and static power consumption and critical-path delay, respectively, when compared with ASAP7. Chenlin Shi, Shinobu Miwa, Tongxin Yang, Ryota Shioya, Hayato Yamaki, Hiroki Honda |
ASP-DAC | 4 |
| 2023 | A Sound and Complete Algorithm for Code Generation in Distance-Based ISAabstractThe single-thread performance of a processor core is essential even in the multicore era. However, increasing the processing width of a core to improve the single-thread performance leads to a super-linear increase in power consumption. To overcome this power consumption issue, an instruction set architecture for general-purpose processors, called STRAIGHT, has been proposed. STRAIGHT adopts a distance-based ISA, in which source operands are specified by the distance between instructions. In STRAIGHT, it is necessary to satisfy constraints on the distance used as operands to generate executable code. However, it is not yet clear how to generate code that satisfies these constraints in the general case. In this paper, we propose three compiling techniques for STRAIGHT code generation and prove that our techniques can reliably generate code that satisfies the distance constraints. We implemented the proposed method on a compiler and evaluated benchmark programs compiled with it through simulation. The evaluation results showed that the proposed method works in all cases, including conditions where the number of registers is small and existing methods fail to generate code. Shu Sugita, Toru Koizumi 0001, Ryota Shioya, Hidetsugu Irie, Shuichi Sakai |
CC | 3 |
| 2023 | Out-of-Step Pipeline for Gather/Scatter InstructionsabstractWider SIMD units suffer from low scalability of gather/scatter instructions that appear in sparse matrix calculations. We address this problem with an out-of-step pipeline which tolerates bank conflicts of a multibank L1D by allowing element operations of SIMD instructions to proceed out of step with each other. We evaluated it with a sparse matrix-vector product kernel for matrices from HPCG and SuiteSparse Matrix Collection. The results show that, for the SIMD width of 1024 bit, it achieves 1.91 times improvement over a model of a conventional pipeline. Yi Ge, Katsuhiro Yoda, Makiko Ito, Toshiyuki Ichiba, Takahide Yoshikawa, Ryota Shioya, Masahiro Goshima |
DATE | 6 |
| 2023 | TURBULENCE: Complexity-effective Out-of-order Execution on GPU with Distance-based ISAabstractA graphic processing unit (GPU) is a processor that achieves high throughput by exploiting data parallelism. We found that many GPU workloads also contain instruction-level parallelism, which can be extracted through out-of-order execution to provide additional performance improvement opportunities. We propose the TURBULENCE architecture for very low-cost out-of-order execution on GPUs. TURBULENCE consists of 1) a novel ISA that introduces the concept of referencing operands by inter-instruction distance instead of register numbers and 2) a novel microarchitecture that executes the novel ISA. Our proposed ISA and microarchitecture enable cost-effective out-of-order execution on GPUs without introducing expensive hardware. Reoma Matsuo, Toru Koizumi 0001, Hidetsugu Irie, Shuichi Sakai, Ryota Shioya |
DATE | 5 |
| 2023 | SurgeFuzz: Surge-Aware Directed Fuzzing for CPU DesignsabstractVarious verification methods have been proposed for bug detection in central processing unit (CPU) designs, yet their effectiveness remains insufficient. We have observed that such CPU bugs often occur in exceptional handling, such as pipeline stalls and flushes. We found that corner cases in such exceptional handling can be effectively verified through situations we term a ‘surge’. A surge refers to a situation where events leading to exceptional handling occur frequently over a short period of time. For instance, a surge caused by frequent queue insertions can eventually fill the capacity, triggering exceptional handling such as a pipeline stall. We propose a novel fuzzing method for CPU designs, named SurgeFuzz, that intentionally generates surges. SurgeFuzz mutates input instruction sequences based on annotations to increase the occurrence of surges. This results in a higher density of event occurrences, thereby enabling efficient verification of corner cases in exceptional handling. We evaluated SurgeFuzz on a large processor design and found several unknown hardware bugs that are difficult to find with existing methods. Yuichi Sugiyama, Reoma Matsuo, Ryota Shioya |
ICCAD | 3 |
| 2023 | An Out-of-Order Superscalar Processor Using STRAIGHT Architecture in 28 nm CMOSabstractThe single-thread performance of a CPU is an essential factor in a computer system. However, increasing the processing width of a CPU to improve performance often results in a super-linear enlargement of the circuit area and, consequently, a massive increase in power consumption. In this paper, we present an out-of-order superscalar processor based on a new architecture, STRAIGHT, which overcomes the circuit area and power consumption problems. We have designed and evaluated the first real processor chip based on the STRAIGHT architecture. The processor chip was fabricated using 28nm CMOS technology, and we confirmed that it could correctly execute real programs. We evaluated its performance, circuit area, and power consumption, and as a result, demonstrated that a large processing width can be achieved in a small area using the new STRAIGHT architecture. Taichi Amano, Junichiro Kadomoto, Satoshi Mitsuno, Toru Koizumi 0001, Ryota Shioya, Hidetsugu Irie, Shuichi Sakai |
ISCAS | 5 |
| 2023 | Clockhands: Rename-free Instruction Set Architecture for Out-of-order ProcessorsabstractOut-of-order superscalar processors are currently the only architecture that speeds up irregular programs, but they suffer from poor power efficiency. To tackle this issue, we focused on how to specify register operands. Specifying operands by register names, as conventional RISC does, requires register renaming, resulting in poor power efficiency and preventing an increase in the front-end width. In contrast, a recently proposed architecture called STRAIGHT specifies operands by inter-instruction distance, thereby eliminating register renaming. However, STRAIGHT has strong constraints on instruction placement, which generally results in a large increase in the number of instructions. Toru Koizumi 0001, Ryota Shioya, Shu Sugita, Taichi Amano, Yuya Degawa, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai |
MICRO | 2 |
| 2023 | Collecting Cyclic Garbage across Foreign Function Interfaces: Who Takes the Last Piece of Cake?abstractA growing number of libraries written in managed languages, such as Python and JavaScript, are bringing about new demand for a foreign language interface (FFI) between two managed languages. Such an FFI allows a host-language program to seamlessly call a library function written in a foreign language and exchange objects. It is often implemented by a user-level library but such implementation cannot reclaim cyclic garbage, or a group of objects with circular references, across the language boundary. This paper proposes Refgraph GC , which enables FFI implementation that can reclaim cyclic garbage. Refgraph GC coordinates the garbage collectors of two languages and it needs to modify the managed runtime of one language only. It does not modify that of the other language. This paper discusses the soundness and completeness of the proposed algorithm and also shows the results of the experiments with our implementation of FFI with Refgraph GC. This FFI allows a Ruby program to access a JavaScript library. Tetsuro Yamazaki, Tomoki Nakamaru, Ryota Shioya, Tomoharu Ugawa, Shigeru Chiba |
Proc. ACM Program. Lang. | 3 |
| 2022 | T-SKID: Predicting When to Prefetch Separately from Address PredictionabstractPrefetching is an important technique for reducing the number of cache misses and improving processor performance, and thus various prefetchers have been proposed. Many prefetchers are focused on issuing prefetches sufficiently earlier than demand accesses to hide miss latency. In contrast, we propose aT-SKID prefetcher, which focuses on delaying prefetching. If a prefetcher issues prefetches for demand accesses too early, the prefetched line will be evicted before it is referenced. We found that existing prefetchers often issue such too-early prefetches, and this observation offers new opportunities to improve performance. To tackle this issue, T-SKID performs timing prediction indepen-dently of address prediction. In addition to issuing prefetches sufficiently early as existing prefetchers do, T-SKID can delay the issue of prefetches until an appropriate time if necessary. We evaluated T-SKID by simulations using SPEC CPU 2017. The result shows that T-SKID achieves a 5.6 % performance improve-ment for multi-core environment, compared to Instruction Pointer Classifier based Prefetching, which is a state-of-the-art prefetcher. Toru Koizumi 0001, Tomoki Nakamura, Yuya Degawa, Hidetsugu Irie, Shuichi Sakai, Ryota Shioya |
DATE | 6 |
| 2021 | Compiling and Optimizing Real-world Programs for STRAIGHT ISAabstractThe renaming unit of a superscalar processor is a very expensive module. It consumes large amounts of power and limits the front-end bandwidth. To overcome this problem, an instruction set architecture called STRAIGHT has been proposed. Owing to its unique manner of referencing operands, STRAIGHT does not cause false dependencies and allows out-of-order execution without register renaming. However, the compiler optimization techniques for STRAIGHT are still immature, and we found that the naive code generators currently available can generate inefficient code with additional instructions. In this paper, we propose two novel compiler optimization techniques and a novel calling convention for STRAIGHT to reduce the number of instructions. We compiled real-world programs with a compiler that implemented these techniques and measured their performance through simulation. The evaluation results show that the proposed methods reduced the number of executed instructions by 15% and improved the performance by 17%. Toru Koizumi 0001, Shu Sugita, Ryota Shioya, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai |
ICCD | 3 |
| 2021 | Accurate and Fast Performance Modeling of Processors with Decoupled Front-endabstractVarious techniques, such as cache replacement algorithms and prefetching, have been studied to prevent instruction cache misses from becoming a bottleneck in the processor frontend. In such studies, the goal of the design has been to reduce the number of instruction cache misses. However, owing to the increasing complexity of modern processors, the correlation between reducing instruction cache misses and reducing the number of executed cycles has become smaller than in previous cases. In this paper, we propose a new guideline for improving the performance of modern processors. In addition, we propose a method for estimating the approximate performance of a design two orders of magnitude faster than a full simulation each time the designers modify their design. Yuya Degawa, Toru Koizumi 0001, Tomoki Nakamura, Ryota Shioya, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai |
ICCD | 4 |
| 2021 | The Granularity Gap Problem: A Hurdle for Applying Approximate Memory to Complex Data LayoutabstractThe main memory access latency has not much improved for more than two decades while the CPU performance had been exponentially increasing until recently.Approximate memory is a technique to reduce the DRAM access latency in return of losing data integrity. It is expected to be beneficial for applications that are robust to noisy input and intermediate data such as artificial intelligence, image/video processing, and big-data analytics. To obtain reasonable outputs from applications on approximate memory, it is crucial to protect critical data while accelerating accesses to non-critical data. We refer the minimum size of a continuous memory region that the same error rate is applied in approximate memory to as the approximation granularity. A fundamental limitation of approximate memory is that the approximation granularity is as large as a few kilo bytes. However, applications may have critical and non-critical data interleaved with smaller granularity. For example, a data structure for graph nodes can have pointers (critical) to neighboring nodes and its score (non-critical, depending on the use-case). This data structure cannot be directly mapped to approximate memory due to the gap between the approximation granularity and the granularity of data criticality. We refer to this issue as the granularity gap problem. In this paper, we first show that many applications potentially suffer from this problem. Then we propose a framework to quantitatively evaluate the performance overhead of a possible method to avoid this problem using known techniques.The evaluation results show that the performance overhead is non-negligible compared to expected benefit from approximate memory,suggesting that the granularity gap problem is a significant concern. Soramichi Akiyama, Ryota Shioya |
ICPE | 2 |
| 2020 | A High-Performance Out-of-Order Soft Processor Without Register RenamingabstractOwing to the growth of FPGA-based systems and the increasing complexity of applications, the demand for high-performance soft processors in FPGAs has increased. The performance of processors is enhanced through out-of-order (OoO) superscalar execution using a register renaming mechanism. However, the register renaming mechanism has two problems. First, it requires a register mapping table (RMT), which usually comprises a RAM with a large number of ports. A multi-port RAM is not suitable for an FPGA. Second, register renaming complicates recovery mechanisms for exceptions, such as branch mispredictions. These problems increase the usage of resources and hinder the improvement of performance. Recently, the STRAIGHT architecture was proposed to solve these problems. STRAIGHT has a unique instruction format and enables OoO execution without register renaming. This approach eliminates the RMT and makes the recovery operation more efficient. In this study, we demonstrate a high-performance OoO STRAIGHT soft processor by implementing several mechanisms for adopting the STRAIGHT architecture and fabricate the first STRAIGHT processor capable of executing practical complex programs. Compared to a state-of-the-art OoO soft processor, our processor consumes approximately 17% fewer LUTs and 10% fewer FlipFlops and achieves 15% higher performance in CoreMark, which is a standard benchmark. Satoshi Mitsuno, Junichiro Kadomoto, Toru Koizumi 0001, Ryota Shioya, Hidetsugu Irie, Shuichi Sakai |
FPL | 4 |
| 2020 | Energy Efficient Runahead Execution on a Tightly Coupled Heterogeneous CoreabstractOut-of-order (OoO) processors generally offer significant performance gains over simpler in-order (InO) processors. However, recent studies have revealed that OoO processors provide little performance benefit in many program phases, and these phases are distributed in fine granularity. Leveraging these fine-grained phases, tightly coupled heterogeneous cores (TCHCs) have been proposed to improve the energy efficiency. A TCHC, which is a processor core that consists of multiple back-ends, each with different characteristics in terms of their performance and energy consumption (e.g., a power-efficient InO back-end and a high-performance OoO back-end), improves the energy efficiency by executing programs by switching to the most energy-efficient back-end with a very small switching penalty. Susumu Mashimo, Ryota Shioya, Koji Inoue |
HPC Asia | 2 |
| 2018 | An Area-Efficient Out-of-Order Soft-Core Processor Without Register RenamingabstractIn this paper, we present an out-of-order soft-core processor adopting STRAIGHT architecture. STRAIGHT has a unique instruction format in which source operands are expressed as distances from producer instructions. This eliminates the need for register renaming and eliminates a register map table (RMT), which usually consists of a large multi-port RAM. That leads to small area, low power consumption, and high scalability of the front-end pipeline width. Moreover, the simplified architecture enables rapid miss-recovery. The prototype is implemented and evaluated on an FPGA. Compared to an out-of-order soft-core processor with a conventional RISC ISA, the proposed soft-core consumes 147-829 fewer LUTs for the front-end pipeline. The evaluation results show that the proposed soft-core is correctly operating on an FPGA, and estimated dynamic power consumption of the soft-core is 0.120 W. Junichiro Kadomoto, Toru Koizumi 0001, Akifumi Fukuda, Reoma Matsuo, Susumu Mashimo, Akifumi Fujita, Ryota Shioya, Hidetsugu Irie, Shuichi Sakai |
FPT | 7 |
| 2018 | Rearranging Random Issue Queue with High IPC and Short DelayabstractSingle-thread performance has remained mostly static for more than a decade. Among structures in a processor, the issue queue (IQ) is a structure that significantly affects the performance. To achieve high performance, high IPC and a short delay are required for the IQ, which have failed to be achieved in conventional IQs. We propose a novel IQ organization that we call the rearranging random issue queue (RRQ). The RRQ realizes an age-aware instruction selection in the IQ where instructions are ordered randomly. The RRQ divides the IQ into small (OQ: old queue) and large portions (MQ: main queue), where instructions in the OQ are prioritized using a simple select logic. To achieve age-aware selection, a small number of the oldest instructions in the MQ are moved to the OQ every cycle. Our implementation of the RRQ does not complicate the IQ circuit, and hardly increases the delay of the IQ. Evaluation results obtained in architectural simulation show that the RRQ achieves IPC as high as the shifting queue with compaction that realizes the perfect age-aware selection. Our evaluation results also show that the performance of the RRQ significantly outweighs that of processors with an IQ that has the age matrix, which suffers a long delay of the IQ. Shinji Sakai, Taishi Suenaga, Ryota Shioya, Hideki Ando |
ICCD | 3 |
| 2018 | STRAIGHT: Hazardless Processor Architecture Without Register RenamingabstractThe single-thread performance of a processor improves the capability of the entire system by reducing the critical path latency of programs. Typically, conventional superscalar processors improve this performance by introducing out-of-order (OoO) execution with register renaming. However, it is also known to increase the complexity and affect the power efficiency. This paper realizes a novel computer architecture called "STRAIGHT" to resolve this dilemma. The key feature is a unique instruction format in which the source operand is given based on the distance from the producer instruction. By leveraging this format, register renaming is completely removed from the pipeline. This paper presents the practical Instruction Set Architecture (ISA) design, the novel efficient OoO microarchitecture, and the compilation algorithm for the STRAIGHT machine code. Because the ISA has sequential execution semantics, as in general CPUs, and is provided with a compiler, programming for the architecture is as easy as that of conventional CPUs. A compiler, an assembler, a linker, and a cycle-accurate simulator are developed to measure the performance. Moreover, an RTL description of STRAIGHT is developed to estimate the power reduction. The evaluation using standard benchmarks shows that the performance of STRAIGHT is 18.8% better than the conventional superscalar processor of the same issue-width and instruction window size. This improvement is achieved by STRAIGHT's rapid miss-recovery. Compilation technology for resolving the possible overhead of the ISA is also revealed. The RTL power analysis shows that the architecture reduces the power consumption by removing the power for renaming. The revealed performance and efficiencies support that STRAIGHT is a novel viable alternative for designing general purpose OoO processors. Hidetsugu Irie, Toru Koizumi 0001, Akifumi Fukuda, Seiya Akaki, Satoshi Nakae, Yutaro Bessho, Ryota Shioya, Takahiro Notsu, Katsuhiro Yoda, Teruo Ishihara, Shuichi Sakai |
MICRO | 7 |
| 2014 | Energy efficiency improvement of renamed trace cache through the reduction of dependent path lengthabstractA renaming logic is a high-cost module in a superscalar processor, and it consumes significant energy. For mitigating this, renamed trace cache (RTC), which caches renamed operands, was proposed. However, conventional RTCs have several problems such as low capacity-efficiency, large hardware overhead and insufficient caching of renamed operands. We propose a semi-global renamed trace cache (SGRTC) that caches only renamed operands whose distances from producers outside traces are short, and it solves the problems of conventional RTCs. Evaluation results show that SGRTC achieves 64% lower energy consumption for renaming with a 0.2% performance overhead compared to a conventional processor. Ryota Shioya, Hideki Ando |
ICCD | 1 |
| 2014 | A Front-End Execution Architecture for High Energy EfficiencyabstractSmart phones and tablets have recently become widespread and dominant in the computer market. Users require that these mobile devices provide a high-quality experience and an even higher performance. Hence, major developers adopt out-of-order superscalar processors as application processors. However, these processors consume much more energy than in-order superscalar processors, because a large amount of energy is consumed by the hardware for dynamic instruction scheduling. We propose a Front-end Execution Architecture (FXA). FXA has two execution units: an out-of-order execution unit (OXU) and an in-order execution unit (IXU). The OXU is the execution core of a common out-of-order superscalar processor. In contrast, the IXU comprises functional units and a bypass network only. The IXU is placed at the processor front end and executes instructions without scheduling. Fetched instructions are first fed to the IXU, and the instructions that are already ready or become ready to execute by the resolution of their dependencies through operand bypassing in the IXU are executed in-order. Not ready instructions go through the IXU as a NOP, thereby, its pipeline is not stalled, and instructions keep flowing. The not-ready instructions are then dispatched to the OXU, and are executed out-of-order. The IXU does not include dynamic scheduling logic, and its energy consumption is consequently small. Evaluation results show that FXA can execute over 50% of instructions using IXU, thereby making it possible to shrink the energy-consuming OXU without incurring performance degradation. As a result, FXA achieves both a high performance and low energy consumption. We evaluated FXA compared with conventional out-of-order/in-order superscalar processors after ARM big. LITTLE architecture. The results show that FXA achieves performance improvements of 67% at the maximum and 7.4% on geometric mean in SPECCPU INT 2006 benchmark suite relative to a conventional superscalar processor (big), while reducing the energy consumption by 86% at the issue queue and 17% in the whole processor. The performance/energy ratio (the inverse of the energy-delay product) of FXA is 25% higher than that of a conventional superscalar processor (big) and 27% higher than that of a conventional in-order superscalar processor (LITTLE). Ryota Shioya, Masahiro Goshima, Hideki Ando |
MICRO | 1 |
| 2010 | Register Cache System Not for Latency Reduction PurposeabstractA register cache has been proposed to solve the problems of the huge register files of recent super scalar processors. The register cache reduces the effective access latency of the register file for IPC improvement, simplifies the bypass network, and reduces the ports of the main register file. Though the primary purpose of the previous works is to improve IPC, the misses on the register cache may degrade the IPC. We propose Non-Latency-Oriented Register Cache System (NORCS). Though the effects of NORCS are the same as the conventional systems, it is free from register cache miss penalties that the conventional systems suffer from. In NORCS, the register cache itself is not different from that of the conventional systems. The difference is that the instruction pipeline has stages to read the main register file, which all instructions go through regardless of register cache hit / miss. Therefore, the instruction pipeline of NORCS is not immediately disturbed by the register cache misses. For a realistic 4-way super scalar processor, NORCS can simplify the bypass network to the same complexity as a 1-cycle-latency register file, and reduce the ports of the main register file from 12 to 4. CACTI simulation shows that the area and power consumption are reduced to 24.9% and 31.9% compared to the baseline model without register cache. Though these results are not different from the conventional systems, IPCs differ greatly. IPC of the conventional system decreases to 83.1% because of the cache miss penalties, while that of NORCS is retained at 98.0%. Ryota Shioya, Kazuo Horio, Masahiro Goshima, Shuichi Sakai |
MICRO | 1 |
| 2009 | String-Wise Information Flow Tracking against Script Injection AttacksabstractNowadays, security of Web applications faces a threat of script injection attacks. DTP (dynamic taint propagation) and DIFT (dynamic information flow tracking) have been established as powerful techniques to detect script injection attacks. However current DTP/DIFT systems still suffer from tradeoff between false positives and negatives.This paper proposes string-wise information flow tracking, SWIFT. SWIFT traces memory access of program execution, detects string access and distinguishes string operations from other memory access. Current DTP/DIFT systems propagate taint from source to destination operands. Instead of that, SWIFT propagates taint information under string operations. This makes SWIFT provide a better accuracy on detection of script injection attacks than current DTP/DIFT systems.We implemented SWIFT on an IA-32 emulator Bochs, executed typical string operations and made injection attacks to some real-world Web applications with known vulnerabilities. As a result, SWIFT shows a high precision in our security experiments. Kunbo Li, Ryota Shioya, Masahiro Goshima, Shuichi Sakai |
PRDC | 2 |
| 2009 | Low-Overhead Architecture for Security TagabstractA security-tagged architecture is one that applies tags on data to detect attack or information leakage, tracking data flow.The previous studies using security-tagged architecture mostly focused on how to utilize tags, not how the tags are implemented. A naive implementation of tags simply adds a tag field to every byte of the cache and the memory. Such technique, however, results in a huge hardware overhead.This paper proposes a low-overhead tagged architecture. We achieve our goal by exploiting some properties of tag, the non-uniformity and the locality of reference. Our design includes a use of uniquely designed multi-level table and various cache-like structures, all contributing to exploit these properties. Under simulation, our method was able to limit the memory overhead to 1.8%, where a naive implementation suffered 12.5% overhead. Ryota Shioya, Daewung Kim, Kazuo Horio, Masahiro Goshima, Shuichi Sakai |
PRDC | 1 |
| 2006 | Base Address Recognition with Data Flow Tracking for Injection Attack DetectionabstractVulnerabilities such as buffer overflows exist in some programs, and such vulnerabilities are susceptible to address injection attacks. The input data tracking method, which was proposed before, prevents I-data, which are the data derived from the input data, being used as addresses. However, the rules to determine address injection attacks are vague, which produces many false-positives and false-negatives in detection results. Generally, the data used as an address consist of a base address and an address offset. We propose an architectural technique to prevent I-data overwriting B-data, which are the data used as base addresses in this paper. It dynamically recognizes the I-data and the B-data. Address injection is detected if I-data that are not B-data are used as addresses. We implemented the proposed technique on a Pentium-based Bochs emulator and investigated its detection capability. We believe that the technique is the most accurate injection detection technique proposed thus far Satoshi Katsunuma, Hiroyuki Kurita, Ryota Shioya, Kazuto Shimizu, Hidetsugu Irie, Masahiro Goshima, Shuichi Sakai |
PRDC | 3 |