Hidetsugu Irie

dblp:43/4691 · DBLP profile ↗
← Back
29ranked-venue papers
1as first author
19since 2021 · last 2026
0000-0002-5678-2377ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 1 first-author · 12 since 2021Software engineering, systems software and programming languages · 8 · 5 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 CoDA: Constraint-Based Distance Adjustment Optimization for Distance-Based ISAs
Shu Sugita, Masumi Aoki, Junichiro Kadomoto, Hidetsugu Irie
Euro-Par (1)4
2026 Position Estimation Method Using Compact Magnetic Resonant Coupled Coils for Shape-Changing User Interfaces
abstract
This paper presents a position estimation method for shape-changing user interfaces composed of multiple small modules. The method employs compact resonant coils for both communication and position estimation based on received signal strength. When a small coil resonates at a high frequency, the magnetic field exhibits both inductive and radiative characteristics. Our method analyzes the magnetic field through electromagnetic simulation and leverages this relationship for position estimation. The proposed method is robust against occlusion, does not require external cameras or sensors, and enables position estimation solely from communication characteristics. Furthermore, compared to conventional electromagnetic induction-based methods, it supports longer-range communication and position estimation while achieving an average error on the order of millimeters.
Kenta Higuchi, Masato Goto, Junichiro Kadomoto, Hidetsugu Irie
TEI4
2025 Biotite: A High-Performance Static Binary Translator using Source-Level Information
abstract
Research on novel Instruction Set Architectures (ISAs) is actively pursued; however, it requires extensive efforts to develop and maintain comprehensive compilation toolchains for each new ISA. Binary translation can provide a practical solution for ISA researchers to port target programs to novel ISAs when such a level of toolchain support is not available. However, to ensure the correct handling of indirect jumps, existing binary translators rely on complex runtime systems, whose implementation on primitive research ISAs demands significant efforts. ISA researchers generally have access to additional source-level information, including the symbol table and source code, when using binary translators. The symbol table can provide potential jump targets for optimizing indirect jumps, and ISA-independent functions in source code can be directly compiled without translation. Leveraging source-level information as additional input, in this paper, we propose Biotite, a high-performance static binary translator that correctly handles arbitrary indirect jumps. Currently, Biotite supports the translation of RV64GC Linux binaries to self-contained LLVM IR. Our evaluation shows that Biotite successfully translates all benchmarks in SPEC CPU 2017 and achieves a 2.346× performance improvement over QEMU for the integer benchmark suite.
Changbin Chen, Shu Sugita, Yotaro Nada, Hidetsugu Irie, Shuichi Sakai, Ryota Shioya
CC4
2025 Design of an Online Surface Code Decoder Using Union-Find Algorithm
abstract
Real-time Quantum Error Correction (QEC) for Fault-Tolerant Quantum Computing (FTQC) demands immediate surface code decoding with both high accuracy and low latency. The highly accurate Union-Find (UF) algorithm has traditionally been limited to batch-processing, as its iterative cluster growth is incompatible with the concurrent initiation of decoding with each syndrome measurement round, required for real-time QEC. This research presents a novel online UF decoder microarchitecture designed to overcome these limitations. Our key innovation decomposes UF cluster growth into incremental, per-cycle steps managed by a dedicated SHIFT stage. ASIC evaluation of our implementation demonstrates a$\mathbf{1 0 2. 3 8}$ns latency for a surface code with a distance of 23 and an approximate 1.4% threshold, meeting critical QEC performance targets. By demonstrating a viable online UF decoder with competitive performance, this research offers a crucial building block for the practical realization of scalable FTQC.
Takuya Kasamura, Junichiro Kadomoto, Hidetsugu Irie
ICCD3
2025 Register Bridging: A Lightweight Microarchitectural Approach for Skipping Overhead Instructions in Distance-Based ISA Processors
abstract
Out-of-order superscalar processors achieve high performance at the cost of control complexity and energy overhead, with register renaming contributing significantly. Distancebased instruction set architectures (ISAs) provide an alternative to avoid register renaming by specifying operands using relative instruction distances. As a representative design, STRAIGHT implements this approach to support out-of-order execution and eliminate false dependencies in a lightweight design. However, to simplify the overall system design and ensure clear instruction semantics, distance-based architectures require additional instructions (e.g., RMOV) to adjust operand distances, which consume execution resources and, more critically, may delay dependent instructions, resulting in performance degradation. In this paper, we propose Register Bridging, a mechanism that redirects semantically equivalent operands to bypass RMOV dependencies, enabling parallel execution of instructions previously constrained by data-flow ordering. Specifically, a circular buffer is introduced to support operand redirecting with low complexity, in contrast to traditional renaming tables. We implemented the proposed method on a cycle-accurate simulator and compiled benchmarks using the optimizing STRAIGHT compiler. Through a series of simulation experiments on both realistic and synthetic benchmarks, we demonstrate that our proposal enables 41.9% of relay instructions to be bypassed on average, as well as improves performance by up to 5.7%, compared to related methods.
Toru Koizumi 0001, Shu Sugita, Yuriko Yamauchi, Ryota Shioya, Junichiro Kadomoto, Hidetsugu Irie
ICCD8
2025 MorphKeys: A Reconfigurable Keyboard System Using Switched NFC Tags
Koki Yamagami, Masato Goto, Junichiro Kadomoto, Hidetsugu Irie
UIST4
2024 Designing a Reactive Programming Language for Shape-Adaptive Computers
abstract
The advent of shape-adaptive computers, which consist of microscale devices that can wirelessly interconnect and dynamically reconfigure their shapes and functions, presents new challenges for software development. Existing programming environments and languages are not well-suited for managing the complex state interactions and asynchronous processes inherent in such systems. In response to these challenges, we propose the design and current status of MorphLang, a declarative programming language specifically designed for shape-adaptive computers. MorphLang abstracts the complexities of device interaction, enabling developers to focus on high-level behavior definitions without the need for intricate state management. Through a practical example, we demonstrate MorphLang's ability to handle dynamic node interactions effectively, paving the way for more efficient and innovative applications in shape-adaptive computing. Our approach not only simplifies the de-velopment process but also lays the groundwork for future advancements in this emerging field.
Yusuke Izawa, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai
APSEC3
2024 Multi-Tree Network Protocol Enabling System Partitioning for Shape-Changeable Computer System
abstract
Shape-changeable computer system is proposed as a system that forms various shapes by communicating wirelessly with many adjacent chips. A network that supports diverse shapes and dynamic chip replacement is important, and methods for constructing ad-hoc wireless networks and enabling dynamic reconfiguration of the system at runtime have been proposed so far. However, there is room for improvement in performance because they are based on up*/down* routing. Moreover, system partitioning is required by applications such as micro-robots and shape-changing user interfaces, and no network that realizes this has been studied yet. In this study, we propose a multi-tree network construction method for improving routing performance and a protocol that enables system partitioning, and we verify them by simulation using a newly-developed simulator.
Shun Nagasaki, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai
CF3
2023 A Sound and Complete Algorithm for Code Generation in Distance-Based ISA
abstract
The single-thread performance of a processor core is essential even in the multicore era. However, increasing the processing width of a core to improve the single-thread performance leads to a super-linear increase in power consumption. To overcome this power consumption issue, an instruction set architecture for general-purpose processors, called STRAIGHT, has been proposed. STRAIGHT adopts a distance-based ISA, in which source operands are specified by the distance between instructions. In STRAIGHT, it is necessary to satisfy constraints on the distance used as operands to generate executable code. However, it is not yet clear how to generate code that satisfies these constraints in the general case. In this paper, we propose three compiling techniques for STRAIGHT code generation and prove that our techniques can reliably generate code that satisfies the distance constraints. We implemented the proposed method on a compiler and evaluated benchmark programs compiled with it through simulation. The evaluation results showed that the proposed method works in all cases, including conditions where the number of registers is small and existing methods fail to generate code.
Shu Sugita, Toru Koizumi 0001, Ryota Shioya, Hidetsugu Irie, Shuichi Sakai
CC4
2023 TURBULENCE: Complexity-effective Out-of-order Execution on GPU with Distance-based ISA
abstract
A graphic processing unit (GPU) is a processor that achieves high throughput by exploiting data parallelism. We found that many GPU workloads also contain instruction-level parallelism, which can be extracted through out-of-order execution to provide additional performance improvement opportunities. We propose the TURBULENCE architecture for very low-cost out-of-order execution on GPUs. TURBULENCE consists of 1) a novel ISA that introduces the concept of referencing operands by inter-instruction distance instead of register numbers and 2) a novel microarchitecture that executes the novel ISA. Our proposed ISA and microarchitecture enable cost-effective out-of-order execution on GPUs without introducing expensive hardware.
Reoma Matsuo, Toru Koizumi 0001, Hidetsugu Irie, Shuichi Sakai, Ryota Shioya
DATE3
2023 An Out-of-Order Superscalar Processor Using STRAIGHT Architecture in 28 nm CMOS
abstract
The single-thread performance of a CPU is an essential factor in a computer system. However, increasing the processing width of a CPU to improve performance often results in a super-linear enlargement of the circuit area and, consequently, a massive increase in power consumption. In this paper, we present an out-of-order superscalar processor based on a new architecture, STRAIGHT, which overcomes the circuit area and power consumption problems. We have designed and evaluated the first real processor chip based on the STRAIGHT architecture. The processor chip was fabricated using 28nm CMOS technology, and we confirmed that it could correctly execute real programs. We evaluated its performance, circuit area, and power consumption, and as a result, demonstrated that a large processing width can be achieved in a small area using the new STRAIGHT architecture.
Taichi Amano, Junichiro Kadomoto, Satoshi Mitsuno, Toru Koizumi 0001, Ryota Shioya, Hidetsugu Irie, Shuichi Sakai
ISCAS6
2023 Clockhands: Rename-free Instruction Set Architecture for Out-of-order Processors
abstract
Out-of-order superscalar processors are currently the only architecture that speeds up irregular programs, but they suffer from poor power efficiency. To tackle this issue, we focused on how to specify register operands. Specifying operands by register names, as conventional RISC does, requires register renaming, resulting in poor power efficiency and preventing an increase in the front-end width. In contrast, a recently proposed architecture called STRAIGHT specifies operands by inter-instruction distance, thereby eliminating register renaming. However, STRAIGHT has strong constraints on instruction placement, which generally results in a large increase in the number of instructions.
Toru Koizumi 0001, Ryota Shioya, Shu Sugita, Taichi Amano, Yuya Degawa, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai
MICRO7
2023 Poster Abstract: Investigation of Distance Sensing Method Using Magnetic Resonant Coupled Coils for Deformable User Interfaces
abstract
We investigated the possibility of using the output voltage of magnetic resonant coupled coils to realize a distance sensing method that can be applied even when the distance between coils is long. We changed the distance between PCB boards with 1 cm coils and measured the output voltage induced in one coil when an AC voltage of the resonant frequency is input to the other coil. As a result, a correlation was confirmed between the distance between the coils and the value of the output voltage. We found that the distance between the coils can be estimated from the output voltage even when that distance is more than five times the coil diameter. This method enables distance sensing between objects simply by placing a coil on the object. This allows for the sensing of positional relationships between components used in deformable user interfaces.
Kenta Higuchi, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai
SenSys3
2023 Poster Abstract: Towards a Tiny Digital Displacement Sensor Utilizing Bit-Error Characteristics of Inter-Chip Wireless Bus
abstract
In this paper, a displacement sensing method utilizing the bit-error characteristics in inter-chip wireless bus is presented. A non-contact displacement sensor is highly demanded in many engineering fields, and its miniaturization contributes to the expansion of further application areas and simplifies implementation. Inter-chip wireless bus is a short-range wireless communication technology that uses a small coupler, and its bit-error characteristics change according to the relative position between couplers. Additionally, the bit-error rate in inter-chip wireless bus can be obtained solely by digital processing through a tiny microcontroller. Therefore, a tiny digital displacement sensor, integrating a small coupler, wireless transceiver circuits, and a microcontroller, can be realized. The principle of proposed sensing method is verified using simulations, and a 1 mm × 1 mm CMOS LSI test chip is designed and fabricated using 180-nm CMOS technology.
Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai
SenSys2
2022 Deformable Chiplet-Based Computer Using Inductively Coupled Wireless Communication
abstract
Research on microrobot swarms and deformable user interfaces has been conducted extensively. Inductively coupled wireless bus technology has been proposed for such applications. This technology uses inductive coupling among on-chip coils to connect multiple chiplets wirelessly. By wirelessly connecting small chiplets, it is possible to construct deformable systems with various chip configurations. The prototype chip, which has a 32-bit RISC-V processor core and a wireless communication interface, is fabricated in 1.18-µm CMOS technology. The prototype validates that inductively coupled wireless data communication can be achieved between two processor chiplets.
Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai
ASP-DAC2
2022 T-SKID: Predicting When to Prefetch Separately from Address Prediction
abstract
Prefetching is an important technique for reducing the number of cache misses and improving processor performance, and thus various prefetchers have been proposed. Many prefetchers are focused on issuing prefetches sufficiently earlier than demand accesses to hide miss latency. In contrast, we propose aT-SKID prefetcher, which focuses on delaying prefetching. If a prefetcher issues prefetches for demand accesses too early, the prefetched line will be evicted before it is referenced. We found that existing prefetchers often issue such too-early prefetches, and this observation offers new opportunities to improve performance. To tackle this issue, T-SKID performs timing prediction indepen-dently of address prediction. In addition to issuing prefetches sufficiently early as existing prefetchers do, T-SKID can delay the issue of prefetches until an appropriate time if necessary. We evaluated T-SKID by simulations using SPEC CPU 2017. The result shows that T-SKID achieves a 5.6 % performance improve-ment for multi-core environment, compared to Instruction Pointer Classifier based Prefetching, which is a state-of-the-art prefetcher.
Toru Koizumi 0001, Tomoki Nakamura, Yuya Degawa, Hidetsugu Irie, Shuichi Sakai, Ryota Shioya
DATE4
2021 Compiling and Optimizing Real-world Programs for STRAIGHT ISA
abstract
The renaming unit of a superscalar processor is a very expensive module. It consumes large amounts of power and limits the front-end bandwidth. To overcome this problem, an instruction set architecture called STRAIGHT has been proposed. Owing to its unique manner of referencing operands, STRAIGHT does not cause false dependencies and allows out-of-order execution without register renaming. However, the compiler optimization techniques for STRAIGHT are still immature, and we found that the naive code generators currently available can generate inefficient code with additional instructions. In this paper, we propose two novel compiler optimization techniques and a novel calling convention for STRAIGHT to reduce the number of instructions. We compiled real-world programs with a compiler that implemented these techniques and measured their performance through simulation. The evaluation results show that the proposed methods reduced the number of executed instructions by 15% and improved the performance by 17%.
Toru Koizumi 0001, Shu Sugita, Ryota Shioya, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai
ICCD5
2021 Accurate and Fast Performance Modeling of Processors with Decoupled Front-end
abstract
Various techniques, such as cache replacement algorithms and prefetching, have been studied to prevent instruction cache misses from becoming a bottleneck in the processor frontend. In such studies, the goal of the design has been to reduce the number of instruction cache misses. However, owing to the increasing complexity of modern processors, the correlation between reducing instruction cache misses and reducing the number of executed cycles has become smaller than in previous cases. In this paper, we propose a new guideline for improving the performance of modern processors. In addition, we propose a method for estimating the approximate performance of a design two orders of magnitude faster than a full simulation each time the designers modify their design.
Yuya Degawa, Toru Koizumi 0001, Tomoki Nakamura, Ryota Shioya, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai
ICCD6
2021 Stochastic Iterative Approximation: Software/hardware techniques for adjusting aggressiveness of approximation
abstract
Approximate computing (AC) reduces power consumption and increases execution speed in exchange for computational accuracy. By adjusting the accuracy of approximation at runtime to reflect the optimal quality of the application, which changes constantly depending on the user’s cognitive ability and attention, AC achieves even higher efficiency. In this paper, we propose stochastic iterative approximation (SIA) that achieves dynamic and rapid control of the aggressiveness of the approximation. SIA executes a single binary code with multiple level of approximate aggressiveness that are dynamically adjusted. We propose a software implementation of SIA and hardware techniques to further improve the performance of SIA. We implement a compiler and a processor simulator for SIA as the dynamic approximation modules of RISC-V and evaluate their performance. Simulation results on six benchmarks show an adjustable trade-off between output quality and execution efficiency depending on the aggressiveness of the approximation in a single binary run.
Tomoki Nakamura, Kazutaka Tomida, Shouta Kouno, Hidetsugu Irie, Shuichi Sakai
ICCD4
2020 An Inductively Coupled Wireless Bus for Chiplet-Based Systems
abstract
A wireless bus for inter-chiplet communication is presented. Utilizing horizontal inductive coupling of on-chip coils, wireless connection between chiplets are established. A test chip prototyped in 0.18 μm CMOS confirms 2.0 Gb/s bus communication between horizontally arranged coils with BER of less than 10-12.
Junichiro Kadomoto, Satoshi Mitsuno, Hidetsugu Irie, Shuichi Sakai
ASP-DAC3
2020 A High-Performance Out-of-Order Soft Processor Without Register Renaming
abstract
Owing to the growth of FPGA-based systems and the increasing complexity of applications, the demand for high-performance soft processors in FPGAs has increased. The performance of processors is enhanced through out-of-order (OoO) superscalar execution using a register renaming mechanism. However, the register renaming mechanism has two problems. First, it requires a register mapping table (RMT), which usually comprises a RAM with a large number of ports. A multi-port RAM is not suitable for an FPGA. Second, register renaming complicates recovery mechanisms for exceptions, such as branch mispredictions. These problems increase the usage of resources and hinder the improvement of performance. Recently, the STRAIGHT architecture was proposed to solve these problems. STRAIGHT has a unique instruction format and enables OoO execution without register renaming. This approach eliminates the RMT and makes the recovery operation more efficient. In this study, we demonstrate a high-performance OoO STRAIGHT soft processor by implementing several mechanisms for adopting the STRAIGHT architecture and fabricate the first STRAIGHT processor capable of executing practical complex programs. Compared to a state-of-the-art OoO soft processor, our processor consumes approximately 17% fewer LUTs and 10% fewer FlipFlops and achieves 15% higher performance in CoreMark, which is a standard benchmark.
Satoshi Mitsuno, Junichiro Kadomoto, Toru Koizumi 0001, Ryota Shioya, Hidetsugu Irie, Shuichi Sakai
FPL5
2020 Design of Shape-Changeable Chiplet-Based Computers Using an Inductively Coupled Wireless Bus Interface
abstract
Research on small-sized microrobot swarms and shape-changeable user interfaces has been conducted extensively. Wireless bus interface technology has been proposed for such applications. This technology uses inductive coupling among on-chip coils to connect multiple chips wirelessly. However, wireless bus technology has a peculiar characteristic of broadcasting data only to the chips arranged adjacently, and it is challenging to apply existing network protocols. In addition, interference with processors and peripheral circuits on the same chip has not been thoroughly investigated. In this study, the network architecture of a shape-changeable computer system utilizing the wireless bus interface is presented. Moreover, we show the measurement results of the first multi-chip processor prototype, which utilizes the wireless bus interface. The implementation method of a processor core and the interface, and an interchip network protocol from the physical layer to the network layer are shown. A deadlock-free routing path can be formed in an irregular network among multiple chips by the proposed protocol that takes into account the characteristics of the physical layer of the wireless bus interface. The prototype chip, which has a RISC-V processor core, and the wireless bus interface fabricated in 0.18-μm CMOS technology validates that wireless data communication can be achieved between two processor chips.
Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai
ICCD2
2019 WiXI: An Inter-Chip Wireless Bus Interface for Shape-Changeable Chiplet-Based Computers
abstract
Herein, we propose a wireless bus interface that can connect multiple chips to form a flexible system. In the proposed bus interface, on-chip coils are formed along the outer periphery of each chip, and high-speed wireless communication between multiple chips is enabled via horizontal inductive coupling between coils. Using the proposed interface, embedded computer systems can be realized by simply combining small chips with different functions as needed and arranging them in an adjacent manner. The proposed bus interface enables variation in the relative angle between adjacent chips during operation, can be implemented in complicated shapes, and facilitates chip replacement post fabrication to achieve flexible and robust computer systems; such systems can be applied in micro-robots and wearable interfaces. In this paper, we present a theoretical analysis of electromagnetic coupling between coils, electromagnetic field simulation results, and circuit simulation results of transmitter and receiver circuits. Through the simulation conducted using 45 nm CMOS technology, we realized high-speed communication of 14.3 Gb/s with a power efficiency of 0.55 pJ/b using the proposed bus interface. We also verified the data collision detection ability of the proposed bus interface using a collision detection circuit and packet transfer based on a SerDes circuit.
Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai
ICCD2
2018 An Area-Efficient Out-of-Order Soft-Core Processor Without Register Renaming
abstract
In this paper, we present an out-of-order soft-core processor adopting STRAIGHT architecture. STRAIGHT has a unique instruction format in which source operands are expressed as distances from producer instructions. This eliminates the need for register renaming and eliminates a register map table (RMT), which usually consists of a large multi-port RAM. That leads to small area, low power consumption, and high scalability of the front-end pipeline width. Moreover, the simplified architecture enables rapid miss-recovery. The prototype is implemented and evaluated on an FPGA. Compared to an out-of-order soft-core processor with a conventional RISC ISA, the proposed soft-core consumes 147-829 fewer LUTs for the front-end pipeline. The evaluation results show that the proposed soft-core is correctly operating on an FPGA, and estimated dynamic power consumption of the soft-core is 0.120 W.
Junichiro Kadomoto, Toru Koizumi 0001, Akifumi Fukuda, Reoma Matsuo, Susumu Mashimo, Akifumi Fujita, Ryota Shioya, Hidetsugu Irie, Shuichi Sakai
FPT8
2018 STRAIGHT: Hazardless Processor Architecture Without Register Renaming
abstract
The single-thread performance of a processor improves the capability of the entire system by reducing the critical path latency of programs. Typically, conventional superscalar processors improve this performance by introducing out-of-order (OoO) execution with register renaming. However, it is also known to increase the complexity and affect the power efficiency. This paper realizes a novel computer architecture called "STRAIGHT" to resolve this dilemma. The key feature is a unique instruction format in which the source operand is given based on the distance from the producer instruction. By leveraging this format, register renaming is completely removed from the pipeline. This paper presents the practical Instruction Set Architecture (ISA) design, the novel efficient OoO microarchitecture, and the compilation algorithm for the STRAIGHT machine code. Because the ISA has sequential execution semantics, as in general CPUs, and is provided with a compiler, programming for the architecture is as easy as that of conventional CPUs. A compiler, an assembler, a linker, and a cycle-accurate simulator are developed to measure the performance. Moreover, an RTL description of STRAIGHT is developed to estimate the power reduction. The evaluation using standard benchmarks shows that the performance of STRAIGHT is 18.8% better than the conventional superscalar processor of the same issue-width and instruction window size. This improvement is achieved by STRAIGHT's rapid miss-recovery. Compilation technology for resolving the possible overhead of the ISA is also revealed. The RTL power analysis shows that the architecture reduces the power consumption by removing the power for renaming. The revealed performance and efficiencies support that STRAIGHT is a novel viable alternative for designing general purpose OoO processors.
Hidetsugu Irie, Toru Koizumi 0001, Akifumi Fukuda, Seiya Akaki, Satoshi Nakae, Yutaro Bessho, Ryota Shioya, Takahiro Notsu, Katsuhiro Yoda, Teruo Ishihara, Shuichi Sakai
MICRO1
2017 Accelerating Integrity Verification on Secure Processors by Promissory Hash
abstract
Most digital content nowadays is protected by a digital rights management (DRM) framework to prevent piracy. Since much content is distributed throughout the world, modern DRM frameworks must protect the confidentiality and integrity of the data, even from rootkits or physical tampering. Secure processors have been proposed to ensure secure executions by performing memory encryption and integrity verification effectively. However, application start-up takes a considerable time because it requires initial memory hash calculation. In this paper, we propose Promissory Hash, a method to skip the initial hash calculation while maintaining the actual integrity. The processor declares Promissory Hash to the external certifier that should represent the initial memory. While the application is executing, any disagreement from the Promissory Hash halts the execution. A detailed implementation is revealed and additional storage requirement is also estimated: an additional 1,028 KB for the 256 MB main memory greatly reduce the start-up latency.
Mizuki Miyanaga, Hidetsugu Irie, Shuichi Sakai
PRDC2
2016 "Stubborn" strategy to mitigate remaining cache misses
abstract
The capacity of cache memory has reached the megabytes magnitude and various smart cache algorithms have been introduced. However, cache systems still suffer from plenty of cache misses, especially for memory-intensive applications. In this paper, we reveal that a vast majority of the remaining cache misses on LLC are caused by long interval, unpredictable re-reference accesses. Based on the analysis, we propose a “Stubborn” strategy to mitigate misses caused by long re-reference interval accesses. Our strategy simply enhances existing smart algorithms by mixing the cache lines that do not evict their holding content for longer than a 10 to 100 mega-instruction interval. We evaluate the strategy that is mixed with 2MB LLC based on LRU or DRRIP. The results show that our Stubborn strategy achieves a reduction in long interval re-reference misses, as expected. It outperforms LRU and DRRIP on the IPC metric showing increases of 24% and 6%, respectively.
Hayato Nomura, Hiroyuki Katchi, Hidetsugu Irie, Shuichi Sakai
ICCD3
2007 Utilization of SECDED for soft error and variation-induced defect tolerance in caches
Luong Dinh Hung, Hidetsugu Irie, Masahiro Goshima, Shuichi Sakai
DATE2
2006 Base Address Recognition with Data Flow Tracking for Injection Attack Detection
abstract
Vulnerabilities such as buffer overflows exist in some programs, and such vulnerabilities are susceptible to address injection attacks. The input data tracking method, which was proposed before, prevents I-data, which are the data derived from the input data, being used as addresses. However, the rules to determine address injection attacks are vague, which produces many false-positives and false-negatives in detection results. Generally, the data used as an address consist of a base address and an address offset. We propose an architectural technique to prevent I-data overwriting B-data, which are the data used as base addresses in this paper. It dynamically recognizes the I-data and the B-data. Address injection is detected if I-data that are not B-data are used as addresses. We implemented the proposed technique on a Pentium-based Bochs emulator and investigated its detection capability. We believe that the technique is the most accurate injection detection technique proposed thus far
Satoshi Katsunuma, Hiroyuki Kurita, Ryota Shioya, Kazuto Shimizu, Hidetsugu Irie, Masahiro Goshima, Shuichi Sakai
PRDC5