VLDB 2026 Research / reviewers in the wild / expert
Shuichi Sakai
dblp:56/2263
· DBLP profile ↗
59ranked-venue papers
8as first author
14since 2021 · last 2025
0000-0002-5471-9091ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 7 first-author · 9 since 2021Software engineering, systems software and programming languages · 14 · 1 first-author · 5 since 2021Security and privacy · 6Graphics, computer vision, multimedia, augmented reality and games · 3Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Biotite: A High-Performance Static Binary Translator using Source-Level InformationabstractResearch on novel Instruction Set Architectures (ISAs) is actively pursued; however, it requires extensive efforts to develop and maintain comprehensive compilation toolchains for each new ISA. Binary translation can provide a practical solution for ISA researchers to port target programs to novel ISAs when such a level of toolchain support is not available. However, to ensure the correct handling of indirect jumps, existing binary translators rely on complex runtime systems, whose implementation on primitive research ISAs demands significant efforts. ISA researchers generally have access to additional source-level information, including the symbol table and source code, when using binary translators. The symbol table can provide potential jump targets for optimizing indirect jumps, and ISA-independent functions in source code can be directly compiled without translation. Leveraging source-level information as additional input, in this paper, we propose Biotite, a high-performance static binary translator that correctly handles arbitrary indirect jumps. Currently, Biotite supports the translation of RV64GC Linux binaries to self-contained LLVM IR. Our evaluation shows that Biotite successfully translates all benchmarks in SPEC CPU 2017 and achieves a 2.346× performance improvement over QEMU for the integer benchmark suite. Changbin Chen, Shu Sugita, Yotaro Nada, Hidetsugu Irie, Shuichi Sakai, Ryota Shioya |
CC | 5 |
| 2024 | Designing a Reactive Programming Language for Shape-Adaptive ComputersabstractThe advent of shape-adaptive computers, which consist of microscale devices that can wirelessly interconnect and dynamically reconfigure their shapes and functions, presents new challenges for software development. Existing programming environments and languages are not well-suited for managing the complex state interactions and asynchronous processes inherent in such systems. In response to these challenges, we propose the design and current status of MorphLang, a declarative programming language specifically designed for shape-adaptive computers. MorphLang abstracts the complexities of device interaction, enabling developers to focus on high-level behavior definitions without the need for intricate state management. Through a practical example, we demonstrate MorphLang's ability to handle dynamic node interactions effectively, paving the way for more efficient and innovative applications in shape-adaptive computing. Our approach not only simplifies the de-velopment process but also lays the groundwork for future advancements in this emerging field. Yusuke Izawa, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai |
APSEC | 4 |
| 2024 | Multi-Tree Network Protocol Enabling System Partitioning for Shape-Changeable Computer SystemabstractShape-changeable computer system is proposed as a system that forms various shapes by communicating wirelessly with many adjacent chips. A network that supports diverse shapes and dynamic chip replacement is important, and methods for constructing ad-hoc wireless networks and enabling dynamic reconfiguration of the system at runtime have been proposed so far. However, there is room for improvement in performance because they are based on up*/down* routing. Moreover, system partitioning is required by applications such as micro-robots and shape-changing user interfaces, and no network that realizes this has been studied yet. In this study, we propose a multi-tree network construction method for improving routing performance and a protocol that enables system partitioning, and we verify them by simulation using a newly-developed simulator. Shun Nagasaki, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai |
CF | 4 |
| 2023 | A Sound and Complete Algorithm for Code Generation in Distance-Based ISAabstractThe single-thread performance of a processor core is essential even in the multicore era. However, increasing the processing width of a core to improve the single-thread performance leads to a super-linear increase in power consumption. To overcome this power consumption issue, an instruction set architecture for general-purpose processors, called STRAIGHT, has been proposed. STRAIGHT adopts a distance-based ISA, in which source operands are specified by the distance between instructions. In STRAIGHT, it is necessary to satisfy constraints on the distance used as operands to generate executable code. However, it is not yet clear how to generate code that satisfies these constraints in the general case. In this paper, we propose three compiling techniques for STRAIGHT code generation and prove that our techniques can reliably generate code that satisfies the distance constraints. We implemented the proposed method on a compiler and evaluated benchmark programs compiled with it through simulation. The evaluation results showed that the proposed method works in all cases, including conditions where the number of registers is small and existing methods fail to generate code. Shu Sugita, Toru Koizumi 0001, Ryota Shioya, Hidetsugu Irie, Shuichi Sakai |
CC | 5 |
| 2023 | TURBULENCE: Complexity-effective Out-of-order Execution on GPU with Distance-based ISAabstractA graphic processing unit (GPU) is a processor that achieves high throughput by exploiting data parallelism. We found that many GPU workloads also contain instruction-level parallelism, which can be extracted through out-of-order execution to provide additional performance improvement opportunities. We propose the TURBULENCE architecture for very low-cost out-of-order execution on GPUs. TURBULENCE consists of 1) a novel ISA that introduces the concept of referencing operands by inter-instruction distance instead of register numbers and 2) a novel microarchitecture that executes the novel ISA. Our proposed ISA and microarchitecture enable cost-effective out-of-order execution on GPUs without introducing expensive hardware. Reoma Matsuo, Toru Koizumi 0001, Hidetsugu Irie, Shuichi Sakai, Ryota Shioya |
DATE | 4 |
| 2023 | An Out-of-Order Superscalar Processor Using STRAIGHT Architecture in 28 nm CMOSabstractThe single-thread performance of a CPU is an essential factor in a computer system. However, increasing the processing width of a CPU to improve performance often results in a super-linear enlargement of the circuit area and, consequently, a massive increase in power consumption. In this paper, we present an out-of-order superscalar processor based on a new architecture, STRAIGHT, which overcomes the circuit area and power consumption problems. We have designed and evaluated the first real processor chip based on the STRAIGHT architecture. The processor chip was fabricated using 28nm CMOS technology, and we confirmed that it could correctly execute real programs. We evaluated its performance, circuit area, and power consumption, and as a result, demonstrated that a large processing width can be achieved in a small area using the new STRAIGHT architecture. Taichi Amano, Junichiro Kadomoto, Satoshi Mitsuno, Toru Koizumi 0001, Ryota Shioya, Hidetsugu Irie, Shuichi Sakai |
ISCAS | 7 |
| 2023 | Clockhands: Rename-free Instruction Set Architecture for Out-of-order ProcessorsabstractOut-of-order superscalar processors are currently the only architecture that speeds up irregular programs, but they suffer from poor power efficiency. To tackle this issue, we focused on how to specify register operands. Specifying operands by register names, as conventional RISC does, requires register renaming, resulting in poor power efficiency and preventing an increase in the front-end width. In contrast, a recently proposed architecture called STRAIGHT specifies operands by inter-instruction distance, thereby eliminating register renaming. However, STRAIGHT has strong constraints on instruction placement, which generally results in a large increase in the number of instructions. Toru Koizumi 0001, Ryota Shioya, Shu Sugita, Taichi Amano, Yuya Degawa, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai |
MICRO | 8 |
| 2023 | Poster Abstract: Investigation of Distance Sensing Method Using Magnetic Resonant Coupled Coils for Deformable User InterfacesabstractWe investigated the possibility of using the output voltage of magnetic resonant coupled coils to realize a distance sensing method that can be applied even when the distance between coils is long. We changed the distance between PCB boards with 1 cm coils and measured the output voltage induced in one coil when an AC voltage of the resonant frequency is input to the other coil. As a result, a correlation was confirmed between the distance between the coils and the value of the output voltage. We found that the distance between the coils can be estimated from the output voltage even when that distance is more than five times the coil diameter. This method enables distance sensing between objects simply by placing a coil on the object. This allows for the sensing of positional relationships between components used in deformable user interfaces. Kenta Higuchi, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai |
SenSys | 4 |
| 2023 | Poster Abstract: Towards a Tiny Digital Displacement Sensor Utilizing Bit-Error Characteristics of Inter-Chip Wireless BusabstractIn this paper, a displacement sensing method utilizing the bit-error characteristics in inter-chip wireless bus is presented. A non-contact displacement sensor is highly demanded in many engineering fields, and its miniaturization contributes to the expansion of further application areas and simplifies implementation. Inter-chip wireless bus is a short-range wireless communication technology that uses a small coupler, and its bit-error characteristics change according to the relative position between couplers. Additionally, the bit-error rate in inter-chip wireless bus can be obtained solely by digital processing through a tiny microcontroller. Therefore, a tiny digital displacement sensor, integrating a small coupler, wireless transceiver circuits, and a microcontroller, can be realized. The principle of proposed sensing method is verified using simulations, and a 1 mm × 1 mm CMOS LSI test chip is designed and fabricated using 180-nm CMOS technology. Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai |
SenSys | 3 |
| 2022 | Deformable Chiplet-Based Computer Using Inductively Coupled Wireless CommunicationabstractResearch on microrobot swarms and deformable user interfaces has been conducted extensively. Inductively coupled wireless bus technology has been proposed for such applications. This technology uses inductive coupling among on-chip coils to connect multiple chiplets wirelessly. By wirelessly connecting small chiplets, it is possible to construct deformable systems with various chip configurations. The prototype chip, which has a 32-bit RISC-V processor core and a wireless communication interface, is fabricated in 1.18-µm CMOS technology. The prototype validates that inductively coupled wireless data communication can be achieved between two processor chiplets. Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai |
ASP-DAC | 3 |
| 2022 | T-SKID: Predicting When to Prefetch Separately from Address PredictionabstractPrefetching is an important technique for reducing the number of cache misses and improving processor performance, and thus various prefetchers have been proposed. Many prefetchers are focused on issuing prefetches sufficiently earlier than demand accesses to hide miss latency. In contrast, we propose aT-SKID prefetcher, which focuses on delaying prefetching. If a prefetcher issues prefetches for demand accesses too early, the prefetched line will be evicted before it is referenced. We found that existing prefetchers often issue such too-early prefetches, and this observation offers new opportunities to improve performance. To tackle this issue, T-SKID performs timing prediction indepen-dently of address prediction. In addition to issuing prefetches sufficiently early as existing prefetchers do, T-SKID can delay the issue of prefetches until an appropriate time if necessary. We evaluated T-SKID by simulations using SPEC CPU 2017. The result shows that T-SKID achieves a 5.6 % performance improve-ment for multi-core environment, compared to Instruction Pointer Classifier based Prefetching, which is a state-of-the-art prefetcher. Toru Koizumi 0001, Tomoki Nakamura, Yuya Degawa, Hidetsugu Irie, Shuichi Sakai, Ryota Shioya |
DATE | 5 |
| 2021 | Compiling and Optimizing Real-world Programs for STRAIGHT ISAabstractThe renaming unit of a superscalar processor is a very expensive module. It consumes large amounts of power and limits the front-end bandwidth. To overcome this problem, an instruction set architecture called STRAIGHT has been proposed. Owing to its unique manner of referencing operands, STRAIGHT does not cause false dependencies and allows out-of-order execution without register renaming. However, the compiler optimization techniques for STRAIGHT are still immature, and we found that the naive code generators currently available can generate inefficient code with additional instructions. In this paper, we propose two novel compiler optimization techniques and a novel calling convention for STRAIGHT to reduce the number of instructions. We compiled real-world programs with a compiler that implemented these techniques and measured their performance through simulation. The evaluation results show that the proposed methods reduced the number of executed instructions by 15% and improved the performance by 17%. Toru Koizumi 0001, Shu Sugita, Ryota Shioya, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai |
ICCD | 6 |
| 2021 | Accurate and Fast Performance Modeling of Processors with Decoupled Front-endabstractVarious techniques, such as cache replacement algorithms and prefetching, have been studied to prevent instruction cache misses from becoming a bottleneck in the processor frontend. In such studies, the goal of the design has been to reduce the number of instruction cache misses. However, owing to the increasing complexity of modern processors, the correlation between reducing instruction cache misses and reducing the number of executed cycles has become smaller than in previous cases. In this paper, we propose a new guideline for improving the performance of modern processors. In addition, we propose a method for estimating the approximate performance of a design two orders of magnitude faster than a full simulation each time the designers modify their design. Yuya Degawa, Toru Koizumi 0001, Tomoki Nakamura, Ryota Shioya, Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai |
ICCD | 7 |
| 2021 | Stochastic Iterative Approximation: Software/hardware techniques for adjusting aggressiveness of approximationabstractApproximate computing (AC) reduces power consumption and increases execution speed in exchange for computational accuracy. By adjusting the accuracy of approximation at runtime to reflect the optimal quality of the application, which changes constantly depending on the user’s cognitive ability and attention, AC achieves even higher efficiency. In this paper, we propose stochastic iterative approximation (SIA) that achieves dynamic and rapid control of the aggressiveness of the approximation. SIA executes a single binary code with multiple level of approximate aggressiveness that are dynamically adjusted. We propose a software implementation of SIA and hardware techniques to further improve the performance of SIA. We implement a compiler and a processor simulator for SIA as the dynamic approximation modules of RISC-V and evaluate their performance. Simulation results on six benchmarks show an adjustable trade-off between output quality and execution efficiency depending on the aggressiveness of the approximation in a single binary run. Tomoki Nakamura, Kazutaka Tomida, Shouta Kouno, Hidetsugu Irie, Shuichi Sakai |
ICCD | 5 |
| 2020 | An Inductively Coupled Wireless Bus for Chiplet-Based SystemsabstractA wireless bus for inter-chiplet communication is presented. Utilizing horizontal inductive coupling of on-chip coils, wireless connection between chiplets are established. A test chip prototyped in 0.18 μm CMOS confirms 2.0 Gb/s bus communication between horizontally arranged coils with BER of less than 10-12. Junichiro Kadomoto, Satoshi Mitsuno, Hidetsugu Irie, Shuichi Sakai |
ASP-DAC | 4 |
| 2020 | A High-Performance Out-of-Order Soft Processor Without Register RenamingabstractOwing to the growth of FPGA-based systems and the increasing complexity of applications, the demand for high-performance soft processors in FPGAs has increased. The performance of processors is enhanced through out-of-order (OoO) superscalar execution using a register renaming mechanism. However, the register renaming mechanism has two problems. First, it requires a register mapping table (RMT), which usually comprises a RAM with a large number of ports. A multi-port RAM is not suitable for an FPGA. Second, register renaming complicates recovery mechanisms for exceptions, such as branch mispredictions. These problems increase the usage of resources and hinder the improvement of performance. Recently, the STRAIGHT architecture was proposed to solve these problems. STRAIGHT has a unique instruction format and enables OoO execution without register renaming. This approach eliminates the RMT and makes the recovery operation more efficient. In this study, we demonstrate a high-performance OoO STRAIGHT soft processor by implementing several mechanisms for adopting the STRAIGHT architecture and fabricate the first STRAIGHT processor capable of executing practical complex programs. Compared to a state-of-the-art OoO soft processor, our processor consumes approximately 17% fewer LUTs and 10% fewer FlipFlops and achieves 15% higher performance in CoreMark, which is a standard benchmark. Satoshi Mitsuno, Junichiro Kadomoto, Toru Koizumi 0001, Ryota Shioya, Hidetsugu Irie, Shuichi Sakai |
FPL | 6 |
| 2020 | Design of Shape-Changeable Chiplet-Based Computers Using an Inductively Coupled Wireless Bus InterfaceabstractResearch on small-sized microrobot swarms and shape-changeable user interfaces has been conducted extensively. Wireless bus interface technology has been proposed for such applications. This technology uses inductive coupling among on-chip coils to connect multiple chips wirelessly. However, wireless bus technology has a peculiar characteristic of broadcasting data only to the chips arranged adjacently, and it is challenging to apply existing network protocols. In addition, interference with processors and peripheral circuits on the same chip has not been thoroughly investigated. In this study, the network architecture of a shape-changeable computer system utilizing the wireless bus interface is presented. Moreover, we show the measurement results of the first multi-chip processor prototype, which utilizes the wireless bus interface. The implementation method of a processor core and the interface, and an interchip network protocol from the physical layer to the network layer are shown. A deadlock-free routing path can be formed in an irregular network among multiple chips by the proposed protocol that takes into account the characteristics of the physical layer of the wireless bus interface. The prototype chip, which has a RISC-V processor core, and the wireless bus interface fabricated in 0.18-μm CMOS technology validates that wireless data communication can be achieved between two processor chips. Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai |
ICCD | 3 |
| 2019 | WiXI: An Inter-Chip Wireless Bus Interface for Shape-Changeable Chiplet-Based ComputersabstractHerein, we propose a wireless bus interface that can connect multiple chips to form a flexible system. In the proposed bus interface, on-chip coils are formed along the outer periphery of each chip, and high-speed wireless communication between multiple chips is enabled via horizontal inductive coupling between coils. Using the proposed interface, embedded computer systems can be realized by simply combining small chips with different functions as needed and arranging them in an adjacent manner. The proposed bus interface enables variation in the relative angle between adjacent chips during operation, can be implemented in complicated shapes, and facilitates chip replacement post fabrication to achieve flexible and robust computer systems; such systems can be applied in micro-robots and wearable interfaces. In this paper, we present a theoretical analysis of electromagnetic coupling between coils, electromagnetic field simulation results, and circuit simulation results of transmitter and receiver circuits. Through the simulation conducted using 45 nm CMOS technology, we realized high-speed communication of 14.3 Gb/s with a power efficiency of 0.55 pJ/b using the proposed bus interface. We also verified the data collision detection ability of the proposed bus interface using a collision detection circuit and packet transfer based on a SerDes circuit. Junichiro Kadomoto, Hidetsugu Irie, Shuichi Sakai |
ICCD | 3 |
| 2018 | An Area-Efficient Out-of-Order Soft-Core Processor Without Register RenamingabstractIn this paper, we present an out-of-order soft-core processor adopting STRAIGHT architecture. STRAIGHT has a unique instruction format in which source operands are expressed as distances from producer instructions. This eliminates the need for register renaming and eliminates a register map table (RMT), which usually consists of a large multi-port RAM. That leads to small area, low power consumption, and high scalability of the front-end pipeline width. Moreover, the simplified architecture enables rapid miss-recovery. The prototype is implemented and evaluated on an FPGA. Compared to an out-of-order soft-core processor with a conventional RISC ISA, the proposed soft-core consumes 147-829 fewer LUTs for the front-end pipeline. The evaluation results show that the proposed soft-core is correctly operating on an FPGA, and estimated dynamic power consumption of the soft-core is 0.120 W. Junichiro Kadomoto, Toru Koizumi 0001, Akifumi Fukuda, Reoma Matsuo, Susumu Mashimo, Akifumi Fujita, Ryota Shioya, Hidetsugu Irie, Shuichi Sakai |
FPT | 9 |
| 2018 | STRAIGHT: Hazardless Processor Architecture Without Register RenamingabstractThe single-thread performance of a processor improves the capability of the entire system by reducing the critical path latency of programs. Typically, conventional superscalar processors improve this performance by introducing out-of-order (OoO) execution with register renaming. However, it is also known to increase the complexity and affect the power efficiency. This paper realizes a novel computer architecture called "STRAIGHT" to resolve this dilemma. The key feature is a unique instruction format in which the source operand is given based on the distance from the producer instruction. By leveraging this format, register renaming is completely removed from the pipeline. This paper presents the practical Instruction Set Architecture (ISA) design, the novel efficient OoO microarchitecture, and the compilation algorithm for the STRAIGHT machine code. Because the ISA has sequential execution semantics, as in general CPUs, and is provided with a compiler, programming for the architecture is as easy as that of conventional CPUs. A compiler, an assembler, a linker, and a cycle-accurate simulator are developed to measure the performance. Moreover, an RTL description of STRAIGHT is developed to estimate the power reduction. The evaluation using standard benchmarks shows that the performance of STRAIGHT is 18.8% better than the conventional superscalar processor of the same issue-width and instruction window size. This improvement is achieved by STRAIGHT's rapid miss-recovery. Compilation technology for resolving the possible overhead of the ISA is also revealed. The RTL power analysis shows that the architecture reduces the power consumption by removing the power for renaming. The revealed performance and efficiencies support that STRAIGHT is a novel viable alternative for designing general purpose OoO processors. Hidetsugu Irie, Toru Koizumi 0001, Akifumi Fukuda, Seiya Akaki, Satoshi Nakae, Yutaro Bessho, Ryota Shioya, Takahiro Notsu, Katsuhiro Yoda, Teruo Ishihara, Shuichi Sakai |
MICRO | 11 |
| 2017 | Accelerating Integrity Verification on Secure Processors by Promissory HashabstractMost digital content nowadays is protected by a digital rights management (DRM) framework to prevent piracy. Since much content is distributed throughout the world, modern DRM frameworks must protect the confidentiality and integrity of the data, even from rootkits or physical tampering. Secure processors have been proposed to ensure secure executions by performing memory encryption and integrity verification effectively. However, application start-up takes a considerable time because it requires initial memory hash calculation. In this paper, we propose Promissory Hash, a method to skip the initial hash calculation while maintaining the actual integrity. The processor declares Promissory Hash to the external certifier that should represent the initial memory. While the application is executing, any disagreement from the Promissory Hash halts the execution. A detailed implementation is revealed and additional storage requirement is also estimated: an additional 1,028 KB for the 256 MB main memory greatly reduce the start-up latency. Mizuki Miyanaga, Hidetsugu Irie, Shuichi Sakai |
PRDC | 3 |
| 2016 | "Stubborn" strategy to mitigate remaining cache missesabstractThe capacity of cache memory has reached the megabytes magnitude and various smart cache algorithms have been introduced. However, cache systems still suffer from plenty of cache misses, especially for memory-intensive applications. In this paper, we reveal that a vast majority of the remaining cache misses on LLC are caused by long interval, unpredictable re-reference accesses. Based on the analysis, we propose a “Stubborn” strategy to mitigate misses caused by long re-reference interval accesses. Our strategy simply enhances existing smart algorithms by mixing the cache lines that do not evict their holding content for longer than a 10 to 100 mega-instruction interval. We evaluate the strategy that is mixed with 2MB LLC based on LRU or DRRIP. The results show that our Stubborn strategy achieves a reduction in long interval re-reference misses, as expected. It outperforms LRU and DRRIP on the IPC metric showing increases of 24% and 6%, respectively. Hayato Nomura, Hiroyuki Katchi, Hidetsugu Irie, Shuichi Sakai |
ICCD | 4 |
| 2014 | A cloud architecture for protecting guest's information from malicious operators with memory managementabstractWe introduce a novel cloud computing architecture that ensures privacy for guest's information and computation. In conventional cloud architecture, a security policy proposed by a provider only ensured the protection of guest's information. This enabled malicious operators to steal or modify guest's information. Our architecture protects guest's information with novel memory management function of hypervisor from malicious operators. Cloud computing generally relies on virtualization, and VMM or hypervisor maintains page table for interfering VM's memory accesses, which is called shadow page table. Our hypervisor regulates memory accesses by management VM by adding a authority bit to shadow page table entry. Our architecture also prohibits a theft of guest's information when it is stored in storage by encrypting data when they leave memory. Koki Murakami, Tsuyoshi Yamada, Rie Shigetomi Yamaguchi, Masahiro Goshima, Shuichi Sakai |
CODASPY | 5 |
| 2010 | Register Cache System Not for Latency Reduction PurposeabstractA register cache has been proposed to solve the problems of the huge register files of recent super scalar processors. The register cache reduces the effective access latency of the register file for IPC improvement, simplifies the bypass network, and reduces the ports of the main register file. Though the primary purpose of the previous works is to improve IPC, the misses on the register cache may degrade the IPC. We propose Non-Latency-Oriented Register Cache System (NORCS). Though the effects of NORCS are the same as the conventional systems, it is free from register cache miss penalties that the conventional systems suffer from. In NORCS, the register cache itself is not different from that of the conventional systems. The difference is that the instruction pipeline has stages to read the main register file, which all instructions go through regardless of register cache hit / miss. Therefore, the instruction pipeline of NORCS is not immediately disturbed by the register cache misses. For a realistic 4-way super scalar processor, NORCS can simplify the bypass network to the same complexity as a 1-cycle-latency register file, and reduce the ports of the main register file from 12 to 4. CACTI simulation shows that the area and power consumption are reduced to 24.9% and 31.9% compared to the baseline model without register cache. Though these results are not different from the conventional systems, IPCs differ greatly. IPC of the conventional system decreases to 83.1% because of the cache miss penalties, while that of NORCS is retained at 98.0%. Ryota Shioya, Kazuo Horio, Masahiro Goshima, Shuichi Sakai |
MICRO | 4 |
| 2009 | Dependable VLSI: device, design and architecture: how should they cooperate?abstractVLSI dependability is one of the most significant issues in the modern world. Here the panelists will discuss the key technologies for it as well as the cost optimization among device, design and architecture. Shuichi Sakai, Hidetoshi Onodera, Hiroto Yasuura, James C. Hoe |
ASP-DAC | 1 |
| 2009 | String-Wise Information Flow Tracking against Script Injection AttacksabstractNowadays, security of Web applications faces a threat of script injection attacks. DTP (dynamic taint propagation) and DIFT (dynamic information flow tracking) have been established as powerful techniques to detect script injection attacks. However current DTP/DIFT systems still suffer from tradeoff between false positives and negatives.This paper proposes string-wise information flow tracking, SWIFT. SWIFT traces memory access of program execution, detects string access and distinguishes string operations from other memory access. Current DTP/DIFT systems propagate taint from source to destination operands. Instead of that, SWIFT propagates taint information under string operations. This makes SWIFT provide a better accuracy on detection of script injection attacks than current DTP/DIFT systems.We implemented SWIFT on an IA-32 emulator Bochs, executed typical string operations and made injection attacks to some real-world Web applications with known vulnerabilities. As a result, SWIFT shows a high precision in our security experiments. Kunbo Li, Ryota Shioya, Masahiro Goshima, Shuichi Sakai |
PRDC | 4 |
| 2009 | Low-Overhead Architecture for Security TagabstractA security-tagged architecture is one that applies tags on data to detect attack or information leakage, tracking data flow.The previous studies using security-tagged architecture mostly focused on how to utilize tags, not how the tags are implemented. A naive implementation of tags simply adds a tag field to every byte of the cache and the memory. Such technique, however, results in a huge hardware overhead.This paper proposes a low-overhead tagged architecture. We achieve our goal by exploiting some properties of tag, the non-uniformity and the locality of reference. Our design includes a use of uniquely designed multi-level table and various cache-like structures, all contributing to exploit these properties. Under simulation, our method was able to limit the memory overhead to 1.8%, where a naive implementation suffered 12.5% overhead. Ryota Shioya, Daewung Kim, Kazuo Horio, Masahiro Goshima, Shuichi Sakai |
PRDC | 5 |
| 2007 | Utilization of SECDED for soft error and variation-induced defect tolerance in caches
Luong Dinh Hung, Hidetsugu Irie, Masahiro Goshima, Shuichi Sakai |
DATE | 4 |
| 2006 | SEVA: A Soft-Error- and Variation-Aware Cache ArchitectureabstractAs SRAM devices are scaled down, the number of variation-induced defective memory cells increases rapidly. Combination of ECC, particularly SECDED, with a redundancy technique can effectively tolerate a high number of defects. While SECDED can repair a defective cell in a block, the block becomes vulnerable to soft errors. This paper proposes SEVA, an original soft-error- and variation-aware cache architecture. SEVA exploits SECDED to tolerate variation-induced defects while preserving high resilience against soft errors. Information about the defectiveness and data dirtiness is maintained for each SECDED block. SEVA allows only the clean data to be stored in defective (but still usable) blocks of a cache. An error occurring in a defective block can be detected and the correct data can be obtained from the lower level of the memory hierarchy. SEVA improves yield and reliability with low overheads Luong Dinh Hung, Masahiro Goshima, Shuichi Sakai |
PRDC | 3 |
| 2006 | Base Address Recognition with Data Flow Tracking for Injection Attack DetectionabstractVulnerabilities such as buffer overflows exist in some programs, and such vulnerabilities are susceptible to address injection attacks. The input data tracking method, which was proposed before, prevents I-data, which are the data derived from the input data, being used as addresses. However, the rules to determine address injection attacks are vague, which produces many false-positives and false-negatives in detection results. Generally, the data used as an address consist of a base address and an address offset. We propose an architectural technique to prevent I-data overwriting B-data, which are the data used as base addresses in this paper. It dynamically recognizes the I-data and the B-data. Address injection is detected if I-data that are not B-data are used as addresses. We implemented the proposed technique on a Pentium-based Bochs emulator and investigated its detection capability. We believe that the technique is the most accurate injection detection technique proposed thus far Satoshi Katsunuma, Hiroyuki Kurita, Ryota Shioya, Kazuto Shimizu, Hidetsugu Irie, Masahiro Goshima, Shuichi Sakai |
PRDC | 7 |
| 2005 | Mitigating Soft Errors in Highly Associative Cache with CAM-based TagabstractContent addressable memories (CAM) are widely used for the tag portions in highly associative caches. Since data are not explicitly read out of tag array in CAM search, the detection of false misses caused by soft errors for such caches is difficult. This paper presents a technique to detect the false miss in highly associative cache with CAM-based tag. The technique involves subdividing the tags and providing backup checking for cases the tags are partially matched. An original tag encoding scheme is proposed to reduce the frequency of back-up checking. Modifications to support the technique do not increase the cache access latency. The performance degradation incurred by additional cycles for false miss checking is very low. Luong Dinh Hung, Masahiro Goshima, Shuichi Sakai |
ICCD | 3 |
| 2005 | Cooking navi: assistant for daily cooking in kitchenabstractWe are developing a cooking navigation system, which helps even a novice user to cook several recipes in parallel without failure, while improving an advanced user's skill further. To realize this, the system optimizes the cooking procedure considering the following restrictions: (1) Duration of cooking, (2) Accuracy of cooking, and (3) Learning effect, by providing appropriate instructions to user's at the right timing, making full use of multimedia information. The users should be able to cook perfectly and comfortably just by following the text, video and audio provided by the system. According to the result of a preliminary experiment, all users from novice to experienced cooks could finish two dishes in parallel while enjoyeing the cooking very much. The result of a questionnaire shows the effectiveness of the multimedia navigation that we propose. Reiko Hamada, Jun Okabe, Ichiro Ide, Shin'ichi Satoh 0001, Shuichi Sakai, Hidehiko Tanaka |
ACM Multimedia | 5 |
| 2003 | Compiler-Assisted Thread Level Control Speculation
Hideyuki Miura, Luong Dinh Hung, Chitaka Iwama, Daisuke Tashiro, Niko Demus Barli, Shuichi Sakai, Hidehiko Tanaka |
Euro-Par | 6 |
| 2003 | Complexity Analysis of a Cache Controller for Speculative Multithreading Chip Multiprocessors
Yoshimitsu Yanagawa, Luong Dinh Hung, Chitaka Iwama, Niko Demus Barli, Shuichi Sakai, Hidehiko Tanaka |
HiPC | 5 |
| 2003 | A note on greedy algorithms for the maximum weighted independent set problem
Shuichi Sakai, Mitsunori Togasaki, Koichi Yamazaki |
Discret. Appl. Math. | 1 |
| 2002 | An object detection method for describing soccer games from videoabstractWe propose a novel object detection and tracking method in order to detect and track objects necessary to describe contents of a soccer game. On the contrary to intensity oriented conventional object detection methods, the proposed method refers to color rarity and local edge property, and integrally evaluates them by a fuzzy function to achieve better detection quality. These image features were chosen considering the characteristics of soccer video images, that most non-object regions are roughly single colored (green) and most objects tend to have locally strong edges. We also propose a simple object tracking method, that could track objects with occlusion with other objects using a color based template matching. The result of an evaluation experiment applied to actual soccer video showed very high detection rate in detecting player regions without occlusion, and promising ability for regions with occlusion. Okihisa Utsumi, Koichi Miura, Ichiro Ide, Shuichi Sakai, Hidehiko Tanaka |
ICME (1) | 4 |
| 2001 | Improving Conditional Branch Prediction on Speculative Multithreading Architectures
Chitaka Iwama, Niko Demus Barli, Shuichi Sakai, Hidehiko Tanaka |
Euro-Par | 3 |
| 1999 | Associating video with related documentsabstractArticle Free Access Share on Associating video with related documents Authors: Reiko Hamada Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, JapanView Profile , Ichiro Ide Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, JapanView Profile , Shuichi Sakai Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, JapanView Profile , Hidehiko Tanaka Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, JapanView Profile Authors Info & Claims MULTIMEDIA '99: Proceedings of the seventh ACM international conference on Multimedia (Part 2)October 1999Pages 17–20https://doi.org/10.1145/319878.319883Published:01 October 1999Publication History 1citation199DownloadsMetricsTotal Citations1Total Downloads199Last 12 Months11Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Reiko Hamada, Ichiro Ide, Shuichi Sakai, Hidehiko Tanaka |
ACM Multimedia (2) | 3 |
| 1999 | Integrated Manipulation: Context-Aware Manipulation of 2D DiagramsabstractDiagram manipulation in conventional CAD systems requires frequent mode switching and explicit placement of the pivot for rotation and scaling. In order to simplify this process, we propose an interaction technique called integrated manipulation, where the user can move, rotate, and scale without mode switching. In addition, the pivot for rotation and scaling automatically snaps to a contact point during moving operation. We performed a user study is performed using our prototype system and a commercial CAD system. The results showed that users could perform a diagram manipulation task much more rapidly using our technique. Masaaki Honda, Takeo Igarashi, Hidehiko Tanaka, Shuichi Sakai |
ACM Symposium on User Interface Software and Technology | 4 |
| 1997 | Virtual control channel and its application to the massively parallel computer RWC-1abstractGlobal operation and system control are important issues in massively parallel systems. The paper discusses virtual control networks (VCNs), which are substitutes for the current dedicated control networks. First we introduce a new mechanism called Virtual Control Channel (VCC) used to conduct control information over data network links. The network nodes have control finite state machines (CFSMs) and a VCN is composed of CFSMs. The VCC performs the role of a connection wire between CFSMs. The mechanisms are applied to the RWC-1 machine. The simulation results reveal that reduction and broadcasting operations are efficiently executed on VCNs by exploiting the tree structure. Takashi Yokota, Hiroshi Matsuoka, Kazuaki Okamoto, Hideo Hirono, Shuichi Sakai |
HiPC | 5 |
| 1997 | Fine-Grain Multithreading with the EM-X MultiprocessorabstractArticle Fine-grain multithreading with the EM-X multiprocessor Share on Authors: Andrew Sohn Computer and Information Science Dept., New Jersey Institute of Technology, Newark, NJ Computer and Information Science Dept., New Jersey Institute of Technology, Newark, NJView Profile , Yuetsu Kodama Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, Japan Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, JapanView Profile , Jui Ku Computer and Information Science Dept., New Jersey Institute of Technology, Newark, NJ Computer and Information Science Dept., New Jersey Institute of Technology, Newark, NJView Profile , Mitsuhisa Sato Real World Computing Tsukuba Research Center, Tsukuba, Ibaraki, 305, Japan Real World Computing Tsukuba Research Center, Tsukuba, Ibaraki, 305, JapanView Profile , Hirofumi Sakane Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, Japan Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, JapanView Profile , Hayato Yamana Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, Japan Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, JapanView Profile , Shuichi Sakai Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, Japan Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, JapanView Profile , Yoshinori Yamaguchi Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, Japan Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, JapanView Profile Authors Info & Claims SPAA '97: Proceedings of the ninth annual ACM symposium on Parallel algorithms and architecturesJune 1997 Pages 189–198https://doi.org/10.1145/258492.258511Online:01 June 1997Publication History 7citation121DownloadsMetricsTotal Citations7Total Downloads121Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Andrew Sohn, Yuetsu Kodama, Jui Ku, Mitsuhisa Sato, Hirofumi Sakane, Hayato Yamana, Shuichi Sakai, Yoshinori Yamaguchi |
SPAA | 7 |
| 1995 | A prototype router for the massively parallel computer RWC-1abstractThe RWC-1 is a massively parallel computer based on a multi-threaded architecture. This architecture requires extremely high communication performance with reasonable hardware cost. ln this paper, we first introduce a new class of direct interconnection networks called MDCE (Multidimensional Directed Cycles Ensemble extension). MDCE has many desirable features for RWC-1 including small degree, low latency, and high throughput. MDCE is thus adopted for a RWC-1 network. We have designed an MDCE router and fabricated an experimental VLSI chip. We explain the design details in this paper. The chip employs operating system support features as well as communication functions, and enables advanced resource management, A prototype chip with about 125,000 gates has been fabricated using 0.6-/spl mu/m CMOS gate array technology. Its clock runs at 50 MHz and a transmission rate of 300 M bytes per second per communication port is achieved. Takashi Yokota, Hiroshi Matsuoka, Kazuaki Okamoto, Hideo Hirono, Atsushi Hori, Shuichi Sakai |
ICCD | 6 |
| 1995 | A Macrotask-level Unlimited Speculative Execution on MultiprocessorsabstractThe purpose of this paper is to propose a new fast EM-4 multiprocessor to decrease the control overhead of macrotasks.Preliminary evaluations show that the control overhead of the proposed scheme is smaller than that of the other control schemes.Moreover, it is confirmed that the distributed control can be implemented by using software when the average macrotssk execution time is larger than 14.4ps on the EM-4 multiprocessor. Hayato Yamana, Mitsuhisa Sato, Yuetsu Kodama, Hirofumi Sakane, Shuichi Sakai, Yoshinori Yamaguchi |
International Conference on Supercomputing | 5 |
| 1995 | The EM-X Parallel Computer: Architecture and Basic PerformanceabstractLatency tolerance is essential in achieving high performance on parallel computers for remote function calls and fine-grained remote memory accesses. EM-X supports interprocessor communication on an execution pipeline with small and simple packets. It can create a packet in one cycle, and receive a packet from the network in the on-chip buffer without interruption. EM-X invokes threads on packet arrival, minimizing the overhead of thread switching. It can tolerate communication latency by using efficient multi-threading and optimizing packet flow of fine grain communication. EM-X also supports the synchronization of two operands, direct remote memory read/write operations and flexible packet scheduling with priority. This paper describes distinctive features of the EM-X architecture and reports the performance of small synthetic programs and larger more realistic programs. Yuetsu Kodama, Hirohumi Sakane, Mitsuhisa Sato, Hayato Yamana, Shuichi Sakai, Yoshinori Yamaguchi |
ISCA | 5 |
| 1995 | Time Space Sharing Scheduling and Architectural Support
Atsushi Hori, Takashi Yokota, Yutaka Ishikawa, Shuichi Sakai, Hiroki Konaka, Munenori Maeda, Takashi Tomokiyo, Jörg Nolte, Hiroshi Matsuoka, Kazuaki Okamoto, Hideo Hirono |
JSSPP | 4 |
| 1995 | Reduced Interprocessor-Communication Architecture and its Implementation on EM-4
Shuichi Sakai, Yuetsu Kodama, Mitsuhisa Sato, Andrew Shaw, Hiroshi Matsuoka, Hideo Hirono, Kazuaki Okamoto, Takashi Yokota |
Parallel Comput. | 1 |
| 1994 | Overview of RWC Massively Parallel Computer ProjectabstractSummary form only given, as follows. The author introduces the massively parallel computer research and development of the Japanese Real World Computing Project (RWC). RWC is a ten-year project whose budget is expected to be 600 million to 700 million dollars (US). We plan to develop an original 16 K PE stand-alone computer with a parallel C++ language, a more abstract object-oriented language, and a massively parallel operating system.> Shuichi Sakai |
HPDC | 1 |
| 1994 | Nonnumeric search results on the EM-4 distributed-memory multiprocessorabstractNumeric scientific problems have been the main focus of supercomputing as their numerous implementations on various multiprocessors indicate. Nonnumeric problems on the other hand have received very little attention for parallel implementation due to their irregular behaviors and a large amount of resource usage. This report presents our experiences implementing difficult nonnumeric problems on the EM-4 multiprocessor, believing that supercomputers should also be able to effectively execute nonnumeric problems if they are to be considered 'supercomputers'. We selected two typical search problems, the Eight-Puzzle and the Tower-of-Hanoi. Two parallel search techniques we used to implement the search problems, unidirectional and bidirectional heuristic search. A total of eight different programs have been implemented on the EM-4 multiprocessor with realistic problem sizes. Execution results demonstrate that the parallel bidirectional heuristic search can solve the tree depth 20 to 40 of the Eight-Puzzle in an optimal or near optimal number of iterations in less than two seconds, and is highly scalable as it gives over 40-fold speedup for both problems on 80 processors.> Andrew Sohn, Mitsuhisa Sato, Shuichi Sakai, Yuetsu Kodama, Yoshinori Yamaguchi |
SC | 3 |
| 1993 | EMC-Y: Parallel Processing Element Optimizing Communication and ComputationabstractEMC-Y is a new processing element for highly parallel computers designed to achieve high performance parallel computation by fusing a dataflow mechanism and a von Neumann execution pipeline. We have already developed EMC-R, which is the processing element used in the EM-4 prototype. EMC-Y improves on EMC-R's packet communication performance, allowing it to tolerate a more network traffic. This paper presents the architecture of EMC-Y, concentrating on the principles of packet communication. EMC-Y uses an output packet buffer and optimal packet routing to improve the performance of packet sending and transferring. EMC-Y changes the memory access priority for input packet buffer operation to improve the performance of receiving packets. Since the EMC-Y processor not only improves the performance of packet input and output but also balances them, it can tolerate a large amount of traffic and can improve the execution performance. We evaluate the improvements of EMC-Y architecture using a clock level simulator. The results show that EMC-Y improves performance by 50% to 70% in several programs over EMC-R at the same clock speed. Yuetsu Kodama, Yasuhito Koumura, Mitsuhisa Sato, Hirohumi Sakane, Shuichi Sakai, Yoshinori Yamaguchi |
International Conference on Supercomputing | 5 |
| 1993 | Super-Threading: Architectural and Software Mechanisms for Optimizing Parallel ComputationabstractThis paper presents super-threading, which generically means the architectural and software mechanisms for optimizing parallel computation. Super-threading includes architectural optimization of a processing element (PE), mechanism for supporting fast communication and computation, techniques of a compiler and a run time system for optimizing thread creation, thread allocation, tuning of granularity and data allocation to physically distributed storage.This paper states what super-threading is and examines some of the technologies belonging to it. The processor architecture based on super-threading is proposed and its implementation on a highly parallel computer EM-4 is shown with performance data. Software issues about super-threading are also examined mainly from the viewpoint of granularity optimization. Dynamic granularity optimization methods are proposed here, and evaluated on EM-4. The performance data indicate that super-threading is a key technology for realizing an efficient massively parallel computer. Shuichi Sakai, Kazuaki Okamoto, Hiroshi Matsuoka, Hideo Hirono, Yuetsu Kodama, Mitsuhisa Sato |
International Conference on Supercomputing | 1 |
| 1993 | Design and Implementation of a Circular Omega Network in the EM-4
Shuichi Sakai, Yuetsu Kodama, Yoshinori Yamaguchi |
Parallel Comput. | 1 |
| 1992 | Thread-based Programming for the EM-4 Hybrid Dataflow MachineabstractIn this paper, we present a thread-based programming model for the EM-4 hybrid dataflow machine, where parallelism and synchronization among threads of sequential execution are described explicitly by the programmer. Although EM-4 was originally designed as a dataflow machine, we demonstrate that it provides effective architectural support for a variety of programming styles, including message passing and distributed data sharing in imperative languages. Our approach allows the programmer to control the parallelism and maintain data locality explicitly to achieve high performance. EM-4 can be thought of as a multi-threaded architecture that can exploit both von Neumann and dataflow compiling technology. Thread-based programming provides the first step to explore better programming/compiling technology for a hybrid dataflow machine as well as EM-4. Mitsuhisa Sato, Yuetsu Kodama, Shuichi Sakai, Yoshinori Yamaguchi, Yasuhito Koumura |
ISCA | 3 |
| 1992 | A priority forwarding scheme for real-time multistage interconnection networksabstractThe authors propose a priority control scheme for packet switching multistage networks, called priority forwarding, which prevents priority inversion, a situation in which higher priority packets are blocked by lower priority packets. In an N*N omega network, the worst case delay of the priority forwarding scheme on the highest priority packet is O(log/sup 2/ N), while that for round-robin arbitration is O(N). Simulation results show that the priority forwarding scheme offers shorter delays for higher priority packets without throughput degradation and fits least-laxity-first control. The hardware implementation cost of this scheme is relatively small and requires no extra signal lines between routers. Consequently, this scheme offered predictability, scalability, and implementation eligibility.> Kenji Toda, Kenji Nishida, Shuichi Sakai, Toshio Shimada |
RTSS | 3 |
| 1992 | A prototype of a highly parallel dataflow machine EM-4 and its preliminary evaluation
Yuetsu Kodama, Shuichi Sakai, Yoshinori Yamaguchi |
Future Gener. Comput. Syst. | 2 |
| 1992 | Methodologies in development and testing of the dataflow machine EM-4
Kazuaki Okamoto, Yuetsu Kodama, Shuichi Sakai, Yoshinori Yamaguchi |
Parallel Comput. | 3 |
| 1991 | Design and Implementation of a Versatile Interconnection Network in the EM-4
Shuichi Sakai, Yuetsu Kodama, Yoshinori Yamaguchi |
ICPP (1) | 1 |
| 1991 | Load balancing by function distribution on the EM-4 prototypeabstractThe EM-4 is a highly parallel daiaflow machine that will eventually have more than 1,000 processing elements (PEs). This paper presents load balancing methods by function distribution in the EM-4 and their evaluations on the EM-4 prototype, which con-sists of 80 PEs. The EM-4 can distribute function in-stances statically by using several different allocation functions, including a method exploiting the locality of the network. Furthermore, the EM-4 can dynami-cally distribute function instances using MLPE packets which circulate through the PEs and detect the local-minimum load PE. These function distn’bution meth-ods are evaluated by executing a divide-and-conquer program and a game tree searching program, and ex-amining its dynamic characteristics on every PE. 1 Yuetsu Kodama, Shuichi Sakai, Yoshinori Yamaguchi |
SC | 2 |
| 1990 | Dataflow computer development in JapanabstractThis paper describes the research activity on dataflow computing in Japan focusing on dataflow computer development at the Electrotechnical Laboratory (ETL). First, the history of dataflow computer development in Japan is outlined. Some distinguished milestones in the history are mentioned in detail. Second, two types of dataflow computing systems developed at ETL, SIGMA-1 and EM-4, are described with their research goals. The fundamental characteristics of the both systems are given and some comparisons are made. Finally, future problems toward new generation computer systems to meet the challenge of Tera FLOPS machines are discussed. Toshitsugu Yuba, Toshio Shimada, Yoshinori Yamaguchi, Kei Hiraki, Shuichi Sakai |
ICS | 5 |
| 1989 | An Architecture of a Dataflow Single Chip ProcessorabstractA highly parallel (more than a thousand) dataflow machine EM-4 is now under development. The EM-4 design principle is to construct a high performance computer using a compact architecture by overcoming several defects of dataflow machines. Constructing the EM-4, it is essential to fabricate a processing element (PE) on a single chip for reducing operation speed, system size, design complexity and cost. In the EM-4, the PE , called EMC-R, has been specially designed using a 50,000-gate gate array chip. This paper focuses on an architecture of the EMC-R. The distinctive features of it are: a strongly connected arc dataflow model; a direct matching scheme; a RISC-based design; a deadlock-free on-chip packet switch; and an integration of a packet-based circular pipeline and a register-based advanced control pipeline. These features are intensively examined, and the instruction set architecture and the configuration architecture which exploit them are described. Shuichi Sakai, Yoshinori Yamaguchi, Kei Hiraki, Yuetsu Kodama, Toshitsugu Yuba |
ISCA | 1 |