EDBT 2026 Demo / reviewers in the wild / expert
Yasuhiko Nakashima
dblp:70/1442
· DBLP profile ↗
33ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0002-9457-5061ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 10 · 4 since 2021Security and privacy · 2Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mutual Information Loss for Enhancing Class-Wise Representation in Spiking Neural Networks
Yiling Yang, Yirong Kan, Yasuhiko Nakashima |
ISCAS | 4 |
| 2026 | MergeSlide: Continual Model Merging and Task-to-Class Prompt-Aligned Inference for Lifelong Learning on Whole Slide ImagesabstractLifelong learning on Whole Slide Images (WSIs) aims to train or fine-tune a unified model sequentially on cancer-related tasks, reducing the resources and effort required for data transfer and processing, especially given the gigabyte-scale size of WSIs. In this paper, we introduce MergeSlide, a simple yet effective framework that treats lifelong learning as a model-merging problem by leveraging a vision–language pathology foundation model. When a new task arrives, it is ❶ defined with class-aware prompts, ❷ fine-tuned for a few epochs in a classifier-free manner, and ❸ merged into a unified model using an orthogonal continual-merging strategy that preserves performance and mitigates catastrophic forgetting. For inference under the class-incremental learning (CLASS-IL) setting, where task identity is unknown, we introduce Task-to-Class Prompt-aligned (TCP) inference. Specifically, TCP first identifies the most relevant task using task-level prompts and then applies the corresponding class-aware prompts to generate predictions. To evaluate MergeSlide, we conduct experiments on a stream of six TCGA datasets. The results show that MergeSlide outperforms both rehearsal-based continual learning and vision-language zero-shot baselines. Code and data are available at https://github.com/caodoanh2001/MergeSlide. Doanh C. Bui, Ba Hung Ngo, Hoai Luan Pham, Khang Nguyen 0001, Maï K. Nguyen, Yasuhiko Nakashima |
WACV | 6 |
| 2026 | FPSpike: A Fully Parallel and Reconfigurable Architecture for Accelerating Spiking Neural Networks With Structured SparsityabstractThis paper introduces FPSpike, a fully-parallel and reconfigurable architecture for accelerating spiking neural networks (SNNs) with structured sparsity. By introducing structured sparse synaptic connections, the neuron computation and weight storage costs are significantly reduced, while the wiring constraints of hardware implementation are alleviated, thus realizing a fully parallel SNN architecture on a single chip. Furthermore, FPSpike achieves high reconfigurability through local connections, allowing the network topology to be flexibly partitioned into multiple distinct hardware cores without redundancy, thereby enabling multi-task spatial parallel processing. In particular, a series of software-hardware co-designs are performed to improve the performance of FPSpike. Experimental results with FPGA-based implementation demonstrate that FPSpike achieves inference throughput improvements of$2.5\times \sim 43.9\times $and$3.2\times \sim 93.8\times $compared to Intel i7-12700K CPU and NVIDIA RTX3060Ti GPU, respectively. FPSpike also achieves performance speedups of$2.0\times \sim 20.2\times $and energy efficiency improvements of$1.3\times \sim 7.7\times $compared to the state-of-the-art FPGA-based SNN accelerators. Yirong Kan, Man Wu, Yasuhiko Nakashima |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2026 | DystoHD: an area-efficient hyperdimensional computing system with dynamic hypervector generation for memory-constrained devices
Yirong Kan, Man Wu, Yasuhiko Nakashima |
J. Supercomput. | 5 |
| 2025 | CTFE: A High-Efficient Heterogeneous Cryptographic CGRA for Diverse Security ApplicationsabstractNowadays, cryptographic computation across various security applications necessitates the development of hardware that is not only fast and power-efficient but also flexible enough to support a range of cryptographic algorithms. Unfortunately, existing computing platforms for cryptography struggle to balance high flexibility, high performance, and low power consumption. To address these issues, this article introduces the crypto-tailored flexible engine (CTFE), a next-generation coarse-grained reconfigurable array (CGRA) for cryptography. Concretely, the CTFE incorporates four innovative ideas to achieve high flexibility and performance with high hardware efficiency: 1) processing element array (PEA) with dual-buffer lanes and multiplexer optimization; 2) high flexibility Hyper-ALU; 3) heterogeneous PEA; and 4) bi-tiered pipeline coordination. Real-time evaluation results on Xilinx ZCU102 FPGA at the System-on-Chip (SoC) level demonstrate that the CTFE is 1.13–14.3 times better in throughput and 55.3–14 232 times better in energy efficiency than state-of-the-art CPUs. Experiments on an ASIC 45 nm CMOS technology show that the CTFE consumes the power of 1.06 W, occupies an area of$2.77 \; \text {mm}^{{2}}$, and operates at the frequency of 510 MHz. In comparison to existing CGRA solutions, CTFE outperforms 1.63–20.23 times in throughput and 1.61–73.2 times in area efficiency. Vu Trung Duong Le, Hoai Luan Pham, Thi Hong Tran, Van Duy Tran, Tuan Hai Vu 0001, Yasuhiko Nakashima |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2025 | MINA: A Hardware-Efficient and Flexible Mini-InceptionNet Accelerator for ECG Classification in Wearable DevicesabstractClassification is a crucial aspect of cardiovascular-related challenges, requiring thorough research and optimization to develop effective solutions for both patients and doctors. Recently, rapid advancements in artificial intelligence, particularly Convolutional Neural Networks (CNNs), have introduced numerous effective methods, significantly improving disease classification in Electrocardiogram (ECG) analysis. However, existing CNN-based accelerators often encounter challenges such as high parameter counts, limited flexibility in handling diverse CNN configurations, and inefficient hardware utilization. To address these issues, this paper proposes the Mini InceptionNet Accelerator (MINA), a hardware-efficient and flexible accelerator designed specifically for one-dimensional (1-D) CNN-based ECG classification. First, a novel 1-D CNN model, Mini InceptionNet, reduces the parameter count by 41.6% compared to the smallest existing 1-D CNN, minimizing memory requirements while maintaining high classification accuracy. Second, a flexible Processing Element Array (PEA) is designed with a Sharing Buffer Allocator (SBA) to support dynamic data coordination across various network topology parameters. Third, each Processing Element (PE) is equipped with four Local Data Memories (LDMs) and an ALU, enabling efficient intermediate data storage and versatile operations for modern CNN models. To demonstrate its effectiveness, MINA has been successfully implemented and verified on the ZCU102 FPGA at the system-on-chip level. FPGA evaluations show that MINA achieves 1.3×-2.9× higher energy efficiency (GOP/s/MeLUT) than state-of-the-art 2-D CNN accelerators. Compared to existing 1-D CNN accelerators, MINA achieves at least 1.53× improvement in the area-delay product (ADP). Additionally, weight pruning is discussed as a supporting strategy, achieving up to 3× faster inference time and a 2.13× improvement in ADP at 70% sparsity. Hoai Luan Pham, Vu Trung Duong Le, Yasuhiko Nakashima |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | A Fully-Parallel Reconfigurable Spiking Neural Network Accelerator with Structured Sparse ConnectionsabstractIn this work, we present a fully parallel reconfigurable spiking neural network (SNN) accelerator for various applications of edge computing. In contrast to conventional fully connected and irregular sparse network topologies, structured sparse synaptic connections are introduced to implement SNN in fully parallel on hardware by significantly reducing neuron computation and weight memory costs. Benefiting from the hardware-friendly symmetric SNN topology, the proposed accelerator is flexibly configured into multiple classifiers without hardware redundancy to support various tasks. Furthermore, to compress the layer depth of individual classifiers to reduce latency, layer-skip connections are introduced to alleviate the vanishing gradient problem in training. Various experiments are conducted to explore the optimal settings of the proposed accelerator. The proposed SNN accelerator is verified on the Xilinx ZCU102 FPGA. The results show that the proposed SNN accelerator achieves accuracy of 97.1% and energy efficiency of 40.1 GSOP/s/W on the MNIST dataset. Yirong Kan, Yasuhiko Nakashima |
ISCAS | 4 |
| 2024 | Fusion synapse by memristor and capacitor for spiking neuromorphic systems
Takumi Kuwahara, Reon Oshio, Mutsumi Kimura, Yasuhiko Nakashima |
Neurocomputing | 5 |
| 2024 | LiCryptor: High-Speed and Compact Multi-Grained Reconfigurable Accelerator for Lightweight CryptographyabstractEmerging modern internet-of-things (IoT) systems require hardware development to support multiple 8/32/64-bit lightweight cryptographic (LWC) algorithms with high speed and energy efficiency to ensure diverse security requirements. Accordingly, a coarse-grained reconfigurable array (CGRA) is considered the most effective architecture for achieving high speed, low power, and high flexibility for implementing LWC algorithms. However, existing CGRA designs for cryptography focus only on improvements to outdated 8/32-bit algorithms, suffer from large area requirements, and have long compilation times. To address these issues, this paper proposes a new CGRA-based accelerator named LiCryptor to support various 8/32/64-bit LWC algorithms with high speed and small area. Three innovative ideas are proposed to enable LiCryptor to achieve these goals: a compact multi-grained processing element array (M-PEA), a shared 8/32/64-bit arithmetic logic unit (ALU), and an assembly-like inline directive (AID) mapping method. The LiCryptor has been successfully implemented and verified on the Xilinx ZCU102 FPGA. Real-time performance evaluation across various LWC algorithms on FPGA shows that LiCryptor is 1.33 to 4 times better in execution time and 3.4 to 153 times better in power-delay products (PDP) compared to today’s most powerful CPUs. Notably, evaluation of AID mapping on the ARM Cortex-A53 CPU of the ZCU102 FPGA shows that its compilation time is less than 1.5 ms for most LWC algorithms, at least 2,333 times faster than CFG mapping in current CGRAs. Moreover, experimental results on 45nm ASIC technology show that the LiCryptor significantly outperforms existing CGRAs and other reconfigurable designs in terms of throughput and area efficiency. Hoai Luan Pham, Vu Trung Duong Le, Van Duy Tran, Tuan Hai Vu 0001, Yasuhiko Nakashima |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | Bisection Neural Network Toward Reconfigurable Hardware ImplementationabstractA hardware-friendly bisection neural network (BNN) topology is proposed in this work for approximately implementing massive pieces of complex functions in arbitrary on-chip configurations. Instead of the conventional reconfigurable fully connected neural network (FC-NN) circuit topology, the proposed hardware-friendly topology performs NN behaviors in a bisection structure, in which each neuron includes two constant synapse connections for both inputs and outputs. Compared with the FC-NN one, the reconfiguration of the BNN circuit topology eliminates the remarkable amount of dummy synapse connections in hardware. As the main target application, this work aims at building a general-purpose BNN circuit topology that offers a great amount of NN regressions. To achieve this target, we prove that the NN behaviors of the FC-NN circuit topologies can be migrated to the BNN circuit topologies equivalently. We introduce two approaches including the refining training algorithm and the inverted-pyramidal strategy to further reduce the number of neurons and synapses. Finally, we conduct the inaccuracy tolerance analysis to suggest the guideline for ultra-efficient hardware implementations. Compared with the state-of-the-art FC-NN circuit topology-based TrueNorth baseline, the proposed design can achieve 17.8- 22.2× hardware reduction and less than 1% inaccuracy. Yirong Kan, Sa Yang, Yasuhiko Nakashima |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Neuromorphic System Using Memcapacitors and Autonomous Local LearningabstractArtificial intelligence is used for various applications and is promising as an indispensable infrastructure in future societies. Neural networks are representative technologies that imitate human brains and exhibit various advantages. However, the size is bulky, the power is huge, and some advantages are not demonstrated because they are executed on Neumann-type computers. Neuromorphic systems are biomimetic systems from the hardware level to implement neuron and synapse elements, and the size is compact, the power is low, and the operation is robust. However, because the conventional ones are not composed of fully optimized hardware, the power is not yet minimal, and extra control circuits must be used. In this article, we developed a neuromorphic system using memcapacitors and autonomous local learning. By using memcapacitors, the power can be minimal, and by using autonomous local learning, the control circuits to handle the synapse elements can be deleted. First, the memcapacitors are completed in a cross-bar array, where the ferroelectric layers are sandwiched between the horizontal and perpendicular electrodes. The polarization and capacitance exhibit hysteresis due to the dielectric polarization. Next, autonomous local learning is introduced as follows. During the training phase, associative patterns to be memorized are directly sent, relatively high voltages are applied, and dielectric polarizations are induced. During the operation phase, relatively low voltages are applied, and input signals are weighted with the capacitances of the memcapacitors, summed, and transferred as the output signals. Finally, the experimental system is set up, and the experimental results are acquired. The memorized patterns during the training phase, distorted patterns as the input signals during the operation phase, and retrieved patterns as the output signals in the operation phase are shown. Researchers found that the retrieved patterns are completely the same as the memorized patterns. This means that the neuromorphic system works as an associative memory. Mutsumi Kimura, Yuma Ishisaki, Yuta Miyabe, Homare Yoshida, Isato Ogawa, Tomoharu Yokoyama, Ken-Ichi Haga, Eisuke Tokumitsu, Yasuhiko Nakashima |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2022 | MuGRA: A Scalable Multi-Grained Reconfigurable Accelerator Powered by Elastic Neural NetworkabstractA massive core computing architecture is developed for accelerating arbitrary calculations in fully parallel with high speed and low cost. The proposed architecture is reconfigurable in fine-grained (arbitrary functions), mid-grained (flexible function feature, accuracy, and number of operands), and coarse-grained (organization of cores). By implementing a large scale of novel bisection neural network (BNN) on hardware, the re-configuration is conducted by partitioning entire BNN into any specific pieces without redundancy. Each piece of BNN retrieves the arbitrary function approximately. By reconfiguring the BNN topology in software, we can easily adjust dimensions of the computing kernel without rewiring, and achieve a wide range of trade-offs between accuracy and efficiency in hardware. In this manner, the multi-grained reconfigurable accelerator (MuGRA) is achieved. Since MuGRA is flexible in all grained levels, various configurations for each validation are demonstrated with rich options of performance-cost matrix. From the FPGA implementation results, compared with other traditional function approximation methods, our method provides fewer parameter storage requirements. The comparison against related works proves that our accelerator effectively reduces the calculation latency with slight accuracy loss. Yirong Kan, Man Wu, Yasuhiko Nakashima |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | DiaNet: An elastic neural network for effectively re-configurable implementationabstractAn elastic neural network is developed and evolved towards effectively re-configurable hardware in fully parallel on chip. The original prototype of DiaNet is organized as a symmetrical bisection neural network, which is feasible to be partitioned into arbitrary pieces of neural networks (NNs) without redundancy. To prevent the depth explosion in implementing complex tasks (complicated pattern recognition for instance), the evolution of DiaNets is investigated in this work. By using the I/O layer integration technology which enables all neurons in the hidden layer of DiaNet to receive inputs, the number of layers is reduced to 8.8% of DiaNet prototype. In this manner, the DiaNet topology is feasible to implement complex NNs without the risk of depth explosion. Moreover, the skip connection technology is proposed to avoid the gradient vanishing due to deep learning, which is significant to DiaNets especially. Compared with the LeNet5 model as state-of-the-art, the evolved DiaNet topology achieves the parameter reduction of 90.86% for MNIST recognition with the negligible loss of accuracy. To reduce hardware utilization, the sensitivity to the decline of computational precision and bit-width is investigated to suggest the guideline for efficient hardware implementations. Finally, the effectiveness of DiaNet is verified by the proposed re-configurable architecture on FPGA with the power reduction of 10.8% compared to state-of-the-art implementations. Man Wu, Yirong Kan, Tati Erlina, Yasuhiko Nakashima |
Neurocomputing | 5 |
| 2021 | Efficient hardware task migration for heterogeneous FPGA computing using HDL-based checkpointing
Hoang Gia Vu, Takashi Nakada, Yasuhiko Nakashima |
Integr. | 3 |
| 2021 | BCA: A 530-mW Multicore Blockchain Accelerator for Power-Constrained Devices in Securing Decentralized NetworksabstractBlockchain distributed ledger technology (DLT) has widespread applications in society 5.0 because it improves service efficiency and significantly reduces labor costs. However, employing blockchain DLT entails considerable energy consumption in the mining process. This paper proposes a blockchain accelerator (BCA) with ultralow power consumption and a high processing rate to address the problem. The BCA focuses on accelerating the double secure hash algorithm (SHA) 256 function required in the mining process at a system-on-chip (SoC) level. We propose three ideas, namely, multiple local memories (multimem), double-cell processing element (D-PE), and nonce autoupdate (NAU), to reduce the external data transfer time and improve the BCA hardware efficiency. We propose a cascaded multiple BCA chip model to enhance the system throughput by several-fold. Our experiments on an ASIC and FPGA prove that the proposed BCA successfully performs the mining process for multiple blockchain networks with much lower power consumption than that of the state-of-the-art CPUs and GPUs. The BCA is laid out with Renesas 65 nm technology with a chip area of$25~mm^{2}$and consumes$530~mW$at 100MHz. The power efficiency of the layout chip is improved by 2428 and 143 times compared with that of the fastest CPU Intel i9-10940X and GPU RTX 3090, respectively. Thi Hong Tran, Hoai Luan Pham, Tri Dung Phan, Yasuhiko Nakashima |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2020 | A Compact and Accuracy-Reconfigurable Univariate RBF Kernel Based on Stochastic LogicabstractThis paper presents a novel architecture for univariate radial basis kernel computation employing Stochastic Computing. The univariate radial basic function is optimized using simple stochastic logic circuits. We validated this approach by comparison with both Bernstein polynomial and two-dimensional finite-state-machine-based implementation. Optimally, the mean absolute error is reduced 40% and 80% compared to two other well-known approaches, Bernstein polynomial and two-dimensional finite-state-machine-based implementation, respectively. In terms of hardware cost, our proposed solution required as much as the Bernstein method did. Moreover, the proposed approach outperforms the two-dimensional finite-state-machine-based mplementation, roughly 54% less hardware cost. Regarding the critical path delay, the proposed approach is 12% less than others on average. Besides, our work also required 70% less power than two-dimensional finite-state-machine-based implementation. Tieu-Khanh Luong, Yasuhiko Nakashima |
ISCAS | 4 |
| 2019 | Run-Length Limited Decoding for Visible Light Communications: A Deep Learning ApproachabstractIn visible light communication (VLC) system, flicker mitigation is considered as an essential requirement that can be met by the use of run-length limited (RLL) code. Among researches about RLL code, soft input soft output (SISO) RLL decoding scheme is proved significantly improving the error correction performance of the entire VLC system. However, it suffers from high computational complexity. To overcome this limitation, this paper proposed a deep learning framework that effectively performs the RLL decoding at a VLC receiver. The proposed framework makes use of a one-layer long short term memory (LSTM) followed by a fully-connected layer to learn the input-output relation of RLL decoder. Additionally, we present numerical results to validate the excellent performance of the proposed RLL decoder. The results exhibit that the combination between FEC and DL-based RLL approach is capable of achieving the bit error rate (BER) performance of concatenated FEC and SISO RLL decoder; while substantially reducing computational complexity. Dinh-Dung Le, Duc Phuc Nguyen, Thi Hong Tran, Yasuhiko Nakashima |
APCC | 4 |
| 2019 | An Efficient Time-based Stochastic Computing Circuitry Employing Neuron-MOSabstractA compact and low energy circuitry of time-based stochastic computing (TBSC) have been designed. In the TBSC theory, stochastic numbers (SNs) are represented by duty-cycle of periodic signals. Additionally, multiplication and addition operations of the SNs require the signals to be in-harmonic and uncorrelated in frequency. In order to dynamically tune the frequency, a current-starved structure is applied in a three-stage inverter chain type of oscillator with the neuron-MOS mechanism. By feeding the pulses to a neuron-MOS based inverter with adjustable switching threshold, arbitrary duty-cycle for representing any specific SNs can be generated with the accuracy of 96%. In this manner, the implementation cost of the stochastic number generator (SNG) is reduced by avoiding the use of complex frequency-programmable-oscillator and comparator which are exploited in the conventional TBSC circuit. From circuit simulation results, multiplication and addition operations can be retrieved by proposed TBCS circuits with the accuracy of about 97%. The entire circuitry utilizes 210 CMOS transistors, and consumes the energy of 2.5pJ for one computation, which is 14% and 36.7% of conventional TBCS circuit, respectively. Tati Erlina, Yasuhiko Nakashima |
ACM Great Lakes Symposium on VLSI | 4 |
| 2019 | Neuro-inspired System with Crossbar Array of Amorphous Metal-Oxide-Semiconductor Thin-Film Devices as Self-plastic Synapse Units
Mutsumi Kimura, Kenta Umeda, Keisuke Ikushima, Toshimasa Hori, Ryo Tanaka, Tokiyoshi Matsuda, Tomoya Kameda, Yasuhiko Nakashima |
ICONIP (2) | 8 |
| 2018 | Hopfield Neural Network with Double-Layer Amorphous Metal-Oxide Semiconductor Thin-Film Devices as Crosspoint-Type Synapse Elements and Working Confirmation of Letter Recognition
Mutsumi Kimura, Kenta Umeda, Keisuke Ikushima, Toshimasa Hori, Ryo Tanaka, Tokiyoshi Matsuda, Tomoya Kameda, Yasuhiko Nakashima |
ICONIP (7) | 8 |
| 2017 | CPRring: A Structure-Aware Ring-Based Checkpointing Architecture for FPGA ComputingabstractIn this paper, we present a new architecture for FPGA checkpointing along with an efficient mechanism. We then provide a static analysis of original HDL source code to reduce the cost of hardware for checkpointing functionality. Our evaluations show that with the proposals, checkpointing hardware causes small degradation in maximum clock frequency (less than10%). The LUT overhead varies from 14.4% (Dijkstra) to 103.84%(Matrix Multiplication). Hoang Gia Vu, Shinya Takamaeda-Yamazaki, Takashi Nakada, Yasuhiko Nakashima |
FCCM | 4 |
| 2017 | Neuromorphic Hardware Using Simplified Elements and Thin-Film Semiconductor Devices as Synapse Elements - Simulation of Hopfield and Cellular Neural Network -
Tomoya Kameda, Mutsumi Kimura, Yasuhiko Nakashima |
ICONIP (6) | 3 |
| 2017 | Cellular neural network formed by simplified processing elements composed of thin-film transistors
Mutsumi Kimura, Ryohei Morita, Sumio Sugisaki, Tokiyoshi Matsuda, Tomoya Kameda, Yasuhiko Nakashima |
Neurocomputing | 6 |
| 2016 | Letter Reproduction Simulator for Hardware Design of Cellular Neural Network Using Thin-Film Synapses - Crosspoint-Type Synapses and Simulation Algorithm
Tomoya Kameda, Mutsumi Kimura, Yasuhiko Nakashima |
ICONIP (2) | 3 |
| 2016 | Simplification of Processing Elements in Cellular Neural Networks - Working Confirmation Using Circuit Simulation
Mutsumi Kimura, Nao Nakamura, Tomoharu Yokoyama, Tokiyoshi Matsuda, Tomoya Kameda, Yasuhiko Nakashima |
ICONIP (2) | 6 |
| 2014 | Emulator-oriented tiny processors for unreliable post-silicon devices: A case studyabstractAlthough various post-silicon devices have been invested in recent years, they still have a major issue of reliability. Because circuit area is an essential factor of reliability, especially for such unreliable post-silicon devices, it is desired to build small circuits which can reuse as many today's application programs as possible even if the performance is not very high. This paper presents the very first work to study novel, efficient techniques of emulating wider-bit guest processors (e.g., 32-bit) on a narrower-bit host processor (e.g., 8-bit) with very limited hardware resources while mitigating performance degradation. We propose three types of emulator-oriented tiny processors varying in available hardware resources and reliability enhancement approaches. Quantitative evaluation and discussions are done for comparing those three processors. We believe that this work will will be a good help of making breakthrough for further development of new device technologies and computers based on them. Yuko Hara-Azumi, Masaya Kunimoto, Yasuhiko Nakashima |
ASP-DAC | 3 |
| 2014 | Better-Than-DMR Techniques for Yield Improvement
Shunichi Sanae, Yuko Hara-Azumi, Shigeru Yamashita, Yasuhiko Nakashima |
FCCM | 4 |
| 2012 | Tensor Rank and Strong Quantum Nondeterminism in Multiparty Communication
Marcos Villagra, Masaki Nakanishi, Shigeru Yamashita, Yasuhiko Nakashima |
TAMC | 4 |
| 2010 | A Minimal Roll-Back Based Recovery Scheme for Fault Toleration in Pipeline ProcessorsabstractIn this paper, we proposed a light-weighted recovery scheme for fault tolerable pipeline processors after error has been detected by redundant executions. A minimal rolling back procedure is designed to schedule the re-execution based recovery in a one-cycle delay. This scheme makes full use of in-fly pipeline working status to aid the recovery, which relieves the recovery from a large checkpoint buffer. Jun Yao 0001, Ryoji Watanabe, Takashi Nakada, Hajime Shimada, Yasuhiko Nakashima, Kazutoshi Kobayashi |
PRDC | 5 |
| 2009 | A Speculative Technique for Auto-Memoization Processor with MultithreadingabstractWe have proposed an auto-memoization processor. This processor automatically and dynamically memoizes both functions and loop iterations, and skips their execution by reusing their results. On the other hand, multi/many-core processors have come into wide use. The number of cores is expected to increase to a hundred or more. However, many programs do not have so much parallelism in them. Therefore it becomes very important to consider how to utilize many cores effectively. This paper describes a speedup technique for auto-memoization processor using speculative multi-threading. Two speculative threads will be forked on reuse test. The one assumes that the reuse test will succeed, and executes the following codes of the reuse target block speculatively. The other assumes that the reuse test will fail, and executes the reuse target block. These two threads conceal the overhead of auto-memoization processor. The result of the experiment with SPEC CPU95 suite benchmarks shows that proposing method improves the maximum speedup from 13.9% to 36.0%. Yushi Kamiya, Tomoaki Tsumura, Hiroshi Matsuo, Yasuhiko Nakashima |
PDCAT | 4 |
| 2008 | A Functional Unit with Small Variety of Highly Reliable CellsabstractRecently, the miniaturization process has brought an increase in transistor variations and in the failure rate at transistors. We propose a small variety of new standard cells. The proposed cells can correct and detect transistor faults. A functional unit with the proposed cells shows better fault tolerance. The area of this unit is approximately 1.4 times that of traditional cells. Kouki Suzuki, Takashi Nakada, Masaki Nakanishi, Shigeru Yamashita, Yasuhiko Nakashima |
PRDC | 5 |
| 2001 | A high-speed dynamic instruction scheduling scheme for superscalar processors
Masahiro Goshima, Kengo Nishino, Toshiaki Kitamura, Yasuhiko Nakashima, Shinji Tomita, Shin-ichiro Mori |
MICRO | 4 |
| 1995 | Scalar Processor of the VPP500 Parallel SupercomputerabstractArticle Free Access Share on Scalar processor of the VPP500 parallel supercomputer Authors: Yasuhiko Nakashima Fujitsu Limited, 1015 Kamikodanaka, Nakahara-ku Kawasaki 211, Japan Fujitsu Limited, 1015 Kamikodanaka, Nakahara-ku Kawasaki 211, JapanView Profile , Toshiaki Kitamura Fujitsu Limited, 1015 Kamikodanaka, Nakahara-ku Kawasaki 211, Japan Fujitsu Limited, 1015 Kamikodanaka, Nakahara-ku Kawasaki 211, JapanView Profile , Hideo Tamura Fujitsu Limited, 1015 Kamikodanaka, Nakahara-ku Kawasaki 211, Japan Fujitsu Limited, 1015 Kamikodanaka, Nakahara-ku Kawasaki 211, JapanView Profile , Masaaki Takiuchi Fujitsu Limited, 1015 Kamikodanaka, Nakahara-ku Kawasaki 211, Japan Fujitsu Limited, 1015 Kamikodanaka, Nakahara-ku Kawasaki 211, JapanView Profile , Kenichi Miura Fujitsu America Inc, 3055 Orchard Dr, M/S 1-5, San Jose, CA Fujitsu America Inc, 3055 Orchard Dr, M/S 1-5, San Jose, CAView Profile Authors Info & Claims ICS '95: Proceedings of the 9th international conference on SupercomputingJuly 1995 Pages 348–356https://doi.org/10.1145/224538.224624Published:03 July 1995Publication History 2citation235DownloadsMetricsTotal Citations2Total Downloads235Last 12 Months23Last 6 weeks7 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Yasuhiko Nakashima, Toshiaki Kitamura, Hideo Tamura, Masaaki Takiuchi, Ken'ichi Miura |
International Conference on Supercomputing | 1 |