Masatoshi Ishii

dblp:17/7201 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
4since 2021 · last 2023
0000-0003-0794-7232ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2023 Architectures and Circuits for Analog-memory-based Hardware Accelerators for Deep Neural Networks (Invited)
abstract
Analog non-volatile memory (NVM)-based accelerators for Deep Neural Networks (DNNs) can achieve high-throughput and energy-efficient multiply-accumulate (MAC) operations by taking advantage of massively parallelized analog compute, implemented with Ohm's law and Kirchhoff's current law on arrays of resistive memory devices. Competitive end-to-end DNN accuracies can be obtained, provided that weights are accurately programmed onto NVM devices and MAC operations are sufficiently linear. In this paper, we report architectural and circuit advances for such Analog NVM-based accelerators. We describe a highly heterogeneous and programmable accelerator architecture for DNN inference that combines analog NVM memory-array “Tiles” for weight-stationary, energy-efficient MAC operations, together with heterogeneous special-function Compute-Cores for auxiliary digital computation. Massively parallel vectors of neuron-activation data are exchanged over short distances using a dense and efficient circuit-switched 2D mesh, enabling a wide range of DNN workloads, including CNNs, LSTMs, and Transformers. We also show a 14-nm inference chip consisting of multiple$\mathbf{512}\times \mathbf{512}$arrays of Phase Change Memory (PCM) devices which implements multiple DNN benchmarks using such a circuit-switched 2D mesh.
Hsinyu Tsai, Pritish Narayanan, Shubham Jain 0004, Stefano Ambrogio, Kohji Hosokawa, Masatoshi Ishii, Charles Mackin, Ching-Tzu Chen, Atsuya Okazaki, Akiyo Nomura, Irem Boybat, Ramachandran Muralidhar, Martin M. Frank, Takeo Yasuda, Alexander M. Friz, Yasuteru Kohda, An Chen 0002, Andrea Fasoli, Malte J. Rasch, Stanislaw Wozniak, Jose Luquin, Vijay Narayanan, Geoffrey W. Burr
ISCAS6
2023 A Heterogeneous and Programmable Compute-In-Memory Accelerator Architecture for Analog-AI Using Dense 2-D Mesh
abstract
We introduce a highly heterogeneous and programmable compute-in-memory (CIM) accelerator architecture for deep neural network (DNN) inference. This architecture combines spatially distributed CIM memory array “tiles” for weight-stationary, energy-efficient multiply–accumulate (MAC) operations, together with heterogeneous special-function compute cores for auxiliary digital computation. Massively parallel vectors of neuron activation data are exchanged over short distances using a dense and efficient circuit-switched 2-D mesh, offering full end-to-end support for a wide range of DNN workloads, including CNNs, long-short-term-memory (LSTM), and transformers. We discuss the design of the “analog fabric”—the 2-D grid of tiles and compute cores interconnected by the 2-D mesh—and address the efficiency in both mapping of DNNs onto the hardware and in pipelining of various DNN workloads across a range of batch sizes. We show, for the first time, system-level assessments using projected component parameters for a realistic “analog AI” system, based on dense crossbar arrays of low-power nonvolatile analog memory elements, while incorporating a single common analog fabric design that can scale to large networks by introducing data transport between multiple analog AI chips. Our performance estimates for several networks, including large LSTM and bidirectional encoder representations from transformers (BERT), show highly competitive throughput while offering$40\times $–$140\times $higher energy efficiency than NVIDIA A100—thus illustrating the strong promise of analog AI and the proposed architecture for DNN inference applications.
Shubham Jain 0004, Hsinyu Tsai, Ching-Tzu Chen, Ramachandran Muralidhar, Irem Boybat, Martin M. Frank, Stanislaw Wozniak, Milos Stanisavljevic, Praneet Adusumilli, Pritish Narayanan, Kohji Hosokawa, Masatoshi Ishii, Vijay Narayanan, Geoffrey W. Burr
IEEE Trans. Very Large Scale Integr. Syst.12
2022 Analog-memory-based 14nm Hardware Accelerator for Dense Deep Neural Networks including Transformers
abstract
Analog non-volatile memory (NVM)-based accelerators for deep neural networks perform high-throughput and energy-efficient multiply-accumulate (MAC) operations (e.g., high TeraOPS/W) by taking advantage of massively parallelized analog MAC operations, implemented with Ohm’s law and Kirchhoff’s current law on array-matrices of resistive devices. While the wide-integer and floating-point operations offered by conventional digital CMOS computing are much more suitable than analog computing for conventional applications that require high accuracy and true reproducibility, deep neural networks can still provide competitive end-to-end results even with modest (e.g., 4-bit) precision in synaptic operations. In this paper, we describe a 14-nm inference chip, comprising multiple 512$\times$ 512 arrays of Phase Change Memory (PCM) devices, which can deliver software-equivalent inference accuracy for MNIST handwritten-digit recognition and recurrent LSTM benchmarks, by using compensation techniques to finesse analog-memory challenges such as conductance drift and noise. We also project accuracy for Natural Language Processing (NLP) tasks performed with a state-of-art large Transformer-based model, BERT, when mapped onto an extended version of this same fundamental chip architecture.
Atsuya Okazaki, Pritish Narayanan, Stefano Ambrogio, Kohji Hosokawa, Hsinyu Tsai, Akiyo Nomura, Takeo Yasuda, Charles Mackin, Alexander M. Friz, Masatoshi Ishii, Yasuteru Kohda, Katie Spoon, An Chen 0002, Andrea Fasoli, Malte J. Rasch, Geoffrey W. Burr
ISCAS10
2021 Analysis of Effect of Weight Variation on SNN Chip with PCM-Refresh Method
Akiyo Nomura, Megumi Ito, Atsuya Okazaki, Masatoshi Ishii, SangBum Kim, Junka Okazawa, Kohji Hosokawa, Wilfried Haensch
Neural Process. Lett.4
2019 Performance Analysis of Spiking RBM with Measurement-Based Phase Change Memory Model
Masatoshi Ishii, Megumi Ito, Wanki Kim, SangBum Kim, Akiyo Nomura, Atsuya Okazaki, Junka Okazawa, Kohji Hosokawa, M. J. BrightSky, Wilfried Haensch
ICONIP (5)1
2019 Training Large-Scale Spiking Neural Networks on Multi-core Neuromorphic System Using Backpropagation
Megumi Ito, Malte J. Rasch, Masatoshi Ishii, Atsuya Okazaki, SangBum Kim, Junka Okazawa, Akiyo Nomura, Kohji Hosokawa, Wilfried Haensch
ICONIP (3)3
2018 NVM Weight Variation Impact on Analog Spiking Neural Network Chip
Akiyo Nomura, Megumi Ito, Atsuya Okazaki, Masatoshi Ishii, SangBum Kim, Junka Okazawa, Kohji Hosokawa, Wilfried Haensch
ICONIP (7)4
2012 Reflective photoplethysmography sensor with ring-shaped photodiode
abstract
This paper reports a fabrication process of pulse wave sensor and detection result of pulse wave using the ring-shaped sensor, which has ring-shaped photodiode. An LED is located at the center of a photodiode. This structure achieved noninvasive detection of the pulse wave and does not have the directivity. Therefore, the ring-shaped sensor is expected to reduce noises of body motion (motion artifacts). We evaluated the characteristics of the photodiode and detected the pulse wave using the fabricated sensor. As a result, the ring-shaped sensor can detect the pulse wave. Furthermore, detection of the pulse wave can be applied to measure the pulse wave velocity (PWV).
Masahide Kano, Masatoshi Ishii, Kensuke Kanda, Takayuki Fujita, Kazusuke Maenaka, Kazuo Kasai, Kohei Higuchi
SMC2