Andrew S. Cassidy

dblp:77/2302 · also Andrew Cassidy · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
1since 2021 · last 2023
0000-0001-7305-4198ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 9 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Emerging computing paradigms · 40% Hardware accelerators and domain-specific architectures · 22% Electronic design automation · 16%
Artificial intelligence
1 paper
Speech recognition and synthesis · 100%

Topics — the 15 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Emerging computing paradigms
neuromorphic computing
0.732016
Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications · SC 2016
TrueNorth: Design and Tool Flow of a 65 mW 1 Million Neuron Programmable Neurosynaptic Chip · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Real-Time Scalable Cortical Computing at 46 Giga-Synaptic OPS/Watt with ~100× Speedup in Time-to-Solution and ~100, 000× Reduction in Energy-to-Solution · SC 2014
Emerging computing paradigms
neuromorphic hardware
0.522017
Always-On Speech Recognition Using TrueNorth, a Reconfigurable, Neurosynaptic Processor · IEEE Trans. Computers 2017
Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications · SC 2016
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
spiking neural network accelerator
0.312017
Always-On Speech Recognition Using TrueNorth, a Reconfigurable, Neurosynaptic Processor · IEEE Trans. Computers 2017
Emerging computing paradigms › neuromorphic computing
brain-inspired computing
0.212016
Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications · SC 2016
Hardware accelerators and domain-specific architectures
neural network mapping
0.212015
TrueNorth: Design and Tool Flow of a 65 mW 1 Million Neuron Programmable Neurosynaptic Chip · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Electronic design automation
physical design
0.212015
TrueNorth: Design and Tool Flow of a 65 mW 1 Million Neuron Programmable Neurosynaptic Chip · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Electronic design automation › physical design
placement
0.212015
TrueNorth: Design and Tool Flow of a 65 mW 1 Million Neuron Programmable Neurosynaptic Chip · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Processor architecture and microarchitecture
chip multiprocessor
0.112012
Beyond Amdahl's Law: An Objective Function That Links Multiprocessor Performance Gains to Delay and Energy · IEEE Trans. Computers 2012
Electronic design automation
design space exploration
0.112012
Beyond Amdahl's Law: An Objective Function That Links Multiprocessor Performance Gains to Delay and Energy · IEEE Trans. Computers 2012
Energy-efficient computing › power-performance tradeoff
energy-delay tradeoff
0.112012
Beyond Amdahl's Law: An Objective Function That Links Multiprocessor Performance Gains to Delay and Energy · IEEE Trans. Computers 2012
Processor architecture and microarchitecture
microarchitecture optimization
0.112012
Beyond Amdahl's Law: An Objective Function That Links Multiprocessor Performance Gains to Delay and Energy · IEEE Trans. Computers 2012
Energy-efficient computing › energy-efficient machine learning
energy-efficient neural network inference
0.112017
Always-On Speech Recognition Using TrueNorth, a Reconfigurable, Neurosynaptic Processor · IEEE Trans. Computers 2017
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable processor
0.112017
Always-On Speech Recognition Using TrueNorth, a Reconfigurable, Neurosynaptic Processor · IEEE Trans. Computers 2017
Hardware accelerators and domain-specific architectures › neural network hardware
brain-inspired computing accelerator
0.112014
Real-Time Scalable Cortical Computing at 46 Giga-Synaptic OPS/Watt with ~100× Speedup in Time-to-Solution and ~100, 000× Reduction in Energy-to-Solution · SC 2014
Energy-efficient computing
power management
0.112014
Real-Time Scalable Cortical Computing at 46 Giga-Synaptic OPS/Watt with ~100× Speedup in Time-to-Solution and ~100, 000× Reduction in Energy-to-Solution · SC 2014

Methods — techniques the papers use, named apart from their topics

deep neural network · 0.6audio feature extraction · 0.6software ecosystem · 0.2scalable systems · 0.2mixed asynchronous-synchronous circuit design · 0.2CAD placement tool adaptation · 0.2event-driven kernel · 0.2chip tiling · 0.2analytical modeling · 0.1amdahl's law · 0.1
YearPublicationVenuePosition
2023 IBM NorthPole Neural Inference Machine
Dharmendra S. Modha, Filipp Akopyan, Alexander Andreopoulos, Rathinakumar Appuswamy, John V. Arthur, Andrew S. Cassidy, Pallab Datta, Michael DeBole, Steven K. Esser, Carlos Ortega Otero, Jun Sawada, Brian Taba, Arnon Amir, Deepika Bablani, Peter J. Carlson, Myron Flickner, Rajamohan Gandhasri, Guillaume Garreau, Megumi Ito, Jennifer L. Klamo, Jeffrey A. Kusnitz, Nathaniel J. McClatchey, Jeffrey L. McKinstry, Yutaka Y. Nakamura, Tapan K. Nayak, William P. Risk, Kai Schleupen, Ben Shaw 0001, Jay Sivagnaname, Daniel F. Smith, Ignacio G. Terrizzano, Takanori Ueda
HCS6
2017 Always-On Speech Recognition Using TrueNorth, a Reconfigurable, Neurosynaptic Processor
abstract
Deep neural networks (DNN) have been shown to be very effective at solving challenging problems in several areas of computing, including vision, speech, and natural language processing. However, traditional platforms for implementing these DNNs are often very power hungry, which has lead to significant efforts in the development of configurable platforms capable of implementing these DNNs efficiently. One of these platforms, the IBM TrueNorth processor, has demonstrated very low operating power in performing visual computing and neural network classification tasks in real-time. The neuron computation, synaptic memory, and communication fabrics are all configurable, so that a wide range of network types and topologies can be mapped to TrueNorth. This reconfigurability translates into the capability to support a wide range of low-power functions in addition to feed-forward DNN classifiers, including for example, the audio processing functions presented here.In this work, we propose an end-to-end audio processing pipeline that is implemented entirely on a TrueNorth processor and designed to specifically leverage the highly-parallel, low-precision computing primitives TrueNorth offers. As part of this pipeline, we develop an audio feature extractor (LATTE) designed for implementation on TrueNorth, and explore the tradeoffs among several design variants in terms of accuracy, power, and performance. We customize the energy-efficient deep neuromorphic networks structures that our design utilizes as the classifier and show how classifier parameters can trade between power and accuracy. In addition to enabling a wide range of diverse functions, the reconfigurability of TrueNorth enables re-training and re-programming the system to satisfy varying energy, speed, area, and accuracy requirements. The resulting system's end-to-end power consumption can be as low as$14.43\text{mW}$, which would give up to 100 hours of continuous usage with button cell batteries (CR3023$1.5\; \text{Whr}$) or 450 hours with cellphone batteries (iPhone 6s$6.55\; \text{Whr}$).
Wei-Yu Tsai, Davis Barch, Andrew S. Cassidy, Michael DeBole, Alexander Andreopoulos, Bryan L. Jackson, Myron Flickner, John V. Arthur, Dharmendra S. Modha, Jack Sampson, Narayanan Vijaykrishnan
IEEE Trans. Computers3
2016 A low-power neurosynaptic implementation of Local Binary Patterns for texture analysis
abstract
We demonstrate how to map Local Binary Patterns (LBP), a class of leading feature extractors, onto a neuromorphic processor such as TrueNorth, a silicon expression of a non-von Neumann, low-power, spiking-based, brain-inspired processor. The application is presented in the form of a texture feature extractor that can process 8-bit grayscale video at 30fps. While consuming less than 140mW of power, this neuromorphic implementation provides a rotation and contrast insensitive characterization of texture, with similar accuracy as a standard von Neumann implementation of the same algorithm. The successful mapping of an important vision routine on a neuromorphic architecture is indicative of an alternative paradigm for addressing the von Neumann bottleneck, which is currently placing severe constraints on the processing speed, power consumption, reliability, scalability, programmability and mobility of vision algorithms. This also introduces a new methodology for the design of vision algorithms for power efficient, asynchronous, mobility-targeted applications.
Alexander Andreopoulos, Rodrigo Alvarez-Icaza, Andrew S. Cassidy, Myron Flickner
IJCNN3
2016 TrueHappiness: Neuromorphic emotion recognition on TrueNorth
abstract
We present an approach to constructing a neuromorphic device that responds to language input by producing neuron spikes in proportion to the strength of the appropriate positive or negative emotional response. Specifically, we perform a fine-grained sentiment analysis task with implementations on two different systems: one using conventional spiking neural network (SNN) simulators and the other one using IBM's Neurosynaptic System TrueNorth. Input words are projected into a high-dimensional semantic space and processed through a fully-connected neural network (FCNN) containing rectified linear units (ReLU) trained via backpropagation. After training, this FCNN is converted to a SNN by substituting the ReLUs with integrate-and-fire neurons. We show that there is practically no performance loss due to conversion to a spiking network on a sentiment analysis test set, i.e. correlations with human annotations differ by less than 0.02 between the original DNN and its spiking equivalent. Additionally, we show that the SNN generated with this technique can be mapped to existing neuromorphic hardware - in our case, the TrueNorth chip. Mapping to the chip involves 4-bit synaptic weight discretization and adjustment of the neuron thresholds. The resulting end-to-end system can take a user input, i.e. a word in a vocabulary of over 300,000 words, and estimate its sentiment on TrueNorth with a power consumption of approximately 50 μW.
Peter U. Diehl, Bruno U. Pedroni, Andrew S. Cassidy, Paul Merolla, Emre Neftci, Guido Zarrella
IJCNN3
2016 LATTE: Low-power Audio Transform with TrueNorth Ecosystem
abstract
With recent advances in silicon technology, previously intractable Deep Neural Network (DNN) solutions to complex visual, auditory, and other sensory perception problems are now practical for real-time, energy constrained systems. One such advancement is IBM's TrueNorth neurosynaptic processor, containing 1 million neurons and 256 million synapses, consuming 65mW of power, and capable of operating in real-time for a variety of applications. In this work, we explore how auditory features can be extracted on the TrueNorth processor using low numerical precision while maintaining algorithmic fidelity for DNN based spoken digit recognition on isolated words from the TIDIGITS dataset. Further, we show that our Low-power Audio Transform with TrueNorth Ecosystem (LATTE) is capable of achieving a 24× reduction in energy for feature extraction over a baseline FPGA implementation using standard MFCC audio features, while only incurring a 3 - 6% accuracy penalty.
Wei-Yu Tsai, Davis Barch, Andrew S. Cassidy, Michael DeBole, Alexander Andreopoulos, Bryan L. Jackson, Myron Flickner, Dharmendra S. Modha, Jack Sampson, Narayanan Vijaykrishnan
IJCNN3
2016 Real-time sensory information processing using the TrueNorth Neurosynaptic System
abstract
Summary form only given. The IBM TrueNorth (TN) Neurosynaptic System, is a chip multi processor with a tightly coupled processor/memory architecture, that results in energy efficient neurocomputing and it is a significant milestone to over 30 years of neuromorphic engineering! It comprises of 4096 cores each core with 65K of local memory (6T SRAM)-synapses- and 256 arithmetic logic units - neurons-that operate on a unary number representation and compute by counting up to a maximum of 19 bits. The cores are event-driven using custom asynchronous and synchronous logic, and they are globally connected through an asynchronous packet switched mesh network on chip (NOC). The chip development board, includes a Zyng Xilinx FPGA that does the housekeeping and provides support for standard communication support through an Ethernet UDP interface. The asynchronous Addressed Event Representation (AER) in the NOC is al so exposed to the user for connection to AER based peripherals through a packet with bundled data full duplex interface. The unary data values represented on the system buses can take on a wide variety of spatial and temporal encoding schemes. Pulse density coding (the number of events Ne represents a number N), thermometer coding, time-slot encoding, and stochastic encoding are examples. Additional low level interfaces are available for communicating directly with the TrueNorth chip to aid programming and parameter setting. A hierarchical, compositional programming language, Corelet, is available to aid the development of TN applications. IBM provides support and a development system as well as “Compass” a scalable simulator. The software environment runs under standard Linux installations (Red Hat, CentOS and Ubuntu) and has standard interfaces to Matlab and to Caffe that is employed to train deep neural network models. The TN architecture can be interfaced using native AER to a number of bio-inspired sensory devices developed over many years of neuromorphic engineering (silicon retinas and silicon cochleas). In addition the architecture is well suited for implementing deep neural networks with many applications in computer vision, speech recognition and language processing. In a sensory information processing system architecture one desires both pattern processing in space and time to extract features in symbolic sub-spaces as well as natural language processing to provide contextual and semantic information in the form of priors. In this paper we discuss results from ongoing experimental work on real-time sensory information processing using the TN architecture in three different areas (i) spatial pattern processing -computer vision(ii) temporal pattern processing -speech processing and recognition(iii) natural language processing -word similarity-. A real-time demonstration will be done at ISCAS 2016 using the TN system and neuromorphic event based sensors for audition (silicon cochlea) and vision (silicon retina).
Andreas G. Andreou, Andrew A. Dykman, Kate D. Fischl, Guillaume Garreau, Daniel R. Mendat, Garrick Orchard, Andrew S. Cassidy, Paul Merolla, John V. Arthur, Rodrigo Alvarez-Icaza, Bryan L. Jackson, Dharmendra S. Modha
ISCAS7
2016 Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications
abstract
Abstract not provided
Jun Sawada, Filipp Akopyan, Andrew S. Cassidy, Brian Taba, Michael DeBole, Pallab Datta, Rodrigo Alvarez-Icaza, Arnon Amir, John V. Arthur, Alexander Andreopoulos, Rathinakumar Appuswamy, Heinz Baier, Davis Barch, David J. Berg, Carmelo di Nolfo, Steven K. Esser, Myron Flickner, Thomas A. Horvath, Bryan L. Jackson, Jeffrey A. Kusnitz, Scott Lekuch, Michael Mastro, Timothy Melano, Paul Merolla, Steven E. Millman, Tapan K. Nayak, Norm Pass, Hartmut Penner, William P. Risk, Kai Schleupen, Ben Shaw 0001, Hayley Wu, Brian Giera, Adam Moody, T. Nathan Mundhenk, Brian Van Essen, Eric X. Wang, David P. Widemann, William E. Murphy, Jamie K. Infantolino, James A. Ross, Dale R. Shires, Manuel M. Vindiola, Raju Namburu, Dharmendra S. Modha
SC3
2015 Gibbs sampling with low-power spiking digital neurons
abstract
Restricted Boltzmann Machines and Deep Belief Networks have been successfully used in a wide variety of applications including image classification and speech recognition. Inference and learning in these algorithms uses a Markov Chain Monte Carlo procedure called Gibbs sampling. A sigmoidal function forms the kernel of this sampler which can be realized from the firing statistics of noisy integrate-and-fire neurons on a neuromorphic VLSI substrate. This paper demonstrates such an implementation on an array of digital spiking neurons with stochastic leak and threshold properties for inference tasks and presents some key performance metrics for such a hardware-based sampler in both the generative and discriminative contexts.
Srinjoy Das, Bruno U. Pedroni, Paul Merolla, John V. Arthur, Andrew S. Cassidy, Bryan L. Jackson, Dharmendra S. Modha, Gert Cauwenberghs, Kenneth Kreutz-Delgado
ISCAS5
2015 TrueNorth: Design and Tool Flow of a 65 mW 1 Million Neuron Programmable Neurosynaptic Chip
abstract
The new era of cognitive computing brings forth the grand challenge of developing systems capable of processing massive amounts of noisy multisensory data. This type of intelligent computing poses a set of constraints, including real-time operation, low-power consumption and scalability, which require a radical departure from conventional system design. Brain-inspired architectures offer tremendous promise in this area. To this end, we developed TrueNorth, a 65 mW real-time neurosynaptic processor that implements a non-von Neumann, low-power, highly-parallel, scalable, and defect-tolerant architecture. With 4096 neurosynaptic cores, the TrueNorth chip contains 1 million digital neurons and 256 million synapses tightly interconnected by an event-driven routing infrastructure. The fully digital 5.4 billion transistor implementation leverages existing CMOS scaling trends, while ensuring one-to-one correspondence between hardware and software. With such aggressive design metrics and the TrueNorth architecture breaking path with prevailing architectures, it is clear that conventional computer-aided design (CAD) tools could not be used for the design. As a result, we developed a novel design methodology that includes mixed asynchronous-synchronous circuits and a complete tool flow for building an event-driven, low-power neurosynaptic chip. The TrueNorth chip is fully configurable in terms of connectivity and neural parameters to allow custom configurations for a wide range of cognitive and sensory perception applications. To reduce the system's communication energy, we have adapted existing application-agnostic very large-scale integration CAD placement tools for mapping logical neural networks to the physical neurosynaptic core locations on the TrueNorth chips. With that, we have successfully demonstrated the use of TrueNorth-based systems in multiple applications, including visual object recognition, with higher performance and orders of magnitude lower power consumption than the same algorithms run on von Neumann architectures. The TrueNorth chip and its tool flow serve as building blocks for future cognitive systems, and give designers an opportunity to develop novel brain-inspired architectures and systems based on the knowledge obtained from this paper.
Filipp Akopyan, Jun Sawada, Andrew S. Cassidy, Rodrigo Alvarez-Icaza, John V. Arthur, Paul Merolla, Nabil Imam, Yutaka Y. Nakamura, Pallab Datta, Gi-Joon Nam, Brian Taba, Michael P. Beakes, Bernard Brezzo, Jente B. Kuang, Rajit Manohar, William P. Risk, Bryan L. Jackson, Dharmendra S. Modha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2014 Real-Time Scalable Cortical Computing at 46 Giga-Synaptic OPS/Watt with ~100× Speedup in Time-to-Solution and ~100, 000× Reduction in Energy-to-Solution
abstract
Drawing on neuroscience, we have developed a parallel, event-driven kernel for neurosynaptic computation, that is efficient with respect to computation, memory, and communication. Building on the previously demonstrated highly optimized software expression of the kernel, here, we demonstrate True North, a co-designed silicon expression of the kernel. True North achieves five orders of magnitude reduction in energy to-solution and two orders of magnitude speedup in time-to solution, when running computer vision applications and complex recurrent neural network simulations. Breaking path with the von Neumann architecture, True North is a 4,096 core, 1 million neuron, and 256 million synapse brain-inspired neurosynaptic processor, that consumes 65mW of power running at real-time and delivers performance of 46 Giga-Synaptic OPS/Watt. We demonstrate seamless tiling of True North chips into arrays, forming a foundation for cortex-like scalability. True North's unprecedented time-to-solution, energy-to-solution, size, scalability, and performance combined with the underlying flexibility of the kernel enable a broad range of cognitive applications.
Andrew S. Cassidy, Rodrigo Alvarez-Icaza, Filipp Akopyan, Jun Sawada, John V. Arthur, Paul Merolla, Pallab Datta, Marc González 0001, Brian Taba, Alexander Andreopoulos, Arnon Amir, Steven K. Esser, Jeffrey A. Kusnitz, Rathinakumar Appuswamy, Chuck Haymes, Bernard Brezzo, Roger Moussalli, Ralph Bellofatto, Christian W. Baks, Michael Mastro, Kai Schleupen, Charles E. Cox, Ken Inoue, Steven E. Millman, Nabil Imam, Emmett McQuinn, Yutaka Y. Nakamura, Ivan Vo, Chen Guok, Don Nguyen, Scott Lekuch, Sameh W. Asaad, Daniel J. Friedman, Bryan L. Jackson, Myron Flickner, William P. Risk, Rajit Manohar, Dharmendra S. Modha
SC1
2013 Cognitive computing programming paradigm: A Corelet Language for composing networks of neurosynaptic cores
abstract
Marching along the DARPA SyNAPSE roadmap, IBM unveils a trilogy of innovations towards the TrueNorth cognitive computing system inspired by the brain's function and efficiency. The sequential programming paradigm of the von Neumann architecture is wholly unsuited for TrueNorth. Therefore, as our main contribution, we develop a new programming paradigm that permits construction of complex cognitive algorithms and applications while being efficient for TrueNorth and effective for programmer productivity. The programming paradigm consists of (a) an abstraction for a TrueNorth program, named Corelet, for representing a network of neurosynaptic cores that encapsulates all details except external inputs and outputs; (b) an object-oriented Corelet Language for creating, composing, and decomposing corelets; (c) a Corelet Library that acts as an ever-growing repository of reusable corelets from which programmers compose new corelets; and (d) an end-to-end Corelet Laboratory that is a programming environment which integrates with the TrueNorth architectural simulator, Compass, to support all aspects of the programming cycle from design, through development, debugging, and up to deployment. The new paradigm seamlessly scales from a handful of synapses and neurons to networks of neurosynaptic cores of progressively increasing size and complexity. The utility of the new programming paradigm is underscored by the fact that we have designed and implemented more than 100 algorithms as corelets for TrueNorth in a very short time span.
Arnon Amir, Pallab Datta, William P. Risk, Andrew S. Cassidy, Jeffrey A. Kusnitz, Steven K. Esser, Alexander Andreopoulos, Theodore M. Wong, Myron Flickner, Rodrigo Alvarez-Icaza, Emmett McQuinn, Ben Shaw 0001, Norm Pass, Dharmendra S. Modha
IJCNN4
2013 Cognitive computing building block: A versatile and efficient digital neuron model for neurosynaptic cores
abstract
Marching along the DARPA SyNAPSE roadmap, IBM unveils a trilogy of innovations towards the TrueNorth cognitive computing system inspired by the brain's function and efficiency. Judiciously balancing the dual objectives of functional capability and implementation/operational cost, we develop a simple, digital, reconfigurable, versatile spiking neuron model that supports one-to-one equivalence between hardware and simulation and is implementable using only 1272 ASIC gates. Starting with the classic leaky integrate-and-fire neuron, we add: (a) configurable and reproducible stochasticity to the input, the state, and the output; (b) four leak modes that bias the internal state dynamics; (c) deterministic and stochastic thresholds; and (d) six reset modes for rich finite-state behavior. The model supports a wide variety of computational functions and neural codes. We capture 50+ neuron behaviors in a library for hierarchical composition of complex computations and behaviors. Although designed with cognitive algorithms and applications in mind, serendipitously, the neuron model can qualitatively replicate the 20 biologically-relevant behaviors of a dynamical neuron model.
Andrew S. Cassidy, Paul Merolla, John V. Arthur, Steven K. Esser, Bryan L. Jackson, Rodrigo Alvarez-Icaza, Pallab Datta, Jun Sawada, Theodore M. Wong, Vitaly Feldman, Arnon Amir, Daniel Ben Dayan Rubin, Filipp Akopyan, Emmett McQuinn, William P. Risk, Dharmendra S. Modha
IJCNN1
2013 Cognitive computing systems: Algorithms and applications for networks of neurosynaptic cores
abstract
Marching along the DARPA SyNAPSE roadmap, IBM unveils a trilogy of innovations towards the TrueNorth cognitive computing system inspired by the brain's function and efficiency. The non-von Neumann nature of the TrueNorth architecture necessitates a novel approach to efficient system design. To this end, we have developed a set of abstractions, algorithms, and applications that are natively efficient for TrueNorth. First, we developed repeatedly-used abstractions that span neural codes (such as binary, rate, population, and time-to-spike), long-range connectivity, and short-range connectivity. Second, we implemented ten algorithms that include convolution networks, spectral content estimators, liquid state machines, restricted Boltzmann machines, hidden Markov models, looming detection, temporal pattern matching, and various classifiers. Third, we demonstrate seven applications that include speaker recognition, music composer recognition, digit recognition, sequence prediction, collision avoidance, optical flow, and eye detection. Our results showcase the parallelism, versatility, rich connectivity, spatio-temporality, and multi-modality of the TrueNorth architecture as well as compositionality of the corelet programming paradigm and the flexibility of the underlying neuron model.
Steven K. Esser, Alexander Andreopoulos, Rathinakumar Appuswamy, Pallab Datta, Davis Barch, Arnon Amir, John V. Arthur, Andrew S. Cassidy, Myron Flickner, Paul Merolla, Shyamal Chandra, Nicola Basilico, Stefano Carpin, Thomas G. Zimmerman, Frank Zee, Rodrigo Alvarez-Icaza, Jeffrey A. Kusnitz, Theodore M. Wong, William P. Risk, Emmett McQuinn, Tapan K. Nayak, Raghavendra Singh, Dharmendra S. Modha
IJCNN8
2013 Design of silicon brains in the nano-CMOS era: Spiking neurons, learning synapses and neural architecture optimization
Andrew S. Cassidy, Julius Georgiou, Andreas G. Andreou
Neural Networks1
2012 Building block of a programmable neuromorphic substrate: A digital neurosynaptic core
abstract
The grand challenge of neuromorphic computation is to develop a flexible brain-inspired architecture capable of a wide array of real-time applications, while striving towards the ultra-low power consumption and compact size of biological neural systems. Toward this end, we fabricated a building block of a modular neuromorphic architecture, a neurosynaptic core. Our implementation consists of 256 integrate-and-fire neurons and a 1,024×256 SRAM crossbar memory for synapses that fits in 4.2mm2using a 45nm SOI process and consumes just 45pJ per spike. The core is fully configurable in terms of neuron parameters, axon types, and synapse states and its fully digital implementation achieves one-to-one correspondence with software simulation models. One-to-one correspondence allows us to introduce an abstract neural programming model for our chip, a contract guaranteeing that any application developed in software functions identically in hardware. This contract allows us to rapidly test and map applications from control, machine vision, and classification. To demonstrate, we present four test cases (i) a robot driving in a virtual environment, (ii) the classic game of pong, (iii) visual digit recognition and (iv) an autoassociative memory.
John V. Arthur, Paul Merolla, Filipp Akopyan, Rodrigo Alvarez-Icaza, Andrew S. Cassidy, Shyamal Chandra, Steven K. Esser, Nabil Imam, William P. Risk, Daniel Ben Dayan Rubin, Rajit Manohar, Dharmendra S. Modha
IJCNN5
2012 Beyond Amdahl's Law: An Objective Function That Links Multiprocessor Performance Gains to Delay and Energy
abstract
Beginning with Amdahl's law, we derive a general objective function that links parallel processing performance gains at the system level, to energy and delay in the subsystem microarchitecture structures. The objective function employs parameterized models of computation and communication to represent the characteristics of processors, memories, and communications networks. The interaction of the latter microarchitectural elements defines global system performance in terms of energy-delay cost. Following the derivation, we demonstrate its utility by applying it to the problem of Chip Multiprocessor (CMP) architecture exploration. Given a set of application and architectural parameters, we solve for the optimal CMP architecture for six different architectural optimization examples. We find the parameters that minimize the total system cost, defined by the objective function under the area constraint of a single die. The analytical formulation presented in this paper is general and offers the foundation for the quantitative and rapid evaluation of computer architectures under different constraints including that of single die area.
Andrew S. Cassidy, Andreas G. Andreou
IEEE Trans. Computers1
2011 A high-level analytical model for application specific CMP design exploration
abstract
We present a high-level analytical model for chip-multiprocessors (CMPs) that encompasses processors, memory, and communication in an area-constrained, global optimization process. Applying this analytical model to the design of a symmetric CMP for speech recognition, we demonstrate a methodology for estimating model parameters prior to design exploration. Then we present an automated approach for finding the optimal high-level CMP architecture. The result is the ability to find the allocation of silicon resources for each architectural element that maximizes overall system performance. This balances the performance gains from parallelism, processor microarchitecture, and cache memory with the energy-delay costs of computation and communication.
Andrew S. Cassidy, Haolang Zhou, Andreas G. Andreou
DATE1
2011 A combinational digital logic approach to STDP
abstract
Spike Timing Dependant Plasticity (STDP) is a biologically-based Hebbian reinforcement learning rule for the unsupervised training of synaptic weights in spiking neural networks. We present a low complexity synthetic implementation of STDP using basic combinational digital logic gates. This approach attains comparable results to more complex implementations while utilizing only a fraction of the area. We use our STDP approach to replicate the experimental results of a balanced excitation experiment.
Andrew S. Cassidy, Andreas G. Andreou, Julius Georgiou
ISCAS1
2011 Evaluating on-chip interconnects for low operating frequency silicon neuron arrays
abstract
We present a quantitative analysis of the limits of the time-multiplexed Address Event Representation (AER) bus for on-chip connectivity of silicon neuron arrays. In particular, we evaluate its potential to support high density and low power neural arrays operating in the subthreshold regime. Our analysis shows that due to low clock frequencies when operating in the subthreshold regime, the traditional single AER bus does not scale to large neural arrays. We find that a switched mesh network improves scalability, however, a crosspoint architecture overcomes the bandwidth limitations altogether. By trading off area for improved performance, it increases the number of neurons that can be supported in a single chip neural array.
Andrew S. Cassidy, Thomas S. Murray, Andreas G. Andreou, Julius Georgiou
ISCAS1
2009 Analyzing features for automatic age estimation on cross-sectional data
abstract
We develop an acoustic feature set for the estimation of a per-son’s age from a recorded speech signal. The baseline features are Mel-frequency cepstral coefficients (MFCCs) which are ex-tended by various prosodic features, pitch and formant frequen-cies. From experiments on the University of Florida Vocal Ag-ing Database we can draw different conclusions. On the one hand, adding prosodic, pitch and formant features to the MFCC baseline leads to relative reductions of the mean absolute error between 4-20%. Improvements are even larger when percep-tual age labels are taken as a reference. On the other hand, reasonable results with a mean absolute error in age estimation of about 12 years are already achieved using a simple gender-independent setup and MFCCs only. Future experiments will evaluate the robustness of the prosodic features against channel variability on other databases and investigate the differences be-tween perceptual and chronological age labels.
Werner Spiegl, Georg Stemmer, Eva Lasarcyk, Varada Kolhatkar, Andrew S. Cassidy, Blaise Potard, Stephen H. Shum, Young Chol Song, Puyang Xu, Peter Beyerlein, James D. Harnsberger, Elmar Nöth
INTERSPEECH5
2009 A Switched Capacitor Implementation of the Generalized Linear Integrate-and-fire Neuron
abstract
In this paper we present the circuits and simulation results for a silicon neuron which is based on a modified version of the Mihalas-Niebur neural model [1]. This silicon neuron produces 15 of the 20 known neural spiking and bursting behaviors. It has low complexity and reliable matching and can thus be easily integrated into more complex neuromorphic systems. Implemented in a 0.15um 1.5V CMOS process, each neuron consumes about 7.5nW of power at 1kHz and occupies an area of 70um by 70um.
Fopefolu O. Folowosele, Andre Harrison, Andrew S. Cassidy, Andreas G. Andreou, Ralph Etienne-Cummings, Stefan Mihalas, Ernst Niebur, Tara J. Hamilton
ISCAS3
2005 High-level modeling and simulation of single-chip programmable heterogeneous multiprocessors
abstract
Heterogeneous multiprocessing is the future of chip design with the potential for tens to hundreds of programmable elements on single chips within the next several years. These chips will have heterogeneous, programmable hardware elements that lead to different execution times for the same software executing on different resources as well as a mix of desktop-style and embedded-style software. They will also have a layer of programming across multiple programmable elements forming the basis of a new kind of programmable system which we refer to as a Programmable Heterogeneous Multiprocessor (PHM). Current modeling approaches use instruction set simulation for performance modeling, but this will become far too prohibitive in terms of simulation time for these larger designs. The fundamental question is what the next higher level of design will be. The high-level modeling, simulation and design required for these programmable systems poses unique challenges, representing a break from traditional hardware design. Programmable systems, including layered concurrent software executing via schedulers on concurrent hardware, are not characterizable with traditional component-based hierarchical composition approaches, including discrete event simulation. We describe the foundations of our layered approach to modeling and performance simulation of PHMs, showing an example design space of a network processor explored using our simulation approach.
JoAnn M. Paul, Donald E. Thomas, Andrew S. Cassidy
ACM Trans. Design Autom. Electr. Syst.3
2003 Layered, Multi-Threaded, High-Level Performance Design
Andrew S. Cassidy, JoAnn M. Paul, Donald E. Thomas
DATE1