EDBT 2026 Demo / reviewers in the wild / expert
Dharmendra S. Modha
dblp:67/1008
· DBLP profile ↗
49ranked-venue papers
12as first author
1since 2021 · last 2023
0009-0005-6714-8127ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-authorSystems, architecture and hardware · 17 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-authorComputer networks · 5 · 1 first-authorTheory of computation · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
14 papers |
Emerging computing paradigms · 47% Hardware accelerators and domain-specific architectures · 19% Memory systems · 10% | |
| Artificial intelligence
5 papers |
Efficient and distributed learning · 75% Deep learning architectures and training · 12% Learning theory · 5% | |
| Databases, data mining, and information retrieval
4 papers |
Data mining · 92% Recommender systems · 8% | |
| Theoretical computer science
6 papers |
Coding theory · 60% Computational complexity · 20% Information theory · 8% |
Topics — the 30 heaviest of 65, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Emerging computing paradigms
neuromorphic computing |
1.2 | 7 | 2016 | Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications · SC 2016 TrueNorth: Design and Tool Flow of a 65 mW 1 Million Neuron Programmable Neurosynaptic Chip · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 Backpropagation for Energy-Efficient Neuromorphic Computing · NIPS 2015 |
Emerging computing paradigms
neuromorphic hardware |
1.0 | 4 | 2017 | Always-On Speech Recognition Using TrueNorth, a Reconfigurable, Neurosynaptic Processor · IEEE Trans. Computers 2017 A Low Power, Fully Event-Based Gesture Recognition System · CVPR 2017 Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications · SC 2016 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
spiking neural network accelerator |
0.5 | 2 | 2017 | Always-On Speech Recognition Using TrueNorth, a Reconfigurable, Neurosynaptic Processor · IEEE Trans. Computers 2017 Backpropagation for Energy-Efficient Neuromorphic Computing · NIPS 2015 |
Machine learning › Efficient and distributed learning
model compression |
0.4 | 1 | 2020 | Learned Step Size quantization · ICLR 2020 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.4 | 1 | 2020 | Learned Step Size quantization · ICLR 2020 |
Machine learning › Efficient and distributed learning › model compression › quantization
quantization-aware training |
0.4 | 1 | 2020 | Learned Step Size quantization · ICLR 2020 |
Emerging computing paradigms › neuromorphic computing
brain-inspired computing |
0.2 | 1 | 2016 | Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications · SC 2016 |
Machine learning › Deep learning architectures and training
backpropagation |
0.2 | 1 | 2015 | Backpropagation for Energy-Efficient Neuromorphic Computing · NIPS 2015 |
Hardware accelerators and domain-specific architectures
neural network mapping |
0.2 | 1 | 2015 | TrueNorth: Design and Tool Flow of a 65 mW 1 Million Neuron Programmable Neurosynaptic Chip · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Electronic design automation
physical design |
0.2 | 1 | 2015 | TrueNorth: Design and Tool Flow of a 65 mW 1 Million Neuron Programmable Neurosynaptic Chip · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Electronic design automation › physical design
placement |
0.2 | 1 | 2015 | TrueNorth: Design and Tool Flow of a 65 mW 1 Million Neuron Programmable Neurosynaptic Chip · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Data mining › clustering
co-clustering |
0.2 | 4 | 2007 | A Generalized Maximum Entropy Approach to Bregman Co-clustering and Matrix Approximation · J. Mach. Learn. Res. 2007 Fully automatic cross-associations · KDD 2004 A generalized maximum entropy approach to bregman co-clustering and matrix approximation · KDD 2004 |
Data mining
clustering |
0.2 | 3 | 2007 | A Generalized Maximum Entropy Approach to Bregman Co-clustering and Matrix Approximation · J. Mach. Learn. Res. 2007 Fully automatic cross-associations · KDD 2004 A generalized maximum entropy approach to bregman co-clustering and matrix approximation · KDD 2004 |
Performance modeling and evaluation
simulation |
0.1 | 1 | 2012 | Compass: a scalable simulator for an architecture for cognitive computing · SC 2012 |
Memory systems
cache |
0.1 | 3 | 2005 | SARC: Sequential Prefetching in Adaptive Replacement Cache · USENIX ATC, General Track 2005 CAR: Clock with Adaptive Replacement · FAST 2004 ARC: A Self-Tuning, Low Overhead Replacement Cache · FAST 2003 |
Data mining
matrix approximation |
0.1 | 2 | 2007 | A Generalized Maximum Entropy Approach to Bregman Co-clustering and Matrix Approximation · J. Mach. Learn. Res. 2007 A generalized maximum entropy approach to bregman co-clustering and matrix approximation · KDD 2004 |
Coding theory › constrained coding › constrained systems
constrained block code |
0.1 | 3 | 2002 | Links between complexity theory and constrained block coding · IEEE Trans. Inf. Theory 2002 Art of constructing low-complexity encoders/decoders for constrained block codes · IEEE J. Sel. Areas Commun. 2001 Links Between Complexity Theory and Constrained Block Coding · CCC 2001 |
High-performance computing
scientific computing systems |
0.1 | 1 | 2009 | The cat is out of the bag: cortical simulations with 109 neurons, 1013 synapses · SC 2009 |
Memory systems › cache management
cache replacement |
0.1 | 2 | 2004 | CAR: Clock with Adaptive Replacement · FAST 2004 ARC: A Self-Tuning, Low Overhead Replacement Cache · FAST 2003 |
Interaction techniques and input › input sensing
gesture recognition |
0.1 | 1 | 2017 | A Low Power, Fully Event-Based Gesture Recognition System · CVPR 2017 |
Energy-efficient computing › energy-efficient machine learning
energy-efficient neural network inference |
0.1 | 1 | 2017 | Always-On Speech Recognition Using TrueNorth, a Reconfigurable, Neurosynaptic Processor · IEEE Trans. Computers 2017 |
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable processor |
0.1 | 1 | 2017 | Always-On Speech Recognition Using TrueNorth, a Reconfigurable, Neurosynaptic Processor · IEEE Trans. Computers 2017 |
Coding theory
constrained coding |
0.1 | 2 | 2001 | Art of constructing low-complexity encoders/decoders for constrained block codes · IEEE J. Sel. Areas Commun. 2001 Maximum transition run codes for generalized partial response channels · IEEE J. Sel. Areas Commun. 2001 |
Hardware accelerators and domain-specific architectures › neural network hardware
brain-inspired computing accelerator |
0.1 | 1 | 2014 | Real-Time Scalable Cortical Computing at 46 Giga-Synaptic OPS/Watt with ~100× Speedup in Time-to-Solution and ~100, 000× Reduction in Energy-to-Solution · SC 2014 |
Energy-efficient computing
power management |
0.1 | 1 | 2014 | Real-Time Scalable Cortical Computing at 46 Giga-Synaptic OPS/Watt with ~100× Speedup in Time-to-Solution and ~100, 000× Reduction in Energy-to-Solution · SC 2014 |
Memory systems › cache management › storage caching
non-volatile cache |
0.1 | 1 | 2005 | WOW: Wise Ordering for Writes - Combining Spatial and Temporal Locality in Non-Volatile Caches · FAST 2005 |
Memory systems › cache
prefetching |
0.1 | 1 | 2005 | SARC: Sequential Prefetching in Adaptive Replacement Cache · USENIX ATC, General Track 2005 |
Memory systems › cache › prefetching
sequential prefetching |
0.1 | 1 | 2005 | SARC: Sequential Prefetching in Adaptive Replacement Cache · USENIX ATC, General Track 2005 |
Memory systems › data locality
spatial and temporal locality |
0.1 | 1 | 2005 | WOW: Wise Ordering for Writes - Combining Spatial and Temporal Locality in Non-Volatile Caches · FAST 2005 |
Recommender systems › collaborative filtering › matrix factorization
boolean matrix factorization |
0.0 | 1 | 2004 | Fully automatic cross-associations · KDD 2004 |
Methods — techniques the papers use, named apart from their topics
spiking neural network · 0.6deep neural network · 0.6convolutional neural network · 0.6audio feature extraction · 0.6probability sampling · 0.4ensemble averaging · 0.4backpropagation · 0.4quantization · 0.4software ecosystem · 0.2scalable systems · 0.2CAD placement tool adaptation · 0.2maximum entropy · 0.1bregman divergence · 0.1finite-state transition diagrams · 0.1mutual information · 0.1information theory · 0.1parameter-free algorithm · 0.0information-theoretic criterion · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | IBM NorthPole Neural Inference Machine
Dharmendra S. Modha, Filipp Akopyan, Alexander Andreopoulos, Rathinakumar Appuswamy, John V. Arthur, Andrew S. Cassidy, Pallab Datta, Michael DeBole, Steven K. Esser, Carlos Ortega Otero, Jun Sawada, Brian Taba, Arnon Amir, Deepika Bablani, Peter J. Carlson, Myron Flickner, Rajamohan Gandhasri, Guillaume Garreau, Megumi Ito, Jennifer L. Klamo, Jeffrey A. Kusnitz, Nathaniel J. McClatchey, Jeffrey L. McKinstry, Yutaka Y. Nakamura, Tapan K. Nayak, William P. Risk, Kai Schleupen, Ben Shaw 0001, Jay Sivagnaname, Daniel F. Smith, Ignacio G. Terrizzano, Takanori Ueda |
HCS | 1 |
| 2020 | Learned Step Size quantization
Steven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy, Dharmendra S. Modha |
ICLR | 5 |
| 2017 | A Low Power, Fully Event-Based Gesture Recognition SystemabstractWe present the first gesture recognition system implemented end-to-end on event-based hardware, using a TrueNorth neurosynaptic processor to recognize hand gestures in real-time at low power from events streamed live by a Dynamic Vision Sensor (DVS). The biologically inspired DVS transmits data only when a pixel detects a change, unlike traditional frame-based cameras which sample every pixel at a fixed frame rate. This sparse, asynchronous data representation lets event-based cameras operate at much lower power than frame-based cameras. However, much of the energy efficiency is lost if, as in previous work, the event stream is interpreted by conventional synchronous processors. Here, for the first time, we process a live DVS event stream using TrueNorth, a natively event-based processor with 1 million spiking neurons. Configured here as a convolutional neural network (CNN), the TrueNorth chip identifies the onset of a gesture with a latency of 105 ms while consuming less than 200 mW. The CNN achieves 96.5% out-of-sample accuracy on a newly collected DVS dataset (DvsGesture) comprising 11 hand gesture categories from 29 subjects under 3 illumination conditions. Arnon Amir, Brian Taba, David J. Berg, Timothy Melano, Jeffrey L. McKinstry, Carmelo di Nolfo, Tapan K. Nayak, Alexander Andreopoulos, Guillaume Garreau, Marcela Mendoza, Jeffrey A. Kusnitz, Michael DeBole, Steven K. Esser, Tobi Delbruck, Myron Flickner, Dharmendra S. Modha |
CVPR | 16 |
| 2017 | Always-On Speech Recognition Using TrueNorth, a Reconfigurable, Neurosynaptic ProcessorabstractDeep neural networks (DNN) have been shown to be very effective at solving challenging problems in several areas of computing, including vision, speech, and natural language processing. However, traditional platforms for implementing these DNNs are often very power hungry, which has lead to significant efforts in the development of configurable platforms capable of implementing these DNNs efficiently. One of these platforms, the IBM TrueNorth processor, has demonstrated very low operating power in performing visual computing and neural network classification tasks in real-time. The neuron computation, synaptic memory, and communication fabrics are all configurable, so that a wide range of network types and topologies can be mapped to TrueNorth. This reconfigurability translates into the capability to support a wide range of low-power functions in addition to feed-forward DNN classifiers, including for example, the audio processing functions presented here.In this work, we propose an end-to-end audio processing pipeline that is implemented entirely on a TrueNorth processor and designed to specifically leverage the highly-parallel, low-precision computing primitives TrueNorth offers. As part of this pipeline, we develop an audio feature extractor (LATTE) designed for implementation on TrueNorth, and explore the tradeoffs among several design variants in terms of accuracy, power, and performance. We customize the energy-efficient deep neuromorphic networks structures that our design utilizes as the classifier and show how classifier parameters can trade between power and accuracy. In addition to enabling a wide range of diverse functions, the reconfigurability of TrueNorth enables re-training and re-programming the system to satisfy varying energy, speed, area, and accuracy requirements. The resulting system's end-to-end power consumption can be as low as$14.43\text{mW}$, which would give up to 100 hours of continuous usage with button cell batteries (CR3023$1.5\; \text{Whr}$) or 450 hours with cellphone batteries (iPhone 6s$6.55\; \text{Whr}$). Wei-Yu Tsai, Davis Barch, Andrew S. Cassidy, Michael DeBole, Alexander Andreopoulos, Bryan L. Jackson, Myron Flickner, John V. Arthur, Dharmendra S. Modha, Jack Sampson, Narayanan Vijaykrishnan |
IEEE Trans. Computers | 9 |
| 2016 | LATTE: Low-power Audio Transform with TrueNorth EcosystemabstractWith recent advances in silicon technology, previously intractable Deep Neural Network (DNN) solutions to complex visual, auditory, and other sensory perception problems are now practical for real-time, energy constrained systems. One such advancement is IBM's TrueNorth neurosynaptic processor, containing 1 million neurons and 256 million synapses, consuming 65mW of power, and capable of operating in real-time for a variety of applications. In this work, we explore how auditory features can be extracted on the TrueNorth processor using low numerical precision while maintaining algorithmic fidelity for DNN based spoken digit recognition on isolated words from the TIDIGITS dataset. Further, we show that our Low-power Audio Transform with TrueNorth Ecosystem (LATTE) is capable of achieving a 24× reduction in energy for feature extraction over a baseline FPGA implementation using standard MFCC audio features, while only incurring a 3 - 6% accuracy penalty. Wei-Yu Tsai, Davis Barch, Andrew S. Cassidy, Michael DeBole, Alexander Andreopoulos, Bryan L. Jackson, Myron Flickner, Dharmendra S. Modha, Jack Sampson, Narayanan Vijaykrishnan |
IJCNN | 8 |
| 2016 | Real-time sensory information processing using the TrueNorth Neurosynaptic SystemabstractSummary form only given. The IBM TrueNorth (TN) Neurosynaptic System, is a chip multi processor with a tightly coupled processor/memory architecture, that results in energy efficient neurocomputing and it is a significant milestone to over 30 years of neuromorphic engineering! It comprises of 4096 cores each core with 65K of local memory (6T SRAM)-synapses- and 256 arithmetic logic units - neurons-that operate on a unary number representation and compute by counting up to a maximum of 19 bits. The cores are event-driven using custom asynchronous and synchronous logic, and they are globally connected through an asynchronous packet switched mesh network on chip (NOC). The chip development board, includes a Zyng Xilinx FPGA that does the housekeeping and provides support for standard communication support through an Ethernet UDP interface. The asynchronous Addressed Event Representation (AER) in the NOC is al so exposed to the user for connection to AER based peripherals through a packet with bundled data full duplex interface. The unary data values represented on the system buses can take on a wide variety of spatial and temporal encoding schemes. Pulse density coding (the number of events Ne represents a number N), thermometer coding, time-slot encoding, and stochastic encoding are examples. Additional low level interfaces are available for communicating directly with the TrueNorth chip to aid programming and parameter setting. A hierarchical, compositional programming language, Corelet, is available to aid the development of TN applications. IBM provides support and a development system as well as “Compass” a scalable simulator. The software environment runs under standard Linux installations (Red Hat, CentOS and Ubuntu) and has standard interfaces to Matlab and to Caffe that is employed to train deep neural network models. The TN architecture can be interfaced using native AER to a number of bio-inspired sensory devices developed over many years of neuromorphic engineering (silicon retinas and silicon cochleas). In addition the architecture is well suited for implementing deep neural networks with many applications in computer vision, speech recognition and language processing. In a sensory information processing system architecture one desires both pattern processing in space and time to extract features in symbolic sub-spaces as well as natural language processing to provide contextual and semantic information in the form of priors. In this paper we discuss results from ongoing experimental work on real-time sensory information processing using the TN architecture in three different areas (i) spatial pattern processing -computer vision(ii) temporal pattern processing -speech processing and recognition(iii) natural language processing -word similarity-. A real-time demonstration will be done at ISCAS 2016 using the TN system and neuromorphic event based sensors for audition (silicon cochlea) and vision (silicon retina). Andreas G. Andreou, Andrew A. Dykman, Kate D. Fischl, Guillaume Garreau, Daniel R. Mendat, Garrick Orchard, Andrew S. Cassidy, Paul Merolla, John V. Arthur, Rodrigo Alvarez-Icaza, Bryan L. Jackson, Dharmendra S. Modha |
ISCAS | 12 |
| 2016 | Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applicationsabstractAbstract not provided Jun Sawada, Filipp Akopyan, Andrew S. Cassidy, Brian Taba, Michael DeBole, Pallab Datta, Rodrigo Alvarez-Icaza, Arnon Amir, John V. Arthur, Alexander Andreopoulos, Rathinakumar Appuswamy, Heinz Baier, Davis Barch, David J. Berg, Carmelo di Nolfo, Steven K. Esser, Myron Flickner, Thomas A. Horvath, Bryan L. Jackson, Jeffrey A. Kusnitz, Scott Lekuch, Michael Mastro, Timothy Melano, Paul Merolla, Steven E. Millman, Tapan K. Nayak, Norm Pass, Hartmut Penner, William P. Risk, Kai Schleupen, Ben Shaw 0001, Hayley Wu, Brian Giera, Adam Moody, T. Nathan Mundhenk, Brian Van Essen, Eric X. Wang, David P. Widemann, William E. Murphy, Jamie K. Infantolino, James A. Ross, Dale R. Shires, Manuel M. Vindiola, Raju Namburu, Dharmendra S. Modha |
SC | 46 |
| 2015 | Brain-Inspired ComputingabstractSummary form only given. I will describe a decade-long, multi-disciplinary, multi-institutional effort spanning neuroscience, supercomputing, and nanotechnology to build and demonstrate a brain-inspired computer and describe the architecture, programming model, and applications. I will also describe future efforts to build, literally, "brain-in-a-box". For more information, see: modha.org. Dharmendra S. Modha |
PACT | 1 |
| 2015 | Gibbs sampling with low-power spiking digital neuronsabstractRestricted Boltzmann Machines and Deep Belief Networks have been successfully used in a wide variety of applications including image classification and speech recognition. Inference and learning in these algorithms uses a Markov Chain Monte Carlo procedure called Gibbs sampling. A sigmoidal function forms the kernel of this sampler which can be realized from the firing statistics of noisy integrate-and-fire neurons on a neuromorphic VLSI substrate. This paper demonstrates such an implementation on an array of digital spiking neurons with stochastic leak and threshold properties for inference tasks and presents some key performance metrics for such a hardware-based sampler in both the generative and discriminative contexts. Srinjoy Das, Bruno U. Pedroni, Paul Merolla, John V. Arthur, Andrew S. Cassidy, Bryan L. Jackson, Dharmendra S. Modha, Gert Cauwenberghs, Kenneth Kreutz-Delgado |
ISCAS | 7 |
| 2015 | Backpropagation for Energy-Efficient Neuromorphic ComputingabstractSolving real world problems with embedded neural networks requires both training algorithms that achieve high performance and compatible hardware that runs in real time while remaining energy efficient. For the former, deep learning using backpropagation has recently achieved a string of successes across many domains and datasets. For the latter, neuromorphic chips that run spiking neural networks have recently achieved unprecedented energy efficiency. To bring these two advances together, we must first resolve the incompatibility between backpropagation, which uses continuous-output neurons and synaptic weights, and neuromorphic designs, which employ spiking neurons and discrete synapses. Our approach is to treat spikes and discrete synapses as continuous probabilities, which allows training the network using standard backpropagation. The trained network naturally maps to neuromorphic hardware by sampling the probabilities to create one or more networks, which are merged using ensemble averaging. To demonstrate, we trained a sparsely connected network that runs on the TrueNorth chip using the MNIST dataset. With a high performance network (ensemble of $64$), we achieve $99.42\%$ accuracy at $121 \mu$J per image, and with a high efficiency network (ensemble of $1$) we achieve $92.7\%$ accuracy at $0.408 \mu$J per image. Steven K. Esser, Rathinakumar Appuswamy, Paul Merolla, John V. Arthur, Dharmendra S. Modha |
NIPS | 5 |
| 2015 | TrueNorth: Design and Tool Flow of a 65 mW 1 Million Neuron Programmable Neurosynaptic ChipabstractThe new era of cognitive computing brings forth the grand challenge of developing systems capable of processing massive amounts of noisy multisensory data. This type of intelligent computing poses a set of constraints, including real-time operation, low-power consumption and scalability, which require a radical departure from conventional system design. Brain-inspired architectures offer tremendous promise in this area. To this end, we developed TrueNorth, a 65 mW real-time neurosynaptic processor that implements a non-von Neumann, low-power, highly-parallel, scalable, and defect-tolerant architecture. With 4096 neurosynaptic cores, the TrueNorth chip contains 1 million digital neurons and 256 million synapses tightly interconnected by an event-driven routing infrastructure. The fully digital 5.4 billion transistor implementation leverages existing CMOS scaling trends, while ensuring one-to-one correspondence between hardware and software. With such aggressive design metrics and the TrueNorth architecture breaking path with prevailing architectures, it is clear that conventional computer-aided design (CAD) tools could not be used for the design. As a result, we developed a novel design methodology that includes mixed asynchronous-synchronous circuits and a complete tool flow for building an event-driven, low-power neurosynaptic chip. The TrueNorth chip is fully configurable in terms of connectivity and neural parameters to allow custom configurations for a wide range of cognitive and sensory perception applications. To reduce the system's communication energy, we have adapted existing application-agnostic very large-scale integration CAD placement tools for mapping logical neural networks to the physical neurosynaptic core locations on the TrueNorth chips. With that, we have successfully demonstrated the use of TrueNorth-based systems in multiple applications, including visual object recognition, with higher performance and orders of magnitude lower power consumption than the same algorithms run on von Neumann architectures. The TrueNorth chip and its tool flow serve as building blocks for future cognitive systems, and give designers an opportunity to develop novel brain-inspired architectures and systems based on the knowledge obtained from this paper. Filipp Akopyan, Jun Sawada, Andrew S. Cassidy, Rodrigo Alvarez-Icaza, John V. Arthur, Paul Merolla, Nabil Imam, Yutaka Y. Nakamura, Pallab Datta, Gi-Joon Nam, Brian Taba, Michael P. Beakes, Bernard Brezzo, Jente B. Kuang, Rajit Manohar, William P. Risk, Bryan L. Jackson, Dharmendra S. Modha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 18 |
| 2014 | Real-Time Scalable Cortical Computing at 46 Giga-Synaptic OPS/Watt with ~100× Speedup in Time-to-Solution and ~100, 000× Reduction in Energy-to-SolutionabstractDrawing on neuroscience, we have developed a parallel, event-driven kernel for neurosynaptic computation, that is efficient with respect to computation, memory, and communication. Building on the previously demonstrated highly optimized software expression of the kernel, here, we demonstrate True North, a co-designed silicon expression of the kernel. True North achieves five orders of magnitude reduction in energy to-solution and two orders of magnitude speedup in time-to solution, when running computer vision applications and complex recurrent neural network simulations. Breaking path with the von Neumann architecture, True North is a 4,096 core, 1 million neuron, and 256 million synapse brain-inspired neurosynaptic processor, that consumes 65mW of power running at real-time and delivers performance of 46 Giga-Synaptic OPS/Watt. We demonstrate seamless tiling of True North chips into arrays, forming a foundation for cortex-like scalability. True North's unprecedented time-to-solution, energy-to-solution, size, scalability, and performance combined with the underlying flexibility of the kernel enable a broad range of cognitive applications. Andrew S. Cassidy, Rodrigo Alvarez-Icaza, Filipp Akopyan, Jun Sawada, John V. Arthur, Paul Merolla, Pallab Datta, Marc González 0001, Brian Taba, Alexander Andreopoulos, Arnon Amir, Steven K. Esser, Jeffrey A. Kusnitz, Rathinakumar Appuswamy, Chuck Haymes, Bernard Brezzo, Roger Moussalli, Ralph Bellofatto, Christian W. Baks, Michael Mastro, Kai Schleupen, Charles E. Cox, Ken Inoue, Steven E. Millman, Nabil Imam, Emmett McQuinn, Yutaka Y. Nakamura, Ivan Vo, Chen Guok, Don Nguyen, Scott Lekuch, Sameh W. Asaad, Daniel J. Friedman, Bryan L. Jackson, Myron Flickner, William P. Risk, Rajit Manohar, Dharmendra S. Modha |
SC | 38 |
| 2013 | Cognitive computing programming paradigm: A Corelet Language for composing networks of neurosynaptic coresabstractMarching along the DARPA SyNAPSE roadmap, IBM unveils a trilogy of innovations towards the TrueNorth cognitive computing system inspired by the brain's function and efficiency. The sequential programming paradigm of the von Neumann architecture is wholly unsuited for TrueNorth. Therefore, as our main contribution, we develop a new programming paradigm that permits construction of complex cognitive algorithms and applications while being efficient for TrueNorth and effective for programmer productivity. The programming paradigm consists of (a) an abstraction for a TrueNorth program, named Corelet, for representing a network of neurosynaptic cores that encapsulates all details except external inputs and outputs; (b) an object-oriented Corelet Language for creating, composing, and decomposing corelets; (c) a Corelet Library that acts as an ever-growing repository of reusable corelets from which programmers compose new corelets; and (d) an end-to-end Corelet Laboratory that is a programming environment which integrates with the TrueNorth architectural simulator, Compass, to support all aspects of the programming cycle from design, through development, debugging, and up to deployment. The new paradigm seamlessly scales from a handful of synapses and neurons to networks of neurosynaptic cores of progressively increasing size and complexity. The utility of the new programming paradigm is underscored by the fact that we have designed and implemented more than 100 algorithms as corelets for TrueNorth in a very short time span. Arnon Amir, Pallab Datta, William P. Risk, Andrew S. Cassidy, Jeffrey A. Kusnitz, Steven K. Esser, Alexander Andreopoulos, Theodore M. Wong, Myron Flickner, Rodrigo Alvarez-Icaza, Emmett McQuinn, Ben Shaw 0001, Norm Pass, Dharmendra S. Modha |
IJCNN | 14 |
| 2013 | Cognitive computing building block: A versatile and efficient digital neuron model for neurosynaptic coresabstractMarching along the DARPA SyNAPSE roadmap, IBM unveils a trilogy of innovations towards the TrueNorth cognitive computing system inspired by the brain's function and efficiency. Judiciously balancing the dual objectives of functional capability and implementation/operational cost, we develop a simple, digital, reconfigurable, versatile spiking neuron model that supports one-to-one equivalence between hardware and simulation and is implementable using only 1272 ASIC gates. Starting with the classic leaky integrate-and-fire neuron, we add: (a) configurable and reproducible stochasticity to the input, the state, and the output; (b) four leak modes that bias the internal state dynamics; (c) deterministic and stochastic thresholds; and (d) six reset modes for rich finite-state behavior. The model supports a wide variety of computational functions and neural codes. We capture 50+ neuron behaviors in a library for hierarchical composition of complex computations and behaviors. Although designed with cognitive algorithms and applications in mind, serendipitously, the neuron model can qualitatively replicate the 20 biologically-relevant behaviors of a dynamical neuron model. Andrew S. Cassidy, Paul Merolla, John V. Arthur, Steven K. Esser, Bryan L. Jackson, Rodrigo Alvarez-Icaza, Pallab Datta, Jun Sawada, Theodore M. Wong, Vitaly Feldman, Arnon Amir, Daniel Ben Dayan Rubin, Filipp Akopyan, Emmett McQuinn, William P. Risk, Dharmendra S. Modha |
IJCNN | 16 |
| 2013 | Cognitive computing systems: Algorithms and applications for networks of neurosynaptic coresabstractMarching along the DARPA SyNAPSE roadmap, IBM unveils a trilogy of innovations towards the TrueNorth cognitive computing system inspired by the brain's function and efficiency. The non-von Neumann nature of the TrueNorth architecture necessitates a novel approach to efficient system design. To this end, we have developed a set of abstractions, algorithms, and applications that are natively efficient for TrueNorth. First, we developed repeatedly-used abstractions that span neural codes (such as binary, rate, population, and time-to-spike), long-range connectivity, and short-range connectivity. Second, we implemented ten algorithms that include convolution networks, spectral content estimators, liquid state machines, restricted Boltzmann machines, hidden Markov models, looming detection, temporal pattern matching, and various classifiers. Third, we demonstrate seven applications that include speaker recognition, music composer recognition, digit recognition, sequence prediction, collision avoidance, optical flow, and eye detection. Our results showcase the parallelism, versatility, rich connectivity, spatio-temporality, and multi-modality of the TrueNorth architecture as well as compositionality of the corelet programming paradigm and the flexibility of the underlying neuron model. Steven K. Esser, Alexander Andreopoulos, Rathinakumar Appuswamy, Pallab Datta, Davis Barch, Arnon Amir, John V. Arthur, Andrew S. Cassidy, Myron Flickner, Paul Merolla, Shyamal Chandra, Nicola Basilico, Stefano Carpin, Thomas G. Zimmerman, Frank Zee, Rodrigo Alvarez-Icaza, Jeffrey A. Kusnitz, Theodore M. Wong, William P. Risk, Emmett McQuinn, Tapan K. Nayak, Raghavendra Singh, Dharmendra S. Modha |
IJCNN | 23 |
| 2013 | Nanoscale electronic synapses using phase change devicesabstractThe memory capacity, computational power, communication bandwidth, energy consumption, and physical size of the brain all tend to scale with the number of synapses, which outnumber neurons by a factor of 10,000. Although progress in cortical simulations using modern digital computers has been rapid, the essential disparity between the classical von Neumann computer architecture and the computational fabric of the nervous system makes large-scale simulations expensive, power hungry, and time consuming. Over the last three decades, CMOS-based neuromorphic implementations of “electronic cortex” have emerged as an energy efficient alternative for modeling neuronal behavior. However, the key ingredient for electronic implementation of any self-learning system—programmable, plastic Hebbian synapses scalable to biological densities—has remained elusive. We demonstrate the viability of implementing such electronic synapses using nanoscale phase change devices. We introduce novel programming schemes for modulation of device conductance to closely mimic the phenomenon of Spike Timing Dependent Plasticity (STDP) observed biologically, and verify through simulations that such plastic phase change devices should support simple correlative learning in networks of spiking neurons. Our devices, when arranged in a crossbar array architecture, could enable the development of synaptronic systems that approach the density (∼10 11 synapses per sq cm) and energy efficiency (consuming ∼1pJ per synaptic programming event) of the human brain. Bryan L. Jackson, Bipin Rajendran, Gregory S. Corrado, Matthew J. Breitwisch, Geoffrey W. Burr, Roger Cheek, Kailash Gopalakrishnan, Simone Raoux, Charles T. Rettner, Alvaro Padilla, Alejandro G. Schrott, Rohit S. Shenoy, Bülent N. Kurdi, Chung Hon Lam, Dharmendra S. Modha |
ACM J. Emerg. Technol. Comput. Syst. | 15 |
| 2012 | Building block of a programmable neuromorphic substrate: A digital neurosynaptic coreabstractThe grand challenge of neuromorphic computation is to develop a flexible brain-inspired architecture capable of a wide array of real-time applications, while striving towards the ultra-low power consumption and compact size of biological neural systems. Toward this end, we fabricated a building block of a modular neuromorphic architecture, a neurosynaptic core. Our implementation consists of 256 integrate-and-fire neurons and a 1,024×256 SRAM crossbar memory for synapses that fits in 4.2mm2using a 45nm SOI process and consumes just 45pJ per spike. The core is fully configurable in terms of neuron parameters, axon types, and synapse states and its fully digital implementation achieves one-to-one correspondence with software simulation models. One-to-one correspondence allows us to introduce an abstract neural programming model for our chip, a contract guaranteeing that any application developed in software functions identically in hardware. This contract allows us to rapidly test and map applications from control, machine vision, and classification. To demonstrate, we present four test cases (i) a robot driving in a virtual environment, (ii) the classic game of pong, (iii) visual digit recognition and (iv) an autoassociative memory. John V. Arthur, Paul Merolla, Filipp Akopyan, Rodrigo Alvarez-Icaza, Andrew S. Cassidy, Shyamal Chandra, Steven K. Esser, Nabil Imam, William P. Risk, Daniel Ben Dayan Rubin, Rajit Manohar, Dharmendra S. Modha |
IJCNN | 12 |
| 2012 | Compass: a scalable simulator for an architecture for cognitive computingabstractInspired by the function, power, and volume of the organic brain, we are developing TrueNorth, a novel modular, non-von Neumann, ultra-low power, compact architecture. TrueNorth consists of a scalable network of neurosynaptic cores, with each core containing neurons, dendrites, synapses, and axons. To set sail for TrueNorth, we developed Compass, a multi-threaded, massively parallel functional simulator and a parallel compiler that maps a network of long-distance pathways in the macaque monkey brain to TrueNorth. We demonstrate near-perfect weak scaling on a 16 rack IBM® Blue Gene®/Q (262144 CPUs, 256 TB memory), achieving an unprecedented scale of 256 million neurosynaptic cores containing 65 billion neurons and 16 trillion synapses running only 388x slower than real time with an average spiking rate of 8.1 Hz. By using emerging PGAS communication primitives, we also demonstrate 2x better real-time performance over MPI primitives on a 4 rack Blue Gene/P (16384 CPUs, 16 TB memory). Robert Preissl, Theodore M. Wong, Pallab Datta, Myron Flickner, Raghavendra Singh, Steven K. Esser, William P. Risk, Horst D. Simon, Dharmendra S. Modha |
SC | 9 |
| 2010 | Binding sparse spatiotemporal patterns in spiking computationabstractImagine a two-dimensional spatial array of detectors temporally driven via an unknown number of mutually overlapping, unknown patterns. One at a time, these patterns are randomly, partially, sparsely and repeatedly presented, superimposed with omnipresent noise. The challenge is to design a scheme for detecting and recalling these patterns in an unsupervised, online and computationally efficient fashion. As our main contribution, we propose a network of spiking neurons consisting of two reciprocally connected layers. The bottom layer receives stimulus from the detector array and serves as input/output. The top layer encodes, detects and recalls specific patterns. Feedforward projections are data-driven, bottom-up, and analytic, while feedback projections are model-driven, top-down, and synthetic. We judiciously select neuron dynamics and spike-timing dependent synaptic learning rules such that these feedforward and feedback views eventually converge to bind together the spatial extent of each pattern into a coherent, temporary assembly. We present simulations demonstrating that our system is able to detect repeating patterns in an input stream with an impressive degree of tolerance to noise and pattern characteristics. Steven K. Esser, Anthony Ndirango, Dharmendra S. Modha |
IJCNN | 3 |
| 2009 | Think Global, Act Local; Projectome Estimation with BlueMatter
Anthony J. Sherbondy, Robert F. Dougherty, Rajagopal Ananthanarayanan, Dharmendra S. Modha, Brian A. Wandell |
MICCAI (1) | 4 |
| 2009 | The cat is out of the bag: cortical simulations with 109 neurons, 1013 synapsesabstractIn the quest for cognitive computing, we have built a massively parallel cortical simulator, C2, that incorporates a number of innovations in computation, memory, and communication. Using C2 on LLNL's Dawn Blue Gene/P supercomputer with 147, 456 CPUs and 144 TB of main memory, we report two cortical simulations -- at unprecedented scale -- that effectively saturate the entire memory capacity and refresh it at least every simulated second. The first simulation consists of 1.6 billion neurons and 8.87 trillion synapses with experimentally-measured gray matter thalamocortical connectivity. The second simulation has 900 million neurons and 9 trillion synapses with probabilistic connectivity. We demonstrate nearly perfect weak scaling and attractive strong scaling. The simulations, which incorporate phenomenological spiking neurons, individual learning synapses, axonal delays, and dynamic synaptic channels, exceed the scale of the cat cortex, marking the dawn of a new era in the scale of cortical simulations. Rajagopal Ananthanarayanan, Steven K. Esser, Horst D. Simon, Dharmendra S. Modha |
SC | 4 |
| 2007 | Anatomy of a cortical simulatorabstractInsights into brain's high-level computational principles will lead to novel cognitive systems, computing architectures, programming paradigms, and numerous practical applications. An important step towards this end is the study of large networks of cortical spiking neurons. Rajagopal Ananthanarayanan, Dharmendra S. Modha |
SC | 2 |
| 2007 | A Generalized Maximum Entropy Approach to Bregman Co-clustering and Matrix Approximation
Arindam Banerjee 0001, Inderjit S. Dhillon, Joydeep Ghosh, Srujana Merugu, Dharmendra S. Modha |
J. Mach. Learn. Res. | 5 |
| 2006 | Making the Correct MistakesabstractWe propose a new sequential, adaptive, quadratic-time algorithm for variable-rate lossy compression of memoryless sources at a fixed distortion. The algorithm uses approximate pattern matching and is modeled after the Lempel-Ziv algorithm. As a key new idea, the algorithm uses lower mutual information to carefully select "good" codewords. For Bernoulli sources with Hamming distortion, we empirically demonstrate that the algorithm (a) discovers the optimal reproduction type, (b) leads to absence of multiple matches, and (c) seems to approach the rate-distortion coding rate. Based on empirical observations, we formulate two conjectures that could imply that the algorithm is asymptotically optimal for memoryless sources. Dharmendra S. Modha, Narayana P. Santhanam |
DCC | 1 |
| 2005 | WOW: Wise Ordering for Writes - Combining Spatial and Temporal Locality in Non-Volatile Caches
Binny S. Gill, Dharmendra S. Modha |
FAST | 2 |
| 2005 | SARC: Sequential Prefetching in Adaptive Replacement Cache
Binny S. Gill, Dharmendra S. Modha |
USENIX ATC, General Track | 2 |
| 2004 | CAR: Clock with Adaptive Replacement
Sorav Bansal, Dharmendra S. Modha |
FAST | 2 |
| 2004 | Finite-state rate-distortion for individual sequencesabstractWe introduce a class of lossy finite-state machines for lossy compression of an individual sequence drawn from a finite alphabet at a fixed distortion, and define a fundamental quantity finite-state rate-distortion that is an asymptotically attainable lower bound on the compression rate of any lossy finite-state machine. For Hamming distortion, we obtain a universal lower bound on the finite-state rate-distortion of any individual sequence. Dharmendra S. Modha, Daniela Pucci de Farias |
ISIT | 1 |
| 2004 | A generalized maximum entropy approach to bregman co-clustering and matrix approximationabstractCo-clustering is a powerful data mining technique with varied applications such as text clustering, microarray analysis and recommender systems. Recently, an information-theoretic co-clustering approach applicable to empirical joint probability distributions was proposed. In many situations, co-clustering of more general matrices is desired. In this paper, we present a substantially generalized co-clustering framework wherein any Bregman divergence can be used in the objective function, and various conditional expectation based constraints can be considered based on the statistics that need to be preserved. Analysis of the co-clustering problem leads to the minimum Bregman information principle, which generalizes the maximum entropy principle, and yields an elegant meta algorithm that is guaranteed to achieve local optimality. Our methodology yields new algorithms and also encompasses several previously known clustering and co-clustering algorithms based on alternate minimization. Arindam Banerjee 0001, Inderjit S. Dhillon, Joydeep Ghosh, Srujana Merugu, Dharmendra S. Modha |
KDD | 5 |
| 2004 | Fully automatic cross-associationsabstractLarge, sparse binary matrices arise in numerous data mining applications, such as the analysis of market baskets, web graphs, social networks, co-citations, as well as information retrieval, collaborative filtering, sparse matrix reordering, etc. Virtually all popular methods for the analysis of such matrices---e.g., k-means clustering, METIS graph partitioning, SVD/PCA and frequent itemset mining---require the user to specify various parameters, such as the number of clusters, number of principal components, number of partitions, and "support." Choosing suitable values for such parameters is a challenging problem.Cross-association is a joint decomposition of a binary matrix into disjoint row and column groups such that the rectangular intersections of groups are homogeneous. Starting from first principles, we furnish a clear, information-theoretic criterion to choose a good cross-association as well as its parameters, namely, the number of row and column groups. We provide scalable algorithms to approach the optimal. Our algorithm is parameter-free, and requires no user intervention. In practice it scales linearly with the problem size, and is thus applicable to very large matrices. Finally, we present experiments on multiple synthetic and real-life datasets, where our method gives high-quality, intuitive results. Deepayan Chakrabarti, Spiros Papadimitriou, Dharmendra S. Modha, Christos Faloutsos |
KDD | 3 |
| 2003 | Codelet Parsing: Quadratic-time, Sequential, Adaptive Algorithms for Lossy CompressionabstractThe codelet parsing algorithms were proposed for lossy compression. The algorithms sequentially parse a given source sequence into phrases, say, sourcelets, and map each sourcelet to a distorted phrase, say, a codelet, such that the pre-letter distortion between the two phrases does not exceed the desired distortion. The algorithms adaptively maintain a codebook, and do not require any a priori knowledge of the source statistics. The algorithms use approximate string matching and, at each epoch, carefully select one of the many approximately matching codewords to balance between the code rates in the current epoch versus the code rate from the resulting codebooks in future epochs. The algorithms are quadratic-time in the length of the source sequence and output a distorted sequence that can be naturally losslessly compressed using the Lempel-Ziv algorithm. Dharmendra S. Modha |
DCC | 1 |
| 2003 | ARC: A Self-Tuning, Low Overhead Replacement Cache
Nimrod Megiddo, Dharmendra S. Modha |
FAST | 2 |
| 2003 | CacheCOW: QoS for Storage System Caches
Pawan Goyal 0001, Divyesh Jadav, Dharmendra S. Modha, Renu Tewari |
IWQoS | 3 |
| 2003 | Information-theoretic co-clusteringabstractTwo-dimensional contingency or co-occurrence tables arise frequently in important applications such as text, web-log and market-basket data analysis. A basic problem in contingency table analysis is co-clustering: simultaneous clustering of the rows and columns. A novel theoretical formulation views the contingency table as an empirical joint probability distribution of two discrete random variables and poses the co-clustering problem as an optimization problem in information theory---the optimal co-clustering maximizes the mutual information between the clustered random variables subject to constraints on the number of row and column clusters. We present an innovative co-clustering algorithm that monotonically increases the preserved mutual information by intertwining both the row and column clusterings at all stages. Using the practical example of simultaneous word-document clustering, we demonstrate that our algorithm works well in practice, especially in the presence of sparsity and high-dimensionality. Inderjit S. Dhillon, Subramanyam Mallela, Dharmendra S. Modha |
KDD | 3 |
| 2003 | CacheCOW: providing QoS for storage system cachesabstractManaged hosting and enterprise wide resource consolidation trends are increasingly leading to sharing of storage resources across multiple classes, corresponding to different applications/customers, each with a possibly different Quality of Service (QoS) requirement. To enable a storage system to meet diverse QoS requirements, we present two algorithms for dynamically allocating cache space among multiple classes of workloads. Our algorithms dynamically adapt the cache space allocated to each class in response to the observed response time, the temporal locality of reference, and the arrival pattern for each class. Using trace driven simulations collected from large storage system installations, we experimentally demonstrate that the algorithms not only meet the QoS requirements, but also increase the throughput by achieving a higher hit rate whenever feasible. Pawan Goyal 0001, Dharmendra S. Modha, Renu Tewari |
SIGMETRICS | 2 |
| 2003 | Feature Weighting in k-Means Clustering
Dharmendra S. Modha, W. Scott Spangler |
Mach. Learn. | 1 |
| 2002 | Links between complexity theory and constrained block codingabstractThe goal of this paper is to establish links between computational complexity theory and the theory and practice of constrained block coding. In particular, the complexities of several fundamental problems in constrained block coding are precisely classified in terms of the existing complexity-theoretic structure. One type of problem studied is that of designing encoder and decoder circuits using minimum or approximately minimum hardware; for our purposes, an "input" to this problem is (i) a deterministic, irreducible finite-state transition diagram (DIF) defining a set of constrained binary sequences, and (ii) a desired rate p:q. Several of these minimum-encoder and minimum-decoder problems are shown to be NP-hard, and more interestingly some are shown to be complete in the second and third levels of the polynomial hierarchy. Another fundamental problem is that of computing the maximum rate of a block code; that is, given a DIF and a codeword length q, find the maximum p such that a rate p:q block code exists for the constraint defined by the DIF. This problem is shown to be NP/sup #P/-complete. Although it is not known whether NP/sup #P/ contains problems of super-polynomial complexity, it lies "higher" in the complexity-class structure than NP in the sense that it is possible, given current knowledge, that NP/sup #P/ contains problems of super-polynomial complexity even if P=NP. Another question studied is whether maximum rate block codes can always be implemented by encoders and decoders of polynomial size. The answer to this question is shown to be closely related to whether the class #P lies "lower" in the complexity-class structure than currently believed-a proof of either answer would have major implications in complexity theory. Larry J. Stockmeyer, Dharmendra S. Modha |
IEEE Trans. Inf. Theory | 2 |
| 2001 | Links Between Complexity Theory and Constrained Block CodingabstractThe goal of this paper is to establish links between computational complexity theory and the theory and practice of constrained block coding. The complexities of several fundamental problems in constrained block coding are shown to be complete in various classes of the existing complexity-theoretic structure. The results include (relatively rare) /spl Sigma//sub 2//sup p/-, /spl Sigma//sub 3//sup p/, and NP/sup PP/-completeness results. Two types of problems are considered: (1) the problem of designing encoder and decoder circuits using minimum or approximately minimum hardware for a given constraint and a given rate; (2) computing the maximum rate of a block code for a given constraint and codeword length. In both cases, a constraint is specified by a deterministic finite state transition diagram. Another question studied is whether maximum-rate block codes can always be implemented by encoders and decoders of polynomial size. The answer to this question is shown to be closely related to the complexity of PP. Larry J. Stockmeyer, Dharmendra S. Modha |
CCC | 2 |
| 2001 | Extended bit-filling and LDPC code designabstractCampello, Modha and Rajagopalan (see Proc. Int. Conf. Communications (ICC), Helsinki, Finland, 2001) proposed a simple-to-implement heuristic, namely, bit-filling, for constructing high rate and high girth LDPC codes. In the present work, we extend bit-filling, and demonstrate that the extended algorithm produces better codes, that is, codes with higher rate/girth and good bit error rate performance. We demonstrate the positive effect of girth on bit error rate performance. Jorge Campello, Dharmendra S. Modha |
GLOBECOM | 2 |
| 2001 | Designing LDPC codes using bit-fillingabstractBipartite graphs of bit nodes and parity check nodes arise as Tanner graphs corresponding to low density parity check codes. Given graph parameters such as the number of check nodes, the maximum check-degree, the bit-degree, and the girth, we consider the problem of constructing bipartite graphs with the largest number of bit nodes, that is, the highest rate. We propose a simple-to-implement heuristic bit-filling algorithm for this problem. As a benchmark, our algorithm yields codes better or comparable to those in MacKay (1999). Jorge Campello, Dharmendra S. Modha, Sridhar Rajagopalan |
ICC | 2 |
| 2001 | Maximum transition run codes for generalized partial response channelsabstractA new twins constraint for maximum transition run (MTR) codes is introduced to eliminate quasi-catastrophic error propagation in sequence detectors for generalized partial response channels with spectral nulls both at dc and at the Nyquist frequency. Two variants of the twins constraint that depend on whether the generalized partial response detector trellis is unconstrained or j-constrained are studied. Deterministic finite-state transition diagrams that present the twins constraint are specified, and the capacity of the new class of MTR constraints is computed. The connection between (G,I) constraints and MTR(j) constraints is clarified. Code design methodologies that are based on look-ahead coding in combination with violation detection/substitution as well as on state splitting are used to obtain several specific constructions of high-rate MTR codes. Roy D. Cideciyan, Evangelos Eleftheriou, Brian H. Marcus, Dharmendra S. Modha |
IEEE J. Sel. Areas Commun. | 4 |
| 2001 | Art of constructing low-complexity encoders/decoders for constrained block codesabstractA rate p : q block encoder is a dataword-to-codeword assignment from 2/sup p/ p-bit datawords to 2/sup p/ q-bit codewords, and the corresponding block decoder is the inverse of the encoder. When designing block encoders/decoders for constrained systems, often, more than 2/sup p/ codewords are available. In this paper, as our main contribution, we propose efficient heuristic computer algorithms to eliminate the excess codewords and to construct low hardware complexity block encoders/decoders. For (0,4/4) and (0,3/6) PRML constraints, block encoders/decoders generated using the proposed algorithms are comparable in complexity to human-generated encoders/decoders, but are significantly simpler than lexicographical encoders/decoders. Dharmendra S. Modha, Brian H. Marcus |
IEEE J. Sel. Areas Commun. | 1 |
| 2001 | Concept Decompositions for Large Sparse Text Data Using Clustering
Inderjit S. Dhillon, Dharmendra S. Modha |
Mach. Learn. | 2 |
| 2000 | Reversible arithmetic coding for quantum data compressionabstractWe study the problem of compressing a block of symbols (a block quantum state) emitted by a memoryless quantum Bernoulli source. We present a simple-to-implement quantum algorithm for projecting, with high probability, the block quantum state onto the typical subspace spanned by the lending eigenstates of its density matrix. We propose a fixed-rate quantum Shannon-Fano code to compress the projected block quantum state using a per-symbol code rate that is slightly higher than the von Neumann (1955) entropy limit. Finally, we propose quantum arithmetic codes to efficiently implement quantum Shannon-Fano (1948) codes. Our arithmetic encoder and decoder have a cubic circuit and a cubic computational complexity in the block size. Both the encoder and decoder are quantum-mechanical inverses of each other, and constitute an elegant example of reversible quantum computation. Isaac L. Chuang, Dharmendra S. Modha |
IEEE Trans. Inf. Theory | 2 |
| 1998 | Prequential and Cross-Validated Regression Estimation
Dharmendra S. Modha, Elias Masry |
Mach. Learn. | 1 |
| 1998 | Memory-Universal Prediction of Stationary Random ProcessesabstractWe consider the problem of one-step-ahead prediction of a real-valued, stationary, strongly mixing random process (Xi)/sub i=-/spl infin///sup /spl infin//. The best mean-square predictor of X/sub 0/ is its conditional mean given the entire infinite past (X/sub i/)/sub i=-/spl infin///sup -1/. Given a sequence of observations X/sub 1/, X/sub 2/, X/sub N/, we propose estimators for the conditional mean based on sequences of parametric models of increasing memory and of increasing dimension, for example, neural networks and Legendre polynomials. The proposed estimators select both the model memory and the model dimension, in a data-driven fashion, by minimizing certain complexity regularized least squares criteria. When the underlying predictor function has a finite memory, we establish that the proposed estimators are memory-universal: the proposed estimators, which do not know the true memory, deliver the same statistical performance (rates of integrated mean-squared error) as that delivered by estimators that know the true memory. Furthermore, when the underlying predictor function does not have a finite memory, we establish that the estimator based on Legendre polynomials is consistent. Dharmendra S. Modha, Elias Masry |
IEEE Trans. Inf. Theory | 1 |
| 1996 | Rate of Convergence in Density Estimation Using Neural NetworksabstractGiven N i.i.d. observations {Xi}Ni=1 taking values in a compact subset of Rd, such that p* denotes their common probability density function, we estimate p* from an exponential family of densities based on single hidden layer sigmoidal networks using a certain minimum complexity density estimation scheme. Assuming that p* possesses a certain exponential representation, we establish a rate of convergence, independent of the dimension d, for the expected Hellinger distance between the proposed minimum complexity density estimator and the true underlying density p*. Dharmendra S. Modha, Elias Masry |
Neural Comput. | 1 |
| 1996 | Minimum complexity regression estimation with weakly dependent observationsabstractThe minimum complexity regression estimation framework (Barron, 1991; Barron and Cover, 1991 and Rissanen, 1989) is a general data-driven methodology for estimating a regression function from a given list of parametric models using independent and identically distributed (i.i.d.) observations. We extend Barron's regression estimation framework to m-dependent observations and to strongly mixing observations. In particular, we propose abstract minimum complexity regression estimators for dependent observations, which may be adapted to a particular list of parametric models, and establish upper bounds on the statistical risks of the proposed estimators in terms of certain deterministic indices of resolvability. Assuming that the regression function satisfies a certain Fourier-transform-type representation, we examine minimum complexity regression estimators adapted to a list of parametric models based on neural networks and by using the upper bounds for the abstract estimators, we establish rates of convergence for the statistical risks of these estimators. Also, as a key tool, we extend the classical Bernstein inequality from i.i.d. random variables to m-dependent processes and to strongly mixing processes. Dharmendra S. Modha, Elias Masry |
IEEE Trans. Inf. Theory | 1 |
| 1994 | A learning law for density estimationabstractProbability density functions are estimated by an exponential family of densities based on multilayer feedforward networks. The role of the multilayer feedforward networks, in the proposed estimator, is to approximate the logarithm of the probability density functions. The method of maximum likelihood is used, as the main contribution, to derive an unsupervised backpropagation learning law to estimate the probability density functions. Computer simulation results demonstrating the use of the derived learning law are presented. Dharmendra S. Modha, Yeshaiahu Fainman |
IEEE Trans. Neural Networks | 1 |