Rajit Manohar

dblp:m/RajitManohar · DBLP profile ↗
← Back
52ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0001-8211-6602ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 37 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 4Theory of computation · 4 · 3 first-authorSecurity and privacy · 2Human-computer interaction and ubiquitous computing · 2Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Maelstrom: A Logic Synthesis Technique for Asynchronous Circuits
abstract
A new synthesis method and corresponding open-source tool, called Maelstrom, is introduced, that synthesizes CHP programs into asynchronous circuits. The method is agnostic to circuit family, and produces circuits that show significant improvements over the state-of-the-art synthesis techniques for asynchronous circuits in terms of energy, delay and area. The method also supports different datapath implementations and communication protocols. Pre-layout SPICE simulations of generated netlists of several CHP programs in a 65nm node indicate significant performance benefits over the current state of the art. Maelstrom has also been used to successfully synthesize and fabricate a chip from an abstract high-level functional description. This represents a qualitative improvement in logic synthesis of asynchronous logic from behavioral descriptions.
Karthi Srinivasan, Rajit Manohar
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 PipeLink: A Pipelined Resource Sharing System for Dataflow High-Level Synthesis
abstract
Dynamically scheduled high-level synthesis (HLS) is an approach to HLS that maps programs into dataflow circuits. These circuits use distributed control for communication and therefore can be automatically pipelined. However, pipelined function call is challenging due to the absence of centralized control, thus general software programs cannot be well-supported by the HLS tools. Traditional solutions to this problem impose restrictions on pipelining when accessing the shared functions. We present PipeLink, a modular synthesis method that decomposes the synthesis into compilation stage and linking stage, which enables the fully-pipelined access to shared functions in dataflow circuits. We develop a complete HLS engine using this approach and use asynchronous circuits as the target implementation. Our engine supports pipelined access to shared function units, as well as pipelined memory access in a unified fashion. Compared to existing HLS tools, PipeLink results in 20X reduction in energy and 1.56 X improvement in throughput.
Rui Li 0103, Lincoln Berkley, Rajit Manohar
DAC3
2025 PipeLink: A Pipelined Resource Sharing System for Dataflow High-Level Synthesis
abstract
High-level synthesis (HLS) has received significant interest in hardware accelerator design due to its ability to reduce design time and make hardware design more accessible to software developers [3]. HLS methodologies can be categorized into two main approaches: statically scheduled HLS [2, 4] and dataflow HLS [5, 6]. Statically scheduled HLS requires precise operator delay information at synthesis time to generate optimized schedules, which often lead to suboptimal performance under dynamic workloads. In contrast, dataflow HLS maps programs into dataflow circuits composed of concurrent, independent elements that communicate using local control protocols (e.g., handshake or ready/valid). While dataflow HLS enables automatic pipelining without requiring global control, the lack of centralized coordination complicates efficient resource sharing.
Rui Li 0103, Rajit Manohar
FPGA2
2025 Dataflow-Specific Algorithms for Resource-Constrained Scheduling and Memory Design
abstract
We introduce the Weighted Red-Blue Pebble Game, an extension of the classic red-blue pebble game with weighted operation costs. This weighted formulation enables constant-factor analysis of highly resource-constrained systems with bounded fast memory, unlimited slow memory, and strict energy and power constraints.
Abhishek Bhattacharjee, Quanquan C. Liu, Rajit Manohar, Raghavendra Pradyumna Pothukuchi, Muhammed Ugur
SPAA3
2024 Designing an Energy-Efficient Fully-Asynchronous Deep Learning Convolution Engine
abstract
In the face of exponential growth in semiconductor energy usage, there is a significant push towards highly energy-efficient microelectronics design. While the traditional circuit designs typically employ clocks to synchronize the computing operations, these circuits incur significant performance and energy overheads due to their data-independent worst-case operation and complex clock tree networks. In this paper, we explore asynchronous or clockless techniques where clocks are replaced by request, acknowledge handshaking signals. To quantify the potential energy and performance gains of asynchronous logic, we design a highly energy -efficient asynchronous deep learning convolution engine, which uses 87 % of total DL accelerator energy. Our asynchronous design shows 5.06x lower energy and 5.09 x lower delay than the synchronous one.
Mattia Vezzoli, Lukas Nel, Kshitij Bhardwaj, Rajit Manohar, Maya B. Gokhale
DATE4
2023 SCALO: An Accelerator-Rich Distributed System for Scalable Brain-Computer Interfacing
abstract
SCALO is the first distributed brain-computer interface (BCI) consisting of multiple wireless-networked implants placed on different brain regions. SCALO unlocks new treatment options for debilitating neurological disorders and new research into brain-wide network behavior. Achieving the fast and low-power communication necessary for real-time processing has historically restricted BCIs to single brain sites. SCALO also adheres to tight power constraints, but enables fast distributed processing. Central to SCALO's efficiency is its realization as a full stack distributed system of brain implants with accelerator-rich compute. SCALO balances modular system layering with aggressive cross-layer hardware-software co-design to integrate compute, networking, and storage. The result is a lesson in designing energy-efficient networked distributed systems with hardware accelerators from the ground up.
Karthik Sriram, Raghavendra Pradyumna Pothukuchi, Michal Gerasimiuk, Muhammed Ugur, Oliver Ye, Rajit Manohar, Anurag Khandelwal, Abhishek Bhattacharjee
ISCA6
2022 SPRoute 2.0: A detailed-routability-driven deterministic parallel global router with soft capacity
abstract
Global routing has become more challenging due to advancements in the technology node and the ever-increasing size of chips. Global routing needs to generate routing guides such that (1) routability of detailed routing is considered and (2) the routing is deterministic and fast. In this paper, we firstly introduce soft capacity which reserves routing space for detailed routing based on the pin density and Rectangular Uniform wire Density (RUDY). Second, we propose a deterministic parallelization approach that partitions the netlist into batches and then bulk-synchronously maze-routes a single batch of nets. The advantage of this approach is that it guarantees determinacy without requiring the nets running in parallel to be disjoint, thus guaranteeing scalability. We then design a scheduler that mitigates the load imbalance and livelock issues in this bulk synchronous execution model. We implement SPRoute 2.0 with the proposed methodology. The experimental results show that SPRoute 2.0 generates good quality of results with 43% fewer shorts, 14% fewer DRCs and a 7.4X speedup over a state-of-the-art global router on the ICCAD2019 contest benchmarks.
Jiayuan He 0003, Udit Agarwal, Yihang Yang, Rajit Manohar, Keshav Pingali
ASP-DAC4
2022 HALO: A Flexible and Low Power Processing Fabric for Brain-Computer Interfaces
abstract
The FDA warns against overheating cellular tissue beyond 1°C ➔ 15-40mW
Abhishek Bhattacharjee, Rajit Manohar
HCS2
2022 General Approach to Asynchronous Circuits Simulation Using Synchronous FPGAs
abstract
Using field-programmable gate arrays (FPGAs) for software and hardware verification and development is a standard step in the digital application-specific integrated circuits (ASICs) design flow. However, asynchronous FPGAs are not available on the market and commercially available FPGAs provide support only for synchronous circuits. Although a lot of research effort has been undertaken in order to use synchronous FPGA to map asynchronous circuits, proposed solutions are typically lacking automation and target-specific circuit style and sometimes specific FPGA vendor. In this work, we present an automated solution for asynchronous circuits mapping onto the synchronous FPGAs. We build a synchronous model of the original asynchronous circuit based on the event-driven simulation concepts. The proposed approach supports a wide range of circuit styles, including those with various timing assumptions and complex circuitry structures incompatible with the standard synchronous flow. We avoid using vendor-specific features so that the model we generate can be implemented on any commercially available FPGA. We provide an extensive evaluation of our solution and demonstrate that our approach results in a speedup factor of$1.3\times 10^{5}$against an asynchronous circuit simulator,$2.8\times 10^{4}$against commercial digital simulators, and is 16.5 times slower than the expected performance of the original asynchronous circuit implemented as an ASIC.
Ruslan Dashkin, Rajit Manohar
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Zerializer: towards zero-copy serialization
abstract
Achieving zero-copy I/O has long been an important goal in the networking community. However, data serialization obviates the benefits of zero-copy I/O, because it requires the CPU to read, transform, and write message data, resulting in additional memory copies between the real object instances and the contiguous socket buffer. Therefore, we argue for offloading serialization logic to the DMA path via specialized hardware. We propose an initial hardware design for such an accelerator, and give preliminary evidence of its feasibility and expected benefits.
Adam Wolnikowski, Stephen Ibanez, Jonathan Stone 0004, Changhoon Kim, Rajit Manohar, Robert Soulé
HotOS5
2020 Dali: A Gridded Cell Placement Flow
abstract
Asynchronous Very-Large-Scale-Integration (VLSI) has several potential benefits over its synchronous counterparts, such as reduced power consumption, elastic pipelining, and robustness to variations. However, the lack of electronic design automation (EDA) support for asynchronous circuits, especially physical layout automation tools, largely limits their adoption. To tackle this challenge, we propose a gridded cell layout methodology for asynchronous circuits, in which the cell height and cell width can be any integer multiple of two grid values. The gridded cell approach combines the shape regularity of standard cells with the size flexibility of custom design, and thus achieves a better space utilization ratio and lower wire-length for asynchronous designs. We present the algorithms and our implementation of Dali, a gridded cell placer, that consists of an analytical global placer, a forward-backward legalizer, an N/P-well legalizer, and a power grid router. We show that the gridded cell placement approach reduces area by 15% without impacting the routability of the design. We have also used Dali to tape out a chip in a 65nm process technology, demonstrating that our placer generates design-rule clean placement.
Yihang Yang, Jiayuan He 0003, Rajit Manohar
ICCAD3
2020 Hardware-Software Co-Design for Brain-Computer Interfaces
abstract
Brain-computer interfaces (BCIs) offer avenues to treat neurological disorders, shed light on brain function, and interface the brain with the digital world. Their wider adoption rests, however, on achieving adequate real-time performance, meeting stringent power constraints, and adhering to FDA-mandated safety requirements for chronic implantation. BCIs have, to date, been designed as custom ASICs for specific diseases or for specific tasks in specific brain regions. General-purpose architectures that can be used to treat multiple diseases and enable various computational tasks are needed for wider BCI adoption, but the conventional wisdom is that such systems cannot meet necessary performance and power constraints. We present HALO (Hardware Architecture for LOw-power BCIs), a general-purpose architecture for implantable BCIs. HALO enables tasks such as treatment of disorders (e.g., epilepsy, movement disorders), and records/processes data for studies that advance our understanding of the brain. We use electrophysiological data from the motor cortex of a non-human primate to determine how to decompose HALO's computational capabilities into hardware building blocks. We simplify, prune, and share these building blocks to judiciously use available hardware resources while enabling many modes of brain-computer interaction. The result is a configurable heterogeneous array of hardware processing elements (PEs). The PEs are configured by a low-power RISC-V micro-controller into signal processing pipelines that meet the target performance and power constraints necessary to deploy HALO widely and safely.
Ioannis Karageorgos, Karthik Sriram, Ján Veselý, Marc Powell, David A. Borton, Rajit Manohar, Abhishek Bhattacharjee
ISCA7
2020 Exact Timing Analysis for Asynchronous Circuits With Multiple Periods
abstract
The timing properties of asynchronous circuits can be summarized using cyclic graphs that capture max-delay constraints between signal transitions. There are many results on the timing analysis problem, but they all make various simplifying assumptions on the connectivity properties of the underlying timing graph. Most results provide approximate timing characteristics, with a few providing exact results on the circuit's timing behavior. In this article, we provide results that exactly characterize the timing properties for a more general class of max-delay constraints. We show that the circuit can be partitioned into regions with different periodicities, and provide an efficient algorithm to compute all the periods of the system.
Rajit Manohar
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 SPRoute: A Scalable Parallel Negotiation-based Global Router
abstract
The complexity of global routing increases rapidly as chip designs grow larger. In many global routers, maze routing is the most time-consuming stage. One way to reduce its runtime is parallelization. Existing parallel maze routers work either by identifying and routing independent nets or by partitioning the chip area into non-overlapping regions. In this paper, we describe a scalable parallel global router called SPRoute that initially exploits net-level parallelism, automatically lowers the parallelism when livelock is identified, and finally switches to fine-grain parallelism to guarantee convergence. We evaluate SPRoute on a 28-core machine on the ISPD 2008 global routing contest benchmark suite. It achieves an average speedup of 11.5 with a wirelength penalty of 0.6% on overflow-free benchmarks, and an average speedup of 4.5 with a total overflow penalty of 7% on hard-to-route benchmarks over sequential SPRoute. Compared to FastRoute 4.1, SPRoute achieves an average speedup of 11.0 and 3.1 on overflow-free benchmarks and hard-to-route benchmarks, respectively.
Jiayuan He 0003, Martin Burtscher, Rajit Manohar, Keshav Pingali
ICCAD3
2019 Braindrop: A Mixed-Signal Neuromorphic Architecture With a Dynamical Systems-Based Programming Model
abstract
Braindrop is the first neuromorphic system designed to be programmed at a high level of abstraction. Previous neuromorphic systems were programmed at the neurosynaptic level and required expert knowledge of the hardware to use. In stark contrast, Braindrop's computations are specified as coupled nonlinear dynamical systems and synthesized to the hardware by an automated procedure. This procedure not only leverages Braindrop's fabric of subthreshold analog circuits as dynamic computational primitives but also compensates for their mismatched and temperature-sensitive responses at the network level. Thus, a clean abstraction is presented to the user. Fabricated in a 28-nm FDSOI process, Braindrop integrates 4096 neurons in 0.65 mm2. Two innovations-sparse encoding through analog spatial convolution and weighted spike-rate summation though digital accumulative thinning-cut digital traffic drastically, reducing the energy Braindrop consumes per equivalent synaptic operation to 381 fJ for typical network configurations.
Alexander Neckar, Sam Fok, Ben Varkey Benjamin, Terrence C. Stewart, Nick N. Oza, Aaron Voelker, Chris Eliasmith, Rajit Manohar, Kwabena Boahen 0001
Proc. IEEE8
2019 QDI Constant-Time Counters
abstract
Counters are a generally useful circuit that appear in many contexts. Because of this, the design space for clocked counters has been widely explored. However, the same cannot be said for robust clockless counters. To resolve this, we designed an array of constant response time counters using the most robust clockless logic family, quasi-delay-insensitive (QDI) circuits. We compare our designs to their closest QDI counterparts from the literature, showing significant improvements in design quality metrics including transistor count, energy per operation, frequency, and latency in a 28-nm process. We also compare our designs against prototypical synchronous counters generated by commercial logic synthesis tools.
Ned Bingham, Rajit Manohar
IEEE Trans. Very Large Scale Integr. Syst.2
2019 Self-Timed Adaptive Digit-Serial Addition
abstract
A fundamental operator in modern computational systems has a multitude of highly optimized implementations. In general-purpose systems, the width has grown to 128 bits, putting pressure on designers to use sophisticated carry lookahead or tree adders to maintain throughput while sacrificing area and energy. However, the typical workload mostly exercises the lower 10-15 bits. This leaves many devices on and unused during normal operation, reducing the overall performance. We hypothesize that bit- or digit-serial implementations for arbitrary-length streams represent an opportunity to decrease the overall energy usage while increasing the throughput/area efficiency of the system and verify this hypothesis by constructing an asynchronous digit-serial adder for comparison against its bit-parallel counterparts.
Ned Bingham, Rajit Manohar
IEEE Trans. Very Large Scale Integr. Syst.2
2019 Operation-Dependent Frequency Scaling Using Desynchronization
abstract
Asynchronous circuits are inherently more robust than their synchronous counterparts. Desynchronization is a way to obtain asynchronous circuits from a synchronous specification using standard design tools while improving circuit for variation tolerance, electromagnetic interference, and resulting in similar area, delay, and energy as the synchronous baseline. This paper proposes a novel operation-dependent desynchronization technique, which desynchronizes the circuit and improves performance beyond the limits of synchronous design. We perform a case study of our proposed technique on RISC-V rocket core and show significant improvement in performance with minimal power and area overheads.
Nitish Kumar Srivastava, Rajit Manohar
IEEE Trans. Very Large Scale Integr. Syst.2
2018 Exact Timing Analysis for Asynchronous Systems
abstract
Analyzing the timing properties of asynchronous systems is essential for characterizing their performance and power. Previous work on timing showed that such systems under and-causality and fixed delay exhibit periodicity properties. We give a different graph-based rigorous proof of the exact timing behavior of more general classes of such systems, and conclude their exact periodicity property, where each of the signal transition will occur with the same period after finite occurrences. We established our results under weaker assumption about system connectivity/topology, and this paper provides the theoretical foundation, for the exact periodicity property to be applied and exploited in circuits containing a combination of synchronous and asynchronous components. We provide simulation-based results for several typical asynchronous circuit topologies to quantify this time period in practical circuits. We also provide an extension of our analysis and methods to the case of bounded delay systems. A key result that is a consequence of our analysis is that asynchronous circuits can be integrated with synchronous logic via a metastability-free interface, thereby eliminating the high-overhead synchronizers when an asynchronous circuit is fully surrounded by synchronous logic.
Wenmian Hua, Rajit Manohar
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2017 Accelerating Face Detection on Programmable SoC Using C-Based Synthesis
Nitish Kumar Srivastava, Steve Dai, Rajit Manohar, Zhiru Zhang
FPGA3
2017 DeepRecon: Dynamically reconfigurable architecture for accelerating deep neural networks
abstract
Deep learning models are computationally expensive and their performance depends strongly on the underlying hardware platform. General purpose compute platforms such as GPUs have been widely used for implementing deep learning techniques. However, with the advent of emerging application domains such as internet of things, developments of custom integrated circuits capable of efficiently implementing deep learning models with low power and form factor are in high demand. In this paper we analyze both the computation and communication costs of common deep networks. We propose a reconfigurable architecture that efficiently utilizes computational and storage resources for accelerating deep learning techniques without loss of algorithmic accuracy.
Tayyar Rzayev, Saber Moradi, David H. Albonesi, Rajit Manohar
IJCNN4
2017 Toolbox for exploration of energy-efficient event processors for human-computer interaction
abstract
The advent of high speed input sensor and display technologies and the drive for faster interactive response suggests that human-computer interaction (HCI) task processing deadlines of a few milliseconds or less may be required in future handheld devices. At the same time, users will expect the same, if not better, battery life than today's devices under these more stringent response requirements. In this paper, we present a toolbox for exploring the design space of HCI event processors. We first describe the simulation platform for interactive environments that runs mobile user interface code with inputs recorded from human users. We validate it against a hardware platform from prior work. Given system-level constraints on latency, we demonstrate how this toolbox can be used to design a custom heterogeneous event processor that maximizes battery life. We show that our toolbox can pick design points that are 1.5-2.5× more energy-efficient than general-purpose big. LITTLE architectures.
Tayyar Rzayev, David H. Albonesi, François Guimbretière, Rajit Manohar, Jaeyeon Kihm
ISPASS4
2017 On Using Time Without Clocks via Zigzag Causality
abstract
Even in the absence of clocks, time bounds on the duration of actions enable the use of time for distributed coordination. This paper initiates an investigation of coordination in such a setting. A new communication structure called a zigzag pattern is introduced, and is shown to guarantee bounds on the relative timing of events in this clockless model. Indeed, zigzag patterns are shown to be necessary and sufficient for establishing that events occur in a manner that satisfies prescribed bounds. We capture when a process can know that an appropriate zigzag pattern exists, and use this to provide necessary and sufficient conditions for timed coordination of events using a full-information protocol in the clockless model.
Asa Dan, Rajit Manohar, Yoram Moses
PODC2
2015 Design of a QDI asynchronous AER serializer/deserializer link in 180nm for event-based sensors for robotic applications
abstract
On-chip AER serialization is a required step for the successful integration of event-driven neuromorphic devices on complex robotic platforms. We propose an architecture and its implementation on AMS 180nm technology, synthesised with the Quasi-Delay Insensitive design principles and with minimum timing assumptions. The serialization of 19 bits completes in 70ns with an average power consumption of 2.34mW.
Giovanni Rovere, Chiara Bartolozzi, Nabil Imam, Rajit Manohar
ISCAS4
2015 Preventing glitches and short circuits in high-level self-timed chip specifications
abstract
Self-timed chip designs are commonly specified in a high-level message-passing language called CHP. This language is closely related to Hoare's CSP except it admits erroneous behavior due to the necessary limitations of efficient hardware implementations. For example, two processes sending on the same channel at the same time causes glitches and short circuits in the physical chip implementation. If a CHP program maintains certain invariants, such as only one process is sending on any given channel at a time, it can guarantee an error-free execution that behaves much like a CSP program would. In this paper, we present an inferable effect system for ensuring that these invariants hold, drawing from model-checking methodologies while exploiting language-usage patterns and domain-specific specializations to achieve efficiency. This analysis is sound, and is even complete for the common subset of CHP programs without data-sensitive synchronization. We have implemented the analysis and demonstrated that it scales to validate even microprocessors.
Stephen Longfield Jr., Brittany Nkounkou, Rajit Manohar, Ross Tate
PLDI3
2015 TrueNorth: Design and Tool Flow of a 65 mW 1 Million Neuron Programmable Neurosynaptic Chip
abstract
The new era of cognitive computing brings forth the grand challenge of developing systems capable of processing massive amounts of noisy multisensory data. This type of intelligent computing poses a set of constraints, including real-time operation, low-power consumption and scalability, which require a radical departure from conventional system design. Brain-inspired architectures offer tremendous promise in this area. To this end, we developed TrueNorth, a 65 mW real-time neurosynaptic processor that implements a non-von Neumann, low-power, highly-parallel, scalable, and defect-tolerant architecture. With 4096 neurosynaptic cores, the TrueNorth chip contains 1 million digital neurons and 256 million synapses tightly interconnected by an event-driven routing infrastructure. The fully digital 5.4 billion transistor implementation leverages existing CMOS scaling trends, while ensuring one-to-one correspondence between hardware and software. With such aggressive design metrics and the TrueNorth architecture breaking path with prevailing architectures, it is clear that conventional computer-aided design (CAD) tools could not be used for the design. As a result, we developed a novel design methodology that includes mixed asynchronous-synchronous circuits and a complete tool flow for building an event-driven, low-power neurosynaptic chip. The TrueNorth chip is fully configurable in terms of connectivity and neural parameters to allow custom configurations for a wide range of cognitive and sensory perception applications. To reduce the system's communication energy, we have adapted existing application-agnostic very large-scale integration CAD placement tools for mapping logical neural networks to the physical neurosynaptic core locations on the TrueNorth chips. With that, we have successfully demonstrated the use of TrueNorth-based systems in multiple applications, including visual object recognition, with higher performance and orders of magnitude lower power consumption than the same algorithms run on von Neumann architectures. The TrueNorth chip and its tool flow serve as building blocks for future cognitive systems, and give designers an opportunity to develop novel brain-inspired architectures and systems based on the knowledge obtained from this paper.
Filipp Akopyan, Jun Sawada, Andrew S. Cassidy, Rodrigo Alvarez-Icaza, John V. Arthur, Paul Merolla, Nabil Imam, Yutaka Y. Nakamura, Pallab Datta, Gi-Joon Nam, Brian Taba, Michael P. Beakes, Bernard Brezzo, Jente B. Kuang, Rajit Manohar, William P. Risk, Bryan L. Jackson, Dharmendra S. Modha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.15
2014 Using asymmetric cores to reduce power consumption for interactive devices with bi-stable displays
abstract
Low power "helper" cores have been increasingly included on application processors to accomplish low intensity tasks such as music playing and motion sensing with minimum energy consumption. Recently, Guimbretière et al. [1] demonstrated that such helper cores can also be used to execute simple user interface tasks. We revisit this approach by implementing a similar system on an off-the-shelf application processor (TI OMAP4). Our study shows that in the case of high event rate interactions (pen inking and virtual keyboard), significant battery life gains (×1.7 and ×2.3 respectively) can be achieved with the helper core executing the interface. Having the helper core only dis-patch input events incurs a 18% penalty relative to the maximum savings rate, but allows for simplified deployment since it merely requires a change in toolkit infrastructure.
Jaeyeon Kihm, François Guimbretière, Julia Karl, Rajit Manohar
CHI4
2014 Removing concurrency for rapid functional verification
abstract
VLSI systems are commonly specified using sequential executable functional specifications, but implemented in a highly concurrent manner. Alhough the methods to transform between the sequential specification and concurrent implementation have been well-studied, there are still substantial difficulties in verifying that the concurrent implementation corresponds to the sequential specification after low-level optimization. The majority of methods for doing this verification have focused on strong semantic models for reasoning about systems and their specifications, but these models can add significant unnecessary complexity. In this paper, we explore a weak but effective method for reasoning about implementation relations. We show how a sequential embedding of a concurrent program can be generated, and how that embedding can be used to dramatically reduce the reachable state space of the verification problem while maintaining the semantic model of interest.
Stephen Longfield Jr., Rajit Manohar
ICCAD2
2014 Real-Time Scalable Cortical Computing at 46 Giga-Synaptic OPS/Watt with ~100× Speedup in Time-to-Solution and ~100, 000× Reduction in Energy-to-Solution
abstract
Drawing on neuroscience, we have developed a parallel, event-driven kernel for neurosynaptic computation, that is efficient with respect to computation, memory, and communication. Building on the previously demonstrated highly optimized software expression of the kernel, here, we demonstrate True North, a co-designed silicon expression of the kernel. True North achieves five orders of magnitude reduction in energy to-solution and two orders of magnitude speedup in time-to solution, when running computer vision applications and complex recurrent neural network simulations. Breaking path with the von Neumann architecture, True North is a 4,096 core, 1 million neuron, and 256 million synapse brain-inspired neurosynaptic processor, that consumes 65mW of power running at real-time and delivers performance of 46 Giga-Synaptic OPS/Watt. We demonstrate seamless tiling of True North chips into arrays, forming a foundation for cortex-like scalability. True North's unprecedented time-to-solution, energy-to-solution, size, scalability, and performance combined with the underlying flexibility of the kernel enable a broad range of cognitive applications.
Andrew S. Cassidy, Rodrigo Alvarez-Icaza, Filipp Akopyan, Jun Sawada, John V. Arthur, Paul Merolla, Pallab Datta, Marc González 0001, Brian Taba, Alexander Andreopoulos, Arnon Amir, Steven K. Esser, Jeffrey A. Kusnitz, Rathinakumar Appuswamy, Chuck Haymes, Bernard Brezzo, Roger Moussalli, Ralph Bellofatto, Christian W. Baks, Michael Mastro, Kai Schleupen, Charles E. Cox, Ken Inoue, Steven E. Millman, Nabil Imam, Emmett McQuinn, Yutaka Y. Nakamura, Ivan Vo, Chen Guok, Don Nguyen, Scott Lekuch, Sameh W. Asaad, Daniel J. Friedman, Bryan L. Jackson, Myron Flickner, William P. Risk, Rajit Manohar, Dharmendra S. Modha
SC37
2014 An asymmetric dual-processor architecture for low-power information appliances
abstract
As users become increasingly conscious of their energy footprint—either to improve battery life or to respect the environment—improved energy efficiency of systems has gained in importance. This is especially important in the context of information appliances such as e-book readers that are meant to replace books, since their energy efficiency impacts how long the appliance can be used on a single charge of the battery. In this article, we present a new software and hardware architecture for information appliances that provides significant advantages in terms of device lifetime. The architecture combines a low-power microcontroller with a high-performance application processor, where the low-power microcontroller is used to handle simple user interactions (e.g., turning pages, inking, entering text) without waking up the main application processor. We demonstrate how this architecture is easily adapted to the traditional way of building user interfaces using a user interface markup language. We report on our initial measurements using an E Ink-based prototype. When comparing our hybrid architecture to a simpler solution we found that we can increase the battery life by a factor of 1.72 for a reading task and by a factor of 3.23 for a writing task. We conclude by presenting design guidelines aimed at optimizing the overall energy signature of information appliances.
François Guimbretière, Shenwei Liu, Rajit Manohar
ACM Trans. Embed. Comput. Syst.4
2013 Neural spiking dynamics in asynchronous digital circuits
abstract
We implement a digital neuron in silicon using delay-insensitive asynchronous circuits. Our design numerically solves the Izhikevich equations with a fixed-point number representation, resulting in a compact and energy-efficient neuron with a variety of dynamical characteristics. A digital implementation results in stable, reliable and highly programmable circuits, while an asynchronous design style leads to energy-efficient clockless neurons and their networks that mimic the event-driven nature of biological nervous systems. In 65 nm CMOS technology at 1 V operating voltage and a 16-bit word length, our neuron can update its state 11,600 times per millisecond while consuming 0.5 nJ per update. The design occupies 29,500 μm2and can be used to construct dense neuromorphic systems. Our neuron exhibits the full repertoire of spiking features seen in biological neurons, resulting in a range of computational properties that can be used in artificial systems running neural-inspired algorithms, in neural prosthetic devices, and in accelerated brain simulations.
Nabil Imam, Kyle Wecker, Jonathan Tse, Robert Karmazin, Rajit Manohar
IJCNN5
2012 Building block of a programmable neuromorphic substrate: A digital neurosynaptic core
abstract
The grand challenge of neuromorphic computation is to develop a flexible brain-inspired architecture capable of a wide array of real-time applications, while striving towards the ultra-low power consumption and compact size of biological neural systems. Toward this end, we fabricated a building block of a modular neuromorphic architecture, a neurosynaptic core. Our implementation consists of 256 integrate-and-fire neurons and a 1,024×256 SRAM crossbar memory for synapses that fits in 4.2mm2using a 45nm SOI process and consumes just 45pJ per spike. The core is fully configurable in terms of neuron parameters, axon types, and synapse states and its fully digital implementation achieves one-to-one correspondence with software simulation models. One-to-one correspondence allows us to introduce an abstract neural programming model for our chip, a contract guaranteeing that any application developed in software functions identically in hardware. This contract allows us to rapidly test and map applications from control, machine vision, and classification. To demonstrate, we present four test cases (i) a robot driving in a virtual environment, (ii) the classic game of pong, (iii) visual digit recognition and (iv) an autoassociative memory.
John V. Arthur, Paul Merolla, Filipp Akopyan, Rodrigo Alvarez-Icaza, Andrew S. Cassidy, Shyamal Chandra, Steven K. Esser, Nabil Imam, William P. Risk, Daniel Ben Dayan Rubin, Rajit Manohar, Dharmendra S. Modha
IJCNN11
2011 Energy-Efficient Pipeline Templates for High-Performance Asynchronous Circuits
abstract
We present two novel energy-efficient pipeline templates for high throughput asynchronous circuits. The proposed templates, called N-P and N-Inverter pipelines, use a single-track handshake protocol. There are multiple stages of logic within each pipeline. The proposed techniques minimize handshake overheads associated with input tokens and intermediate logic nodes within a pipeline template. Each template can pack a significant amount of logic in a single stage, while still maintaining a fast cycle time of only 18 transitions. Noise and timing robustness constraints of our pipelined circuits are quantified across all process corners. We present completion detection scheme based on wide NOR gates, which results in significant latency and energy savings especially as the number of outputs increase. To fully quantify all design trade-offs, three separate pipeline implementations of an 8x8-bit Booth-encoded array multiplier are presented. Compared to a standard QDI pipeline implementation, the N-Inverter and N-P pipeline implementations reduced the energy-delay product by 38.5% and 44% respectively. The overall multiplier latency was reduced by 20.2% and 18.7%, while the total transistor width was reduced by 35.6% and 46% with N-Inverter and N-P pipeline templates respectively.
Basit Riaz Sheikh, Rajit Manohar
ACM J. Emerg. Technol. Comput. Syst.2
2007 Utilizing Dynamically Coupled Cores to Form a Resilient Chip Multiprocessor
abstract
Aggressive CMOS scaling will make future chip multiprocessors (CMPs) increasingly susceptible to transient faults, hard errors, manufacturing defects, and process variations. Existing fault-tolerant CMP proposals that implement dual modular redundancy (DMR) do so by statically binding pairs of adjacent cores via dedicated communication channels and buffers. This can result in unnecessary power and performance losses in cases where one core is defective (in which case the entire DMR pair must be disabled), or when cores exhibit different frequency/leakage characteristics due to process variations (in which case the pair runs at the speed of the slowest core). Static DMR also hinders power density/thermal management, as DMR pairs running code with similar power/thermal characteristics are necessarily placed next to each other on the die. We present dynamic core coupling (DCC), an architectural technique that allows arbitrary CMP cores to verify each other's execution while requiring no static core binding at design time or dedicated communication hardware. Our evaluation shows that the performance overhead of DCC over a CMP without fault tolerance is 3% on SPEC2000 benchmarks, and is within 5% for a set of scalable parallel scientific and data mining applications with up to eight threads (16 processors). Our results also show that DCC has the potential to significantly outperform existing static DMR schemes.
Christopher LaFrieda, Engin Ipek, José F. Martínez, Rajit Manohar
DSN4
2006 Yield enhancement of asynchronous logic circuits through 3-dimensional integration technology
abstract
This paper presents a systematic design for yield enhancement of asynchronous logic circuits using 3-D (3-Dimensional) integration technology. In this design, the target asynchronous circuits on one planar device layer which is fabricated with aggressive technology, are built on fault tolerant graph models with extra spare resources, and can be reconfigured by autonomous reconfiguration logic on another planar device layer which is fabricated with conservative technology, in the presence of hard errors. The yield analysis shows that this method can result in 20--30% overall yield enhancement. This design methodology can be conveniently applied to clocked designs without significant changes.
Rajit Manohar
ACM Great Lakes Symposium on VLSI2
2005 A High-Performance Asynchronous FPGA: Test Results
abstract
We report test results from a prototype asynchronous FPGA (AFPGA) implemented in TSMC's 0.18 /spl mu/m CMOS process. The AFPGA uses SRAM-based configuration bits with pipelined logic blocks and switch boxes. Test results demonstrate a throughput of 674 MHz at 1.8 V.
David Fang, John Teifel, Rajit Manohar
FCCM3
2005 Automated synthesis for asynchronous FPGAs
abstract
We present an automatic logic synthesis method targeted for high-performance asynchronous FPGA (AFPGA) architectures. Our method transforms sequential programs as well as high-level descriptions of asynchronous circuits into fine-grain asynchronous process netlists suitable for an AFPGA. The resulting circuits are inherently pipelined, and can be physically mapped onto our AFPGA with standard partitioning and place-and-route algorithms. For a wide variety of benchmarks, our automatic synthesis method not only yields comparable logic densities and performance to those achieved by hand placement, but also attains a throughput close to the peak performance of the FPGA.
David Fang, John Teifel, Rajit Manohar
FPGA4
2005 Fault Tolerant Asynchronous Adder through Dynamic Self-reconfiguration
abstract
This paper presents a systematic method for the design of a self-healing asynchronous adder. We propose a graph-based model for the design of a fault-tolerant linear array with external inputs and outputs with a minimum number of spare resources. A K-fault-tolerant asynchronous adder design is presented based on this analysis, together with the necessary support logic for dynamic self-reconfiguration. Experimental evaluations show that our method incurs both low hardware cost and small performance overhead compared to traditional approaches to fault-tolerance.
Rajit Manohar
ICCD2
2004 An ultra low-power processor for sensor networks
abstract
We present a novel processor architecture designed specifically for use in low-power wireless sensor-network nodes. Our sensor network asynchronous processor (SNAP/LE) is based on an asynchronous data-driven 16-bit RISC core with an extremely low-power idle state, and a wakeup response latency on the order of tens of nanoseconds. The processor instruction set is optimized for sensor-network applications, with support for event scheduling, pseudo-random number generation, bitfield operations, and radio/sensor interfaces. SNAP/LE has a hardware event queue and event coprocessors, which allow the processor to avoid the overhead of operating system software (such as task schedulers and external interrupt servicing), while still providing a straightforward programming interface to the designer. The processor can meet performance levels required for data monitoring applications while executing instructions with tens of picojoules of energy.We evaluate the energy consumption of SNAP/LE with several applications representative of the workload found in data-gathering wireless sensor networks. We compare our architecture and software against existing platforms for sensor networks, quantifying both the software and hardware benefits of our approach.
Virantha N. Ekanayake, Clinton Kelly IV, Rajit Manohar
ASPLOS3
2004 Fault Detection and Isolation Techniques for Quasi Delay-Insensitive Circuits
abstract
This paper presents a circuit fault detection and isolation technique for quasi delay-insensitive asynchronous circuits. We achieve fault isolation by a combination of physical layout and circuit techniques. The asynchronous nature of quasi delay-insensitive circuits combined with layout techniques makes the design tolerant to delay faults. Circuit techniques are used to make sections of the design robust to nondelay faults. The combination of these is an asynchronous defect-tolerant circuit where a large class of faults are tolerated, and the remaining faults can be both detected easily and isolated to a small region of the design.
Christopher LaFrieda, Rajit Manohar
DSN2
2004 Highly pipelined asynchronous FPGAs
abstract
We present the design of a high-performance, highly pipelined asynchronous FPGA. We describe a very fine-grain pipelined logic block and routing interconnect architecture, and show how asynchronous logic can efficiently take advantage of this large amount of pipelining. Our FPGA, which does not use a clock to sequence computations, automatically self-pipelines" its logic without the designer needing to be explicitly aware of all pipelining details. This property makes our FPGA ideal for throughput-intensive applications and we require minimal place and route support to achieve good performance. Benchmark circuits taken from both the asynchronous and clocked design communities yield throughputs in the neighborhood of 300--400 MHz in a TSMC 0.25m process and 500--700 MHz in a TSMC 0.18m process.
John Teifel, Rajit Manohar
FPGA2
2004 An Asynchronous Dataflow FPGA Architecture
abstract
We discuss the design of a high-performance field programmable gate array (FPGA) architecture that efficiently prototypes asynchronous (clockless) logic. In this FPGA architecture, low-level application logic is described using asynchronous dataflow functions that obey a token-based compute model. We implement these dataflow functions using finely pipelined asynchronous circuits that achieve high computation rates. This asynchronous dataflow FPGA architecture maintains most of the performance benefits of a custom asynchronous design, while also providing postfabrication logic reconfigurability. We report results for two asynchronous dataflow FPGA designs that operate at up to 400 MHz in a typical TSMC 0.25 /spl mu/m CMOS process.
John Teifel, Rajit Manohar
IEEE Trans. Computers2
2003 An Event-Synchronization Protocol for Parallel Simulation of Large-Scale Wireless Networks
abstract
We present a new conservative event-synchronization protocol, a time-based synchronization, for parallel discrete-event simulation of mobile ad hoc wireless networks. Simulators that use our protocol proceed at a scaled version of real time and send messages that correspond only to transmissions in the simulated network. We show that such simulators can maintain a constant execution time even as the sizes of the networks that they simulate grow. Moreover, we show that these simulators, when executed on a custom parallel architecture, are capable of simulating many networks faster than real time.
Clinton Kelly IV, Rajit Manohar
DS-RT2
2003 Programmable Asynchronous Pipeline Arrays
John Teifel, Rajit Manohar
FPL2
2003 Power optimal routing in wireless networks
abstract
Reducing power consumption and increasing battery life of nodes in an ad-hoc network requires an integrated power control and routing strategy. Power optimal routing selects the multi-hop links that require the minimum total power cost for data transmission under a constraint on the link quality. This paper studies optimal power routing under the constraint of a fixed end-to-end probability of error and compares the power optimal routes obtained with this criterion with those from the more commonly used fixed per hop error rate constraint. The comparison is carried out by looking at the properties of the power optimal graph, formed by the union of all the power optimal routes. The paper also provides algorithms to determine the power optimal routes.
Rajit Manohar, Anna Scaglione
ICC1
2002 Scalable formal design methods for asynchronous VLSI
abstract
This lecture will provide an overview of the field of asynchronous VLSI, and show how formal methods have played a critical role in the design of complex asynchronous systems. In particular, I will talk about program transformations and their application to asynchronous VLSI, as well as describe a simple language that I developed to describe these circuits and aid in their validation.
Rajit Manohar
POPL1
1999 Joining Specification Statements
K. Rustan M. Leino, Rajit Manohar
Theor. Comput. Sci.2
1999 The entropy of traces in parallel computation
abstract
The following problem arises in the context of parallel computation: how many bits of information are required to specify any one element from an arbitrary (non-empty) k-subset of a set? We characterize optimal coding techniques for this problem. We calculate the asymptotic behavior of the amount of information necessary, and construct an algorithm that specifies an element from a subset in an optimal manner.
Rajit Manohar
IEEE Trans. Inf. Theory1
1998 Slack Elasticity in Concurrent Computing
Rajit Manohar, Alain J. Martin
MPC1
1998 Asynchronous Parallel Prefix Computation
abstract
The prefix problem is to compute all the products x/sub 1//spl otimes/x/sub 2//spl otimes/.../spl otimes/x/sub k/, for 1/spl les/k/spl les/n, where /spl otimes/ is an associative binary operation. We start with an asynchronous circuit to solve this problem with O(log n) latency and O(n log n) circuit size, with O(n) /spl otimes/-operations in the circuit. Our contributions are: (1) a modification to the circuit that improves its average-case latency from O(log n) to O(log log n) time, and (2) a further modification that allows the circuit to run at full-throughput, i.e., with constant response time. The construction can be used to obtain an asynchronous adder with O(log n) worst-case latency and O(log log n) average-case latency.
Rajit Manohar, José A. Tierno
IEEE Trans. Computers1
1997 Performance and Portability of an Air Quality Model
Donald Dabdub, Rajit Manohar
Parallel Comput.2
1995 Conditional Composition
abstract
Abstract Generalizing the notion of function composition, we introduce the concept of conditional function composition and present a theory of such compositions. We use the theory to describe the semantics of a programming language with exceptions, and to relate exceptions to the IF statement.
Rajit Manohar, K. Rustan M. Leino
Formal Aspects Comput.1