EDBT 2026 Demo / reviewers in the wild / expert
Charles Augustine
dblp:41/5248
· DBLP profile ↗
17ranked-venue papers
1as first author
3since 2021 · last 2023
0000-0001-8892-6286ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Eidetic: An In-Memory Matrix Multiplication Accelerator for Neural NetworksabstractThis paper presents theEideticarchitecture, which is an SRAM-based ASIC neural network accelerator that eliminates the need to continuously load weights from off-chip, while also minimizing the need to go off chip for intermediate results. Using in-situ arithmetic in the SRAM arrays, this architecture can supports a variety of precision types allowing for effective inference. We also present different data mapping policies for matrix-vector based networks (RNN and MLP) on theEideticarchitecture and describe the tradeoffs involved. With this architecture, multiple layers of a network can be concurrently mapped, storing both the layer weights and intermediate results on-chip, removing the energy and latency penalty of off-chip memory accesses. We evaluateEideticon Google's Neural Machine Translation System (GNMT) encoder and demonstrate a 17.20× increase in throughput and 7.77× reduction in average latency over a single TPUv2 chip. Charles Eckert, Arun Subramaniyan 0001, Xiaowei Wang 0005, Charles Augustine, Ravi R. Iyer 0001, Reetuparna Das |
IEEE Trans. Computers | 4 |
| 2021 | Compute-Capable Block RAMs for Efficient Deep Learning Acceleration on FPGAsabstractThe density of FPGA on-chip memory has been continuously increasing with modern FPGAs having thousands of block RAMs (BRAMs) distributed across their reconfigurable fabric. These distributed BRAMs can provide a tremendous amount of on-chip bandwidth for efficient acceleration of data-intensive applications. In this work, we propose enhancing the ubiquitous FPGA BRAMs with in-memory compute-capabilities. As a result, BRAMs can act as normal storage units or their bitlines can be re-purposed as SIMD lanes executing bit-serial arithmetic operations. Our proposed architectural change results in 1.6× and 2.3× increase in the peak multiply-accumulate throughput of a large Stratix 10 FPGA, at a minimal cost of only 1.8% increase in the FPGA die size and no change to the BRAM's interface to the programmable routing. Then, we present RIMA, a reconfigurable in-memory accelerator architecture for deep learning (DL) inference. RIMA exploits the proposed compute-capable BRAMs and the FPGA's reconfigurability to achieve 1.25× and 3× higher performance compared to the state-of-the-art Brainwave DL soft processor for 8-bit integer and block floating-point precisions, respectively. In addition, RIMA implemented on a Stratix 10 FPGA enhanced with compute-capable BRAMs can achieve an order of magnitude higher performance compared to a same-generation GPU. Xiaowei Wang 0005, Vidushi Goyal, Jiecao Yu, Valeria Bertacco, Andrew Boutros, Eriko Nurvitadhi, Charles Augustine, Ravi R. Iyer 0001, Reetuparna Das |
FCCM | 7 |
| 2021 | Cache Compression with Efficient in-SRAM Data ComparisonabstractWe present a novel cache compression method that leverages the fine-grained data duplication across cache lines. We leverage the XOR operation of the in-SRAM bit-line computing peripherals, to search for compressible data over a wide range of data locations on cache, reducing the data movement requirements. To reduce the decompression latency, we design specialized compression schemes by fetching the data with the same parallelism as the original cache, according to the architecture of the last-level cache slice. The proposed compression method achieves a 2.05× compression ratio on average (up to 67×), and 4.73% of speedup on average (up to 29%), over the SPEC2006 benchmarks. Xiaowei Wang 0005, Charles Augustine, Eriko Nurvitadhi, Ravi R. Iyer 0001, Li Zhao 0002, Reetuparna Das |
NAS | 2 |
| 2019 | Bit Prudent In-Cache Acceleration of Deep Convolutional Neural NetworksabstractWe propose Bit Prudent In-Cache Acceleration of Deep Convolutional Neural Networks - an in-SRAM architecture for accelerating Convolutional Neural Network (CNN) inference by leveraging network redundancy and massive parallelism. The network redundancy is exploited in two ways. First, we prune and fine-tune the trained network model and develop two distinct methods - coalescing and overlapping - to run inferences efficiently with sparse models. Second, we propose an architecture for network models with a reduced bit width by leveraging bit-serial computation. Our proposed architecture achieves a 17.7×/3.7× speedup over server class CPU/GPU, and a 1.6× speedup compared to the relevant in-cache accelerator, with 2% area overhead each processor die, and no loss on top-1 accuracy for AlexNet. With a relaxed accuracy limit, our tunable architecture achieves higher speedups. Xiaowei Wang 0005, Jiecao Yu, Charles Augustine, Ravi R. Iyer 0001, Reetuparna Das |
HPCA | 3 |
| 2017 | Event-driven random backpropagation: Enabling neuromorphic deep learning machinesabstractAn ongoing challenge in neuromorphic computing is to devise general and computationally efficient models of inference and learning which are compatible with the spatial and temporal constraints of the brain. The gradient descent back-propagation rule is a powerful algorithm that is ubiquitous in deep learning, but it relies on the immediate availability of network-wide information stored with high-precision memory. However, recent work shows that exact backpropagated weights are not essential for learning deep representations. Here, we demonstrate an event-driven random backpropagation (eRBP) rule that uses an error-modulated synaptic plasticity rule for learning deep representations in neuromorphic computing hardware. The rule is very suitable for implementation in neuromorphic hardware using a two-compartment leaky integrate & fire neuron and a membrane-voltage modulated, spike-driven plasticity rule. Our results show that using eRBP, deep representations are rapidly learned without using backpropagated gradients, achieving nearly identical classification accuracies compared to artificial neural network simulations on GPUs, while being robust to neural and synaptic state quantizations during learning. Emre Neftci, Charles Augustine, Somnath Paul, Georgios Detorakis |
ISCAS | 2 |
| 2016 | Synaptic sampling in hardware spiking neural networksabstractUsing a neural sampling approach, networks of stochastic spiking neurons, interconnected with plastic synapses, have been used to construct computational machines such as Restricted Boltzmann Machines (RBMs). Previous work towards building such networks achieved lower performances than traditional RBMs. More recently, Synaptic Sampling Machines (SSMs) were shown to outperform equivalent RBMs. In Synaptic Sampling Machines (SSMs), the stochasticity for the sampling is generated at the synapse. Stochastic synapses play the dual role of a regularizer during learning and an efficient mechanism for implementing stochasticity in neural networks over a wide dynamic range. In this paper we show that SSMs with stochastic synapses implemented in FPGA-based spiking neural networks can obtain a high accuracy in classifying MNIST handwritten digit database. We compare classification accuracy for different bit precision for stochastic and non-stochastic synapses and further argue that stochastic synapses have the same effect as synapses with higher bit precision but require significantly lower computational resources. Sadique Sheik, Somnath Paul, Charles Augustine, Chinnikrishna Kothapalli, Muhammad M. Khellah, Gert Cauwenberghs, Emre Neftci |
ISCAS | 3 |
| 2016 | Cache Design with Domain Wall MemoryabstractDomain wall memory (DWM) is a recently developed spin-based memory technology in which several bits of data are densely packed into the domains of a ferromagnetic wire. DWM has shown great promise in enabling non-volatile memory with very high density and energy efficiency, and has been explored for secondary storage and off-chip memory. In this work, we explore the use of DWM within the on-chip cache hierarchy of general purpose computing platforms. Our work is motivated by the fact that DWMs enable much higher density compared to SRAM, DRAM, and other spin-based memory technologies such as STT-MRAM. However, DWMs also pose the unique challenge of serial access to the bits stored in a cell, leading to large and variable access latencies. In addition, DWMs share the inherent write inefficiency of other spin-based memories. We propose TapeCache, a DWM-based cache design that employs device, circuit, and architectural techniques to address these challenges. At the device level, we perform write optimization by employing a new write mechanism based on domain wall shifts to achieve fast, energy-efficient writes in DWM. At the circuit level, we propose different DWM bit-cell designs that are tailored to the distinct architectural requirements of different levels in the cache hierarchy. At the architecture level, we propose a new cache organization and suitable management policies that mitigate the performance penalty arising from serial access to bits in a DWM cell. We show that the holistic device-circuit-architecture co-design enables all the levels in the cache hierarchy to be realized using DWM and benefit from its improved density. Over a wide range of SPEC CPU 2006 benchmarks, TapeCache achieves an average energy improvement of 7.5x, with virtually identical performance and 7.8x improvement in area, compared to an iso-capacity SRAM cache. Compared to an iso-capacity STT-MRAM cache, TapeCache obtains 3.1x improvement in area and 2x average energy savings along with 1.1 percent performance improvement. Rangharajan Venkatesan, Vivek Joy Kozhikkottu, Mrigank Sharad, Charles Augustine, Arijit Raychowdhury, Kaushik Roy 0001, Anand Raghunathan |
IEEE Trans. Computers | 4 |
| 2013 | Minimum supply voltage for sequential logic circuits in a 22nm technologyabstractThe minimum supply voltage (Vmin) is explored for sequential logic circuits by statistically simulating the impact of within-die process variations and gate-dielectric soft breakdown on data retention and hold time. As supply voltage (Vcc) scales, statistical circuit simulations demonstrate that hold time increases faster than circuit delay or cycle time, consequently the required number of min-delay buffers increases. For this reason, a new hold-time violation metric defines Vminas the Vccin which the hold time exceeds a target percentage of the cycle time. Simulation results in a 22nm tri-gate CMOS technology indicate a data-retention Vminof 0.61Vnorm and a hold-time Vminof 0.73Vnorm, where Vnormrepresents a normalized voltage for the process technology node. A key insight reveals that upsizing the first clock inverter in the sequential circuit reduces the hold-time Vminby 18% and the overall Vminby 16%. Keith A. Bowman, Charles Augustine, Zhengya Zhang, James W. Tschanz |
ISLPED | 3 |
| 2013 | Dual pillar spin-transfer torque MRAMs for low power applicationsabstractElectron-spin based data storage for on-chip memories has the potential for ultra-high density, low power consumption, very high endurance, and reasonably low read/write latency. In this article, we discuss the design challenges associated with spin-transfer torque (STT) MRAM in its state-of-the-art configuration. We propose an alternative bit cell configuration and three new genres of magnetic tunnel junction (MTJ) structures to improve STT-MRAM bit cell stabilities, write endurance, and reduce write energy consumption. The proposed multi-port, multi-pillar MTJ structures offer the unique possibility of electrical and spatial isolation of memory read and write. In order to realize ultralow power under process variations, we propose device, bit-cell and architecture level design techniques. Such design alternatives at multiple levels of design abstraction has been found to achieve substantially enhanced robustness, density, reliability and low power as compared to their charge-based counterparts for future embedded applications. Niladri Narayan Mojumder, Xuanyao Fong, Charles Augustine, Sumeet Kumar Gupta, Sri Harsha Choday, Kaushik Roy 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2012 | Cognitive computing with spin-based neural networksabstractWe model a step transfer function neuron with lateral spin valve (LSV) and propose its application in low power neural network hardware. The computational task in such a network is performed by nano-magnets, metal channels and programmable conductive elements, that constitute the neuron-synapse units and operate at a terminal voltage of ~20 mV. CMOS transistors provide peripheral support in the form of clocking, power gating and inter-neuron signaling. Simulations for cognitive as well as Boolean computation applications show more than 94% improvement in power consumption as compared to a conventional CMOS design at the same technology node. Mrigank Sharad, Charles Augustine, Georgios Panagopoulos, Kaushik Roy 0001 |
DAC | 2 |
| 2012 | A framework for simulating hybrid MTJ/CMOS circuits: Atoms to system approachabstractA simulation framework that can comprehend the impact of material changes at the device level to the system level design can be of great value, especially to evaluate the impact of emerging devices on various applications. To that effect, we have developed a SPICE-based hybrid MTJ/CMOS (magnetic tunnel junction) simulator, which can be used to explore new opportunities in large scale system design. In the proposed simulation framework, MTJ modeling is based on Landau-Lifshitz-Gilbert (LLG) equation, incorporating both spin-torque and external magnetic field(s). LLG along with heat diffusion equation, thermal variations, and electron transport are implemented using SPICE-in built voltage dependent current sources and capacitors. The proposed simulation framework is flexible since the device dimensions such as MgO thickness and area, are user defined parameters. Furthermore, we have benchmarked this model with experiments in terms of switching current density (JC), switching time (TSWITCH) and tunneling magneto-resistance (TMR). Finally, we used our framework to simulate STT-MRAMs and magnetic flip-flops (MFF). Georgios Panagopoulos, Charles Augustine, Kaushik Roy 0001 |
DATE | 2 |
| 2012 | Spin based neuron-synapse module for ultra low power programmable computational networksabstractA spin-CMOS hybrid design for neural network is presented. We employ spin torque switched nano-magnets for realizing ultra low power, high speed neuron and domain wall magnets for compact, programmable synapses. The spin based neuron-synapse units operate locally at ultra low supply voltage of 30 mV resulting in low computation power. CMOS based longer distance, inter neuron communication achieves high integration. We corroborate circuit operation with physics based models developed for the spin devices. Simulation results for a benchmark application shows 95% improvement in power consumption as compared to 45 nm CMOS design. Mrigank Sharad, Charles Augustine, Georgios Panagopoulos, Kaushik Roy 0001 |
IJCNN | 2 |
| 2012 | TapeCache: a high density, energy efficient cache based on domain wall memoryabstractDomain Wall Memory (DWM) is a recently developed spin-based memory technology in which several bits of data are densely packed into the domains of a ferromagnetic wire. DWM has shown great promise in enabling non-volatile memory with unprecedented density and high energy efficiency. In this work, we propose TapeCache, a first attempt to employ DWMs as last-level caches in general purpose computing platforms. DWMs enable much higher density compared to SRAM, DRAM, and other spin-based memory technologies such as STT-MRAM. However, they also pose unique challenges such as serial access to the bits stored in a DWM cell, leading to variable access latencies. We propose a novel circuit-architecture co-design for TapeCache, consisting of (i) a multi-port DWM macro-cell optimized for read operations considering the asymmetry in applications' read/write characteristics, and (ii) a new cache organization and suitable management policies that mitigate the performance penalty arising from serial access to bits in a macro-cell. Over a wide range of SPEC 2006 benchmarks, TapeCache achieves 7.8X improvement in area, an average energy improvement of 7.3X, and an average performance improvement of 1.2% compared to an iso-capacity SRAM cache. Compared to an iso-capacity STT-MRAM cache, TapeCache obtains 2.3X improvement in area and 1.4X average energy savings with virtually identical performance. Rangharajan Venkatesan, Vivek Joy Kozhikkottu, Charles Augustine, Arijit Raychowdhury, Kaushik Roy 0001, Anand Raghunathan |
ISLPED | 3 |
| 2010 | A self-consistent model to estimate NBTI degradation and a comprehensive on-line system lifetime enhancement techniqueabstractLifetime reliability and the resultant temporal performance degradation due to Negative Bias Temperature Instability (NBTI) has emerged as a critical challenge in design and test of integrated circuits in nanometer technology nodes. In this work, we have developed a model that self-consistently estimates the NBTI degradation by considering the impact on the circuit lifetime of inter-dependent parameters such as Vddand temperature simultaneously. Using the proposed model, we observed that a circuit with lower Vddcan provide better lifetime performance than with higher Vdd. This interesting observation can be attributed to the reduction of electric field in the transistor along with the circuit power/temperature reduction that leads to lesser NBTI degradation. Based on this observation we have developed a on-line detection and mitigation scheme that allows Vddscaling to enhance system lifetime. The proposed scheme was applied to various arithmetic units and results in 45nm IBM process technology show 18% lifetime improvement with 57% reduction in power compared to conventional mitigation techniques. We also show that by using existing NBTI estimation models, the error in delay estimation can be as large as 7.6%. Georgios Karakonstantis, Charles Augustine, Kaushik Roy 0001 |
IOLTS | 2 |
| 2009 | A design methodology and device/circuit/architecture compatible simulation framework for low-power magnetic quantum cellular automata systemsabstractCMOS device scaling is facing a daunting challenge with increased parameter variations and exponentially higher leakage current every new technology generation. Thus, researchers have started looking at alternative technologies. Magnetic Quantum Cellular Automata (MQCA) is such an alternative with switching energy close to thermal limits and scalability down to 5nm. In this paper, we present a circuit/architecture design methodology using MQCA. Novel clocking techniques and strategies are developed to improve computation robustness of MQCA systems. We also developed an integrated device/circuit/system compatible simulation framework to evaluate the functionality and the architecture of an MQCA based system and conducted a feasibility/comparison study to determine the effectiveness of MQCAs in digital electronics. Simulation results of an 8-bit MQCA-based Discrete Cosine Transform (DCT) with novel clocking and architecture show up to 290X and 46X improvement (at iso-delay and optimistic assumption) over 45nm CMOS in energy consumption and area, respectively. Charles Augustine, Behtash Behin-Aein, Xuanyao Fong, Kaushik Roy 0001 |
ASP-DAC | 1 |
| 2008 | Modeling of failure probability and statistical design of spin-torque transfer magnetic random access memory (STT MRAM) array for yield enhancementabstractSpin-Torque Transfer Magnetic RAM (STT MRAM) is a promising candidate for future universal memory. It combines the desirable attributes of current memory technologies such as SRAM, DRAM and flash memories. It also solves the key drawbacks of conventional MRAM technology: poor scalability and high write current. In this paper, we analyzed and modeled the failure probabilities of STT MRAM cells due to parameter variations. Based on the model, we developed an efficient simulation tool to capture the coupled electro/magnetic dynamics of spintronic device, leading to effective prediction for memory yield. We also developed a statistical optimization methodology to minimize the memory failure probability. The proposed methodology can be used at an early stage of the design cycle to enhance memory yield. Jing Jane Li, Charles Augustine, Sayeef S. Salahuddin, Kaushik Roy 0001 |
DAC | 2 |
| 1992 | Testing generality in JANUS: a multi-lingual speech translation systemabstractFor speech translation to be practical and useful, speech translation systems should be portable to multiple languages without substantial modification. The authors present results of expanding the English-based JANUS speech translation system to translate from spoken German sentences to English and Japanese utterances. The authors also report the results of implementing part of the linked predictive neural network (LPNN) speech recognition module on a massively parallel machine. The JANUS approach generalizes well, with overall system performance of 97%. This surpasses English-based JANUS performance.> Louise Osterholtz, Charles Augustine, Arthur E. McNair, Ivica Rogina, Hiroaki Saito 0001, Tilo Sloboda, Joe Tebelskis, Alex Waibel |
ICASSP | 2 |