EDBT 2026 Demo / reviewers in the wild / expert
Joycee Mekie
dblp:84/499 · also Joycee M. Mekie
· DBLP profile ↗
30ranked-venue papers
0as first author
21since 2021 · last 2025
0000-0001-9646-1941ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 19 since 2021Software engineering, systems software and programming languages · 8 · 7 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VeriBench: Benchmarking Large Language Models for Verilog Code Generation and Design SynthesisabstractIn the rapidly advancing field of hardware design, Electronic Design Automation (EDA) tools can be significantly improved using Machine Learning. This study evaluates the efficacy of various Large Language Models (LLMs) for automating Electronic Design Automation for Verilog design, testbench generation, and Formal Verification (FV) assertion synthesis by comparing 3 closed-source LLMs and 14 Open-Source LLM variants. In our setup of 33 Verilog designs, ChatGPT-4 generates 22 synthesizable Verilog designs in one-shot without feedback, while the Llama 3 (8B) model generates 20. Both models generate all testbenches correctly, 9 of which are given in our setup. For generating Formal Verification properties, ChatGPT-4 generates all properties correctly, whereas Llama 3 synthesizes 7 out of 9 properties correctly. Of the sample synthesized in Vivado, ChatGPT-4 codes result into power-efficient designs as compared to Llama-3, whereas in Genus there is no clear winner. These results underscore the efficacy of open-source models, which perform competitively despite having significantly fewer parameters (8 billion) compared to closed-source models such as ChatGPT-4. This study demonstrates the potential of parameter-efficient, open-source models for hardware design and verification tasks. Mihir Agarwal, Zaqi Momin, Kailash Prasad, Joycee Mekie |
ISCAS | 4 |
| 2025 | DTQ-16T: Double Node Upset Tolerant Quadruple SRAM for Space ApplicationsabstractThe high-energy particles in space cause SRAM failures. The vulnerability of SRAM increases at lower technology, and it flips the SRAM cell’s data due to single-event multi-node-upset. Various state-of-the-art radiation hardened by design SRAMs have been proposed; however, most designs tackle Single Node Upset (SNU). This paper presents DNU Tolerant Quadruple-16T (DTQ-16T) SRAM with no read disturb. The most important feature of the proposed design is its immunity towards radiation, where it recovers from all possible upsets, whether SNU, DNU, Triple Node Upset (TNU), or Quadruple Node Upset (QNU) for storage ‘1’. On top of it, the proposed design gives very high read stability, Write Access Time, and Wordline Write Trip Voltage (WWTV) than most of the existing radiation-hardened SRAMs. Finally, the post-layout and Monte Carlo simulations validate the efficiency of the proposed SRAM in commercial CMOS 28nm technology. Pramod Kumar Bharti, Govind Prasad, Mukku Pavan Kumar, Joycee Mekie |
ISCAS | 4 |
| 2024 | CANSim: When to Utilize Synchronous and Asynchronous Routers in Large and Complex NoCsabstractAsynchronous routers offer benefits in latency, energy, and area over synchronous routers, but their large system design adoption is limited due to inadequate tool support for performance quantification. We introduce CANSim, a fast and accurate simulator for complex asynchronous and synchronous NoCs. Verified against synthesis models, CANSim models any graph-representable topology, asymmetric links, data-dependent delays, meta-stability, and supports synthetic and real-world traffic. It also accommodates power gating, hierarchical networks, and GALS-like networks. Our comparisons reveal that asynchronous NoCs provide lower packet latency at no-load conditions, but synchronous NoCs may outperform near saturation. On average, asynchronous NoCs offer up to 36% latency and 52% power benefits, with power-gating potentially reducing static power by 90%. Tom Glint, Manu Awasthi, Joycee Mekie |
ASPDAC | 3 |
| 2024 | Hardware-Software Co-Design of a Collaborative DNN Accelerator for 3D Stacked Memories with Multi-Channel DataabstractHardware accelerators are preferred over general-purpose processors for processing Deep Neural Networks (DNN) as the later suffer from power and memory walls. However, hardware accelerators designed as a separate logic chip from the memory still suffer from memory wall. Processing-in-memory accelerators, which try to overcome this memory wall by developing the compute elements as part of the memory structures, are highly constrained due to the memory manufacturing process. Near-data-processing (NDP) based hardware accelerator design is an alternative paradigm that could combine the benefit of high bandwidth, low access energy of processing-in-memory, and design flexibility of separate logic chip. However, NDP has area, data flow and thermal constraints, hindering high throughput designs. In this work, we propose an HBM3-based NDP accelerator that tackles the constraints of NDP with a hardware-software co-design approach. The proposed design takes only 50% area, delivers a speed-up of $3\times$, and is about $6 \times$ more energy efficient than state-of-the-art NDP hardware accelerator for inferencing workloads such as AlexNet, MobileNet, ResNet, and VGG without loss of accuracy. Tom Glint, Manu Awasthi, Joycee Mekie |
ASPDAC | 3 |
| 2024 | DeepFrack: A Comprehensive Framework for Layer Fusion, Face Tiling, and Efficient Mapping in DNN Hardware AcceleratorsabstractDeepFrack is a novel framework developed for enhancing energy efficiency and reducing latency in deep learning workloads executed on hardware accelerators. By optimally fusing layers and implementing an asymmetric tiling strategy, DeepFrack addresses the limitations of traditional layer-by-layer scheduling. The computational efficiency of our method is underscored by significant performance improvements seen across various deep neural network architectures such as AlexNet, VGG, and ResNets when run on Eyeriss and Simba accelerators. The reduction in latency (30 % to 40 %) and energy consumption (30 % to 50 %) are further enhanced by the efficient usage of the on-chip buffer and reduction of external memory bandwidth bottleneck. This work contributes to the ongoing efforts in designing more efficient hardware accelerators for machine learning workloads. Tom Glint, Mithil Pechimuthu, Joycee Mekie |
DATE | 3 |
| 2023 | Hardware-Software Codesign of DNN Accelerators Using Approximate Posit MultipliersabstractEmerging data intensive AI/ML workloads encounter memory and power wall when run on general-purpose compute cores. This has led to the development of a myriad of techniques to deal with such workloads, among which DNN accelerator architectures have found a prominent place. In this work, we propose a hardware-software co-design approach to achieve system-level benefits. We propose a quantized data-aware POSIT number representation that leads to a highly optimized DNN accelerator. We demonstrate this work on SOTA SIMBA architecture, extendable to any other accelerator. Our proposal reduces the buffer/storage requirements within the architecture and reduces the data transfer cost between the main memory and the DNN accelerator. We have investigated the impact of using integer, IEEE floating point, and posit multipliers for LeNet, ResNet and VGG NNs trained and tested on MNIST, CIFAR10 and ImageNet datasets, respectively. Our system-level analysis shows that the proposed approximate-fixed-posit multiplier when implemented on SIMBA architecture, achieves on average ~2.2× speed up, consumes ~3.1× less energy and requires ~3.2× less area, respectively, against the baseline SOTA architecture, without loss of accuracy (~±1%) Tom Glint, Kailash Prasad, Jinay Dagli, Krishil Gandhi, Vrajesh Patel, Joycee Mekie |
ASP-DAC | 8 |
| 2023 | PVC-RAM:Process Variation Aware Charge Domain In-Memory Computing 6T-SRAM for DNNsabstractThis work introduces PVC-RAM, a process variation aware in-memory computing (IMC) static random-access memory (SRAM) macro designed for efficient convolutional neural network (CNN) inference. PVC-RAM is a charge-domain based compact IMC and is the first fully analog IMC for 4b-weights/4b-inputs MAC operation for deep neural networks in 6T-SRAM to the best of our knowledge. Further, PVC-RAM fully computes 4-bit MAC in the analog domain and requires fewer invocations of the ADCs. Implemented in 28nm technology, PVC-RAM achieves a bitwise throughput of 6964.48 TOPS, which is 1.5× higher than the SOTA, and a bitwise energy efficiency of 75.12 TOPS/W, which is 1.3× higher than the SOTA. Sai Shubham, Shubham Pandit, Kailash Prasad, Joycee Mekie |
DAC | 4 |
| 2023 | REDRAW: Fast and Efficient Hardware Accelerator with Reduced Reads And Writes for 3D UNetabstractHardware Accelerators (HAs) proposed so far have been designed with a focus on 2D Convolutional Neural Networks (CNNs) and 3D CNNs using temporal data. To the best of our knowledge, there is no existing HA for 3D CNNs using spatial data. 3D UNet is a 3D CNN with significant applications in the medical domain. However, the total on-chip buffer size (>20 MB) required for the complete stationery approach of processing 3D UNet is cost prohibitive. In this work, we analyze the 3D UNet workload and propose a HA with an optimized memory hierarchy with a total on-chip buffer of less than 4 MB while conceding near theoretical minimum memory accesses required for processing 3D UNet. We demonstrate the efficiency of the proposed HA by comparing it with SOTA Simba architecture with the same number of MAC Units and show a 1.3× increase in TOPSwatt for an ISO-area design. Further, we revise the proposed architecture to increase the ratio of compute operations to memory operations and to meet the latency requirement of 3D UNet-based embedded applications. The revised architecture, compared against a dual instance of Simba, has similar latency. Against the dual instance of Simba, the proposed architecture achieves a 1.8 × increase in TOPS/watt in a similar area. Tom Glint, Manu Awasthi, Joycee Mekie |
DATE | 3 |
| 2023 | Analysis of Quantization Across DNN Accelerator Architecture ParadigmsabstractQuantization techniques promise to significantly reduce the latency, energy, and area associated with multiplier hardware. This work, to the best of our knowledge, for the first time, shows the system-level impact of quantization on SOTA DNN accelerators from different digital accelerator paradigms. Based on the placement of data and compute site, we identify SOTA designs from Conventional Hardware Accelerators (CHA), Near Data Processors (NDP), and Processing-in-Memory (PIM) paradigms and show the impact of quantization when inferencing CNN and Fully Connected Layer (FCL) workloads. We show that the 32-bit implementation of SOTA from PIM consumes less energy than the 8-bit implementation of SOTA from CHA for FCL, while the trend reverses for CNN workloads. Further, PIM has stable latency while scaling the word size while CHA and NDP suffer 20% to$2\times$slow down for doubling word size. Tom Glint, Chandan Kumar Jha 0001, Manu Awasthi, Joycee Mekie |
DATE | 4 |
| 2023 | PIC-RAM: Process-Invariant Capacitive Multiplier Based Analog In Memory Computing in 6T SRAMabstractIn-Memory Computing (IMC) is a promising approach to enabling energy-efficient Deep Neural Network-based applications on edge devices. However, analog domain dot product and multiplication suffers accuracy loss due to process variations. Furthermore, wordline degradation limits its minimum pulsewidth, creating additional non-linearity and limiting IMC's dynamic range and precision. This work presents a complete end-to-end process invariant capacitive multiplier based IMC in 6T-SRAM (PIC-RAM). The proposed architecture employs the novel idea of two-step multiplication in column-major IMC to support 4-bit multiplication. The PIC-RAM uses an operational amplifier-based capacitive multiplier to reduce bitline discharge allowing good enough WL pulse width. Further, it employs process tracking voltage reference and fuse capacitor to tackle dynamic and post-fabrication process variations, respectively. Our design is compute-disturb free and provides a high dynamic range. To the best of our knowledge, PIC-RAM is the first analog SRAM IMC approach to tackle process variation with a focus on its practical implementation. PIC-RAM has a high energy efficiency of about 25.6 TOPS/W for$4-\text{bit}\times 4-\text{bit}$multiplication and has only 0.5% area overheads due to the use of the capacitance multiplier. We obtain 409 bit-wise TOPS/W, which is about 2× better than state-of-the-art. PIC-RAM shows the TOP-1 accuracy for ResNet-18 on CIFAR10 and MNIST is 89.54% and 98.80% for$4bit\times 4bit$multiplication. Kailash Prasad, Aditya Biswas, Arpita Kabra, Joycee Mekie |
DATE | 4 |
| 2023 | Process Variation Resilient Current-Domain Analog In Memory ComputingabstractIn-Memory Computing (IMC) has emerged as one of the energy-efficient solutions for data and compute-intensive machine learning applications. Analog IMC architectures have high throughput, but limited bit precision. Process variation further degrades the bit-precision. This work proposes an efficient way to track process variation and compensate for it to achieve high bit-resolution, which, to the best of our knowledge, is first such proposal. PV tracking is achieved by using an additional SRAM column and compensation by a non-conventional word-line driver. The proposed circuit can be augmented to any analog IMC architecture to make it resilient to process variations. To demonstrate the versatility of the proposal, we have implemented and analyzed 2-bit dot product operation in IMC architectures with six different SRAM cell configurations, and 2-bit, 4-bit, and 8-bit dot product on 6T SRAM IMC. For these, we report a reduction of$4\times$to$14\times$in the standard deviation of statistical variations in bit-line voltage for different SRAM cells, increase in the bit-resolution from 2 bits to 4 bits or 6 bits. Kailash Prasad, Sai Shubham, Aditya Biswas, Joycee Mekie |
DATE | 4 |
| 2023 | Impact of Optimal Design Point on Performance Metrics of DNN accelerators in FPGAabstractDue to their flexibility, FPGAs are used to deploy Deep Neural Network Accelerators (DA) at various compute locations such as servers and edge computes. In this work, we show the impact of choosing architectural parameters on performance metrics for SOTA DA implemented on FPGA, and the optimal design point for maximum compute throughput. Tom Glint, Daniel Giftson, Gaurav Shah, Vrajesh Patel, Ruchit Chudasama, Sukanya More, Joycee Mekie |
ISPASS | 8 |
| 2023 | Analysis of Conventional, Near-Memory, and In-Memory DNN AcceleratorsabstractVarious DNN accelerators based on Conventional compute Hardware Accelerator (CHA), Near-Data-Processing (NDP) and Processing-in-Memory (PIM) paradigms have been proposed to meet the challenges of inferencing Deep Neural Networks (DNNs). To the best of our knowledge, this work aims to perform the first quantitative as well as qualitative comparison among the state-of-the-art accelerators from each digital DNN accelerator paradigm. Our study provides insights into selecting the best architecture for a given DNN workload. We have used workloads of the MLPerf Inference benchmark. We observe that for Fully Connected Layer (FCL) DNNs, PIM-based accelerator is 21× and 3× faster than CHA and NDP-based accelerator respectively. However, NDP is 9× and 2.5× more energy efficient than CHA and PIM for FCL. For Convolutional Neural Network (CNN) workloads, CHA is 10% and 5× faster than NDP and PIM-based accelerator respectively. Further, CHA is 1.5× and 6× more energy efficient than NDP and PIM-based accelerators respectively. Tom Glint, Chandan Kumar Jha 0001, Manu Awasthi, Joycee Mekie |
ISPASS | 4 |
| 2023 | HyGain: High-performance, Energy-efficient Hybrid Gain Cell-based Cache HierarchyabstractIn this article, we propose a “full-stack” solution to designing high-apacity and low-latency on-chip cache hierarchies by starting at the circuit level of the hardware design stack. We propose a novel half V DD precharge 2T Gain Cell (GC) design for the cache hierarchy. The GC has several desirable characteristics, including ~50% higher storage density and ~50% lower dynamic energy as compared to the traditional 6T SRAM, even after accounting for peripheral circuit overheads. We also demonstrate data retention time of 350 us (~17.5× of eDRAM) at 28 nm technology with V DD = 0.9V and temperature = 27°C that, combined with optimizations like staggered refresh, makes it an ideal candidate to architect all levels of on-chip caches. We show that compared to 6T SRAM, for a given area budget, GC-based caches, on average, provide 30% and 36% increase in IPC for single- and multi-programmed workloads, respectively, on contemporary workloads, including SPEC CPU 2017. We also observe dynamic energy savings of 42% and 34% for single- and multi-programmed workloads, respectively. Finally, in a quest to utilize the best of all worlds, we combine GC with STT-RAM to create hybrid hierarchies. We show that a hybrid hierarchy with GC caches at L1 and L2 and an LLC split between GC and STT-RAM is able to provide a 46% benefit in energy-delay product (EDP) as compared to an all-SRAM design, and 13% as compared to an all-GC cache hierarchy, averaged across multi-programmed workloads. Sarabjeet Singh, Neelam Surana, Kailash Prasad, Pranjali Jain, Joycee Mekie, Manu Awasthi |
ACM Trans. Archit. Code Optim. | 5 |
| 2023 | Single Exact Single Approximate Adders and Single Exact Dual Approximate AddersabstractIn this article, we present the design of approximate adders which provide dynamic runtime configurability between exact and approximate modes at the circuit level. We propose the single exact single approximate (SESA) adders that allow for fine grain configurability between exact and approximate modes. We also propose the single exact dual approximate (SEDA) adder that allows for coarse grain configurability between exact and approximate modes. Unlike SESA adders, the SEDA adder allows for two approximate computations at a time. Both the SESA and SEDA adders have a maximum bounded error. We implemented SESA and SEDA adders using UMC 28-nm technology node and evaluated them using the Cadence virtuoso tool. On average, SESA and SEDA adders consume 40% and 51% lesser energy when compared with the exact mirror adder when used in approximate mode. We have evaluated our result on image addition and image enhancement using 16-bit SESA and SEDA adders. We also evaluated 32-bit SESA and SEDA adders on the Moby benchmarks to highlight their use in approximate processors. Chandan Kumar Jha 0001, Ankita Nandi, Joycee Mekie |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | RHSCC-16T: Radiation Hardened Sextuple Cross Coupled Robust SRAM Design for Radiation Prone EnvironmentsabstractThe vulnerability of radiation-induced Single Node Upset (SNU) and Double Node Upset (DNU) on SRAM increase at lower technology nodes. Various state-of-the-art solutions, such as Quatro-10T, DICE cell, etc., have been proposed. Quatro-10T recovers from SNU using the negative feedback networks connected to storage nodes. However, Quatro-10T is not entirely immune to SNU. DICE cell is resistant to SNU, but the design suffers from DNU. In this paper, we propose a Radiation Hardened Sextuple Cross Coupled-16T (RHSCC-16T) SRAM, which is immune to SNU for all the cases and DNU for many cases with less area requirement as compared to DICE cell. In addition, it gives better read access time, write access time, read static noise margin, and wordline write trip voltage than many compared SRAMs. The proposed design possesses 1.54 ×, 1.54 ×, 1.54×, and 1.36× shorter read access time than STD-6T, Quatro-10T, RHM-12T, and RSP-14T. It also exhibits 11.25 ×, and 1.03 × higher read static noise margin than RHM-12T ×and STD-6T and 2.01× higher wordline write trip voltage than Quatro-10T @ VDD = 0.9 V at CMOS 28nm Technology. Pramod Kumar Bharti, Joycee Mekie |
ICCD | 2 |
| 2022 | Compute-In-Memory Using 6T SRAM for a Wide Variety of WorkloadsabstractThis paper presents a split wordline 6T SRAM based compute-in-memory subarray with variable multi-bit precision for input operands and outputs. The split wordlines of the 6T cell enable sign segregation, thus allowing arbitrary sign/magnitude multiply and accumulate (MAC) operations. The arbitrary signmagnitude MAC operation extends the usage of the MAC array for DSP workloads as well. On top of that, the split wordline 6T cell is more resilient to write-disturb, which is a major concern for conventional 6T cell based compute-in-memory operation. The proposed system was designed and implemented using 65nm UMC technology and the energy efficiency and throughput are found to be 80.1 TOPS/W and 496 GOPS respectively, for maximum precision of inputs(4 bits)/outputs(5 bits). The maximum throughput and energy efficiency of the design are found to be 780 GOPS and 94.8 TOPS/W respectively. Hand-written digit recognition application mapped on the proposed system showed a maximum accuracy degradation of 0.2% as compared to that obtained from software with the same input/output quantization. Pramod Kumar Bharti, Kamlesh R. Pillai, Sagar Varma Sayyaparaju, Gurpreet S. Kalsi, Joycee Mekie, Sreenivas Subramoney |
ISCAS | 6 |
| 2022 | An Automated Approach to Compare Bit Serial and Bit Parallel In-Memory Computing for DNNsabstractThis paper presents an exhaustive comparison of two different techniques for In-Memory Computing in SRAM: bit-serial arithmetic (BSA), and bit-parallel arithmetic (BPA). We have modeled both BSA and BPA and integrated them with CACTI for the CMOS 28nm technology node. The results are analyzed for both approaches with ten different sub-array configurations ranging from 128x128 to 2048x2048. We performed the convolution operation on the ImageNet dataset for comparison. The key observation is that the BPA begins to yield at least 25% better (lower) Energy Delay Product (EDP) as compared to that in BSA for large (2048x2048) sub-array sizes. The BSA yields $4 \times$ lower delay on an average, while the BPA yields $\sim 6 \times$ lower dynamic energy. Hence, the choice of the IMC architecture needs to be made depending on the application need (low energy/high performance). We have also simulated 12 different multi-bank IMC arrangements and show that just modifying the memory array structure can improve the EDP by up to $8 \times$. Alok Parmar, Kailash Prasad, Nanditha P. Rao, Joycee Mekie |
ISCAS | 4 |
| 2022 | Fast and Low-Power Quantized Fixed Posit High-Accuracy DNN ImplementationabstractThis brief compares quantized float-point representation in posit and fixed-posit formats for a wide variety of pre-trained deep neural networks (DNNs). We observe that fixed-posit representation is far more suitable for DNNs as it results in a faster and low-power computation circuit. We show that accuracy remains within the range of 0.3% and 0.57% of top-1 accuracy for posit and fixed-posit quantization. We further show that the posit-based multiplier requires higher power-delay-product (PDP) and area, whereas fixed-posit reduces PDP and area consumption by 71% and 36%, respectively, compared to (Devnathet al., 2020) for the same bit-width. Sumit Walia, Bachu Varun Tej, Arpita Kabra, Joydeep Kumar Devnath, Joycee Mekie |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2021 | FPCAM: Floating Point Configurable Approximate Multiplier for Error Resilient ApplicationsabstractIn this paper, we propose the design of a power-efficient floating point configurable approximate multiplier (FPCAM) suitable for error resilient applications. FPCAM allows systematic approximation, and the amount of approximation can be configured at run-time, depending on the error-tolerance of the applications. We show that compared to the existing state of the art multipliers, FPCAM on average consumes 62% lesser power and has 69% less power delay product. FPCAM also has 66% less area as compared to state of the art approximate multipliers. We have analyzed FPCAM for three different multimedia applications and random inputs to show that we achieve similar output quality as compared to existing multipliers while benefiting in power, area, and power delay product. Chandan Kumar Jha 0001, Sumit Walia, Gagan Kanojia, Joycee Mekie |
ISCAS | 4 |
| 2021 | Zero Aware Configurable Data Encoding by Skipping Transfer for Error Resilient ApplicationsabstractData transfer across DRAM channels accounts for nearly a quarter of the total energy consumption of DDR4 DRAMs. Modern applications with high bandwidth requirements further increase channel energy consumption. However, channel energy consumption is dependent on data being transferred. Pseudo Open Drain (POD) asymmetric termination, used in current DDR4 systems, consumes energy only when 1's are being transmitted over the channels. Many modern applications, including AI/ML ones are resilient to errors in data, and can work well with approximate data. This resilience can vary widely across and within applications, which provides a number of ways for exploiting these characteristics to save data transfer energy across the DRAM channel. However, all DRAM data encoding schemes have been targeted towards applications that require exact data and are not approximation resilient. In this paper, we propose Zero Aware Configurable Data Encoding by Skipping Transfer (ZAC-DEST), a data encoding scheme to reduce the energy consumption of DRAM channels, specifically targeted towards approximate computing and error resilient applications. ZAC-DEST exploits the similarity between recent data transfers across channels and information about error resilience behaviour of applications to reduce on-die termination and switching energy by reducing the number of 1's transmitted over the channels. ZAC-DEST also provides a number of knobs for trading off application's accuracy for energy savings, and vice versa, and can be applied to both training and inference. We apply ZAC-DEST to five machine learning applications. On average, across all applications and configurations, we observed a reduction of 40% in termination energy and 37% in switching energy as compared to the state of the art data encoding technique BD-Coder with an average output quality loss of 10%. We show that if both training and testing are done assuming the presence of ZAC-DEST, the output quality of the applications can be improved upto 9× as compared to when ZAC-DEST is only applied during testing leading to energy savings during training and inference with increased output quality. Chandan Kumar Jha 0001, Shreyas Singh, Riddhi Thakker, Manu Awasthi, Joycee Mekie |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2020 | Robust and High-Performance 12-T Interlocked SRAM for In-Memory ComputingabstractIn this paper, we analyze the existing SRAM based In-Memory Computing(IMC) proposals and show through exhaustive simulations that they fail under process variations. 6-T SRAM, 8-T SRAM, and 10-T SRAM based IMC architectures suffer from compute-disturb (stored data flips during IMC), compute-failure (provides false computation results), and half-select failures, respectively. To circumvent these issues, we propose a novel 12-T Dual Port Dual Interlockedstorage Cell (DPDICE) SRAM. DPDICE SRAM based IMC architecture(DPDICE-IMC) can perform essential boolean functions successfully in a single cycle and can perform basic arithmetic operations such as add and multiply. The most striking feature is that DPDICE-IMC architecture can perform IMC on two datasets simultaneously, thus doubling the throughput. Cumulatively, the proposed DPDICE-IMC is 26.7%, 8×, and 28% better than 6-T SRAM, 8-T SRAM, and 10-T SRAM based IMC architectures, respectively. Neelam Surana, Mili Lavania, Abhishek Barma, Joycee Mekie |
DATE | 4 |
| 2020 | ANSim: A Fast and Versatile Asynchronous Network-On-Chip SimulatorabstractIn a large chip, an asynchronous Network-on-Chip (NoC) is a suitable candidate for establishing an interconnection network between varied components. Architectural level simulation is an accepted methodology for evaluating such systems. In this paper, we propose a fast and versatile Asynchronous Network-on-Chip (NoC) Simulator - ANSim, which brings down the simulation time by 25 ×, compared to the state of the art simulators. It can model and analyze all synchronous, asynchronous, and mixed synchronous-asynchronous system of cores connected through NoC. ANSim can model routers with different delays, routers with asynchronous arbitration, connected in a wide range of topologies. ANSim supports individual routers modeled to have varying timing constraints. Further, it supports synthetic and real-workloads, and produces system-level latency, throughput, power, power-gating, and arbitration reports. ANSim has been verified against RTL models of NoCs, and other RTL verified simulators. An open-source synchronous NoC router, TNoC, and its asynchronous derivative are used to demonstrate ANSim's usefulness and features. Tom Glint, Jitesh Sah, Manu Awasthi, Joycee Mekie |
ICCD | 4 |
| 2020 | FPAD: A Multistage Approximation Methodology for Designing Floating Point Approximate DividersabstractApproximate computing has emerged as a unique proposition for error-resilient applications such as image/video processing, neural networks, and the like, where both performance and power can be simultaneously reduced by trading off output quality. In this paper we propose a multistage approximation methodology for designing IEEE 754 floating point approximate dividers (FPADs). We propose a number of FPADs with varying upper bounds on error. For the same mean error, FPAD is 2.84× better in terms of power-delay product (PDP) as compared to state of the art approximate floating point divider. Further, when applied on applications, such as image enhancement, mean filtering and JPEG compression, FPAD outperforms the existing state-of-the-art approximate divider in terms of PDP by 25%. We also show that FPAD gives same PDP benefits in Alexnet convolutional neural network without noticeable drop in top-5 and top-1 accuracy. Chandan Kumar Jha 0001, Kailash Prasad, Vibhor Kumar Srivastava, Joycee Mekie |
ISCAS | 4 |
| 2020 | SEDAAF: FPGA Based Single Exact Dual Approximate Adders for Approximate ProcessorsabstractApproximate circuits for ASICs have gained immense traction in recent years due to the benefits obtained in both energy and performance with little or no loss in output quality. Approximation in FPGAs remain a challenge due to the higher level of granularity at which logic is implemented on FPGAs. The smallest configurable blocks in FPGAs used for implementing logic consists of the look up tables (LUTs). In this paper, we exploit the inherent structures available in the FPGAs to implement SEDAAF. SEDAAF is a runtime configurable approximate adder that can perform a one-bit exact addition or two-bit approximate addition using the same hardware. SEDAAF also has a maximum bounded error, i.e. for an n-bit adder if m-bits are approximated the maximum error is 2m- 1. SEDAAF consumes 25% lesser power and has a 17% lesser power delay product as compared to existing designs. SEDAAF outperforms the existing state of the art designs in terms of output quality for Sobel edge detection application and can be used in approximate processors for performing both exact and approximate additions. Chandan Kumar Jha 0001, Kailash Prasad, Arun Singh Tomar, Joycee Mekie |
ISCAS | 4 |
| 2020 | A Low-Voltage Split Memory Architecture for Binary Neural NetworksabstractThis paper performs an in-depth study of error-resiliency of Neural Networks(NNs). Our investigation has resulted into two important findings. First, we found that Binary Neural Networks (BNNs) are more error-tolerant than 32-bits NNs. Second, in BNNs the network accuracy is more sensitive to errors in Batch Normalization Parameters(BNPs) than that in binary weights. A detailed discussion on the same is presented in the paper. Based on these findings, we propose a split memory architecture for low power BNNs, suitable for IoTs. In the proposed split memory architecture, weights are stored in area-efficient 6T SRAM, and BNPs are stored in robust 12T SRAM. The proposed split memory architecture for BNNs synthesized in UMC 28nm is highly energy efficient as the Vmin(minimum operating voltage) can be reduced to 0.36 V, 0.52 V, and 0.52 V for the MNIST, CIFAR10, and ImageNet datasets respectively, with accuracy drop of less than 1%. Joydeep Kumar Devnath, Neelam Surana, Joycee Mekie |
ISCAS | 3 |
| 2019 | SEDA - Single Exact Dual Approximate Adders for Approximate ProcessorsabstractApproximate computing has gained a lot of popularity due to its energy benefits in a variety of error-tolerant applications. In this paper we are proposing an adder which can perform n-bit single exact addition or dual approximate addition (SEDA), and is suitable for processors. The conversion from exact to approximate addition can be dynamically done at runtime. The maximum error is bounded for SEDA adders as carry is not approximated. Our proposed design consumes 48% lesser energy, has 32% lesser delay, occupies 24% lesser area as compared to exact mirror adder. Chandan Kumar Jha 0001, Joycee Mekie |
DAC | 2 |
| 2019 | Design of Novel CMOS Based Inexact Subtractors and Dividers for Approximate Computing: An In-Depth Comparison with PTL Based DesignsabstractMultimedia applications consume an immense amount of energy. These applications have division as one of the fundamental operations. Division is also one of the costliest operations in terms of energy consumption. Thus, various works have been done to address the issue of energy consumption in multimedia applications by using approximate dividers based on pass transistor logic (PTL). Since these applications have resilience towards erroneous computations huge energy benefits are obtained as a result of approximate computations with similar output quality. In this paper, we have shown that PTL based designs are not suitable for lower technology nodes. We performed an in-depth analysis using UMC 65nm and UMC 28nm to highlight the adverse effects of technology scaling on energy consumption and delay in PTL based design as compared to CMOS based designs. We also propose four different inexact CMOS subtractor (ICS) designs, as they are the basic repeated module in inexact restoring array dividers (IRADs). Our proposed ICS designs consume ~ 2× lesser dynamic energy, ~ 3× lesser static power and have ~ 2.5× lesser delay as compared to the existing PTL based designs in UMC 65nm. These benefits increase for UMC 28nm, which shows PTL based designs further worsens at lower technology nodes. IRADs also give about 50% reduction in energy consumption with only 3% degradation in Structural Similarity (SSIM) Index, an image quality metric in multimedia applications like change detection, background removal, and JPEG compression, as compared to exact restoring array divider (ERAD). Chandan Kumar Jha 0001, Joycee Mekie |
DSD | 2 |
| 2019 | PANE: Pluggable Asynchronous Network-on-Chip SimulatorabstractCommunication between different IP cores in MPSoCs and HMPs often results in clock domain crossing. Asynchronous network on chip (NoC) support communication in such heterogeneous set-ups. While there are a large number of tools to model NoCs for synchronous systems, there is very limited tool support to model communication for multi-clock domain NoCs and analyse them. In this article, we propose the P luggable A synchronous NE twork on Chip (PANE) simulator, which allows system-level simulation of asynchronous network on chip (NoC). PANE allows design space exploration of synchronous, asynchronous, and mixed synchronous-asynchronous(heterogeneous) NoC for various system-level NoC parameters such as packet latencies, throughput, network saturation point and power analysis. PANE supports a large range of NoC configurations—routing algorithms, topologies, network sizes, and so on—for both synthetic and real traffic patterns. We demonstrate the application of PANE by using synchronous routers, asynchronous routers, and a mix of asynchronous and synchronous routers. One of the key advantages of PANE is that it allows a seamless transition from synchronous to asynchronous NoC simulators while keeping pace with the developments in synchronous NoC tools as they can be integrated with PANE. Sneha N. Ved, Sarabjeet Singh, Joycee Mekie |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2006 | Reasoning about synchronization in GALS systems
Supratik Chakraborty, Joycee Mekie, Dinesh K. Sharma |
Formal Methods Syst. Des. | 2 |