VLDB 2026 Research / reviewers in the wild / expert
Rajiv V. Joshi
dblp:42/5976
· DBLP profile ↗
59ranked-venue papers
10as first author
18since 2021 · last 2026
0009-0007-7486-1531ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 58 · 10 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mixed-Mode ReRAM Architecture with Digital Weights and Inputs for High-Precision In-Memory Computing
Nithin Krishnan, Rajiv V. Joshi |
ISCAS | 2 |
| 2025 | Energy-Efficient Multi-Operand XOR Logic-Based CIM Accelerator using RRAM technologyabstractRecent advances in Resistive RAM (RRAM) based Computation-In-Memory (CIM) architectures highlight significant potential for accelerating data-intensive computing tasks. However, non-idealities in RRAM devices, such as variability, result in small sensing margins that can significantly affect the computational efficiency. This issue becomes even more pronounced when dealing with complex multi-operand logic operations. This paper introduces a circuit-level scheme for CIM-based multi-operand XOR logic operations, leveraging a Voltage-To-Time converter (VTC) to perform multi-phased XORs in a single clock cycle. In this approach, we exploit bitline capacitances for voltage-based sensing during computation, generating an output voltage that is linearly proportional to the operand values. This voltage is then converted into the desired logic output using the VTC design. Furthermore, low-power techniques are employed in the deployment of sense amplifiers, such as regulating power consumption during operation and disabling the amplifiers once the decision is made. Simulation results for a post-layout extracted 512x512 (256Kb) RRAM-based CIM array show that up to 16-operand XOR operation can be accurately and reliably performed as opposed to a maximum of three operands supported by state-of-the-art solutions, while offering up to 49× better figure-of-merit combining energy-efficiency and throughput. Abhairaj Singh, Konstantinos Stavrakakis, Rajendra Bishnoi, Rajiv V. Joshi, Said Hamdioui |
ICCAD | 4 |
| 2025 | Design Techniques for Ultra-low Power Cryogenic CMOS for Quantum Computing ApplicationsabstractThis paper describes design techniques for ultra-low power cryogenic CMOS circuits for next generation quantum computing applications. Specifically, it discusses two generations of fully integrated analog front-end circuits realized using end-to-end current mode design techniques to realize RF control of qubits at 4K ambient temperatures, presenting design The demonstrated current mode qubit state controller designs were implemented in 14nm FinFET CMOS technology; these designs consume 23.1mW and 12.8mW respectively. Sudipto Chakraborty, Pat Rosno, John F. Bulzacchelli, David J. Frank, Rajiv V. Joshi, Daniel J. Friedman |
ISLPED | 5 |
| 2024 | SRAM-Based Hybrid Analog Compute-In-memory Architecture to Enhance the Signal MarginabstractThis manuscript proposes an SRAM-based hybrid analog compute-in-memory (CIM) architecture to enhance the signal margin. This hybrid architecture presents fully differential current-based and C-2C charge-sharing-based multiplication and accumulation (MAC) CIM schemes for 4-bit MAC operation. The MAC operation of the filter weight's least significant bits (w0and w1) is implemented in the current-based CIM. However, the MAC operation of the most significant bits (w2and w3) is implemented in the charge-based CIM. The proposed architecture achieves a 4.37× enhancement in signal margin compared to the state-of-the-art. The energy efficiency and throughput of the proposed architecture are 1551.5 TOPS/W and 512 GOPS, respectively, at 0.9 V supply voltage and 250 MHz frequency. A convolutional neural network (CNN) is implemented on the proposed architecture, and the inference accuracy for the MNIST and CIFAR-10 data sets is 98.6 % and 86 %, respectively. The proposed architecture is scalable for multi-bit MAC operation and implemented in 28 nm CMOS technology. Dinesh Kushwaha, Rajiv V. Joshi, Sudeb Dasgupta, Bulusu Anand |
ISCAS | 2 |
| 2024 | AFSRAM-CIM: Adder Free SRAM-Based Digital Computation-in-Memory for BNNabstractBinary Neural Networks (BNNs) have demonstrated significant advantages in reducing computation and memory costs, all while maintaining acceptable accuracy on various image detection tasks. Thus, BNNs have the potential to support practical cognitive tasks on resource-constrained platforms, such as edge computing devices. To realize this, SRAM-based digital Computation-in-Memory (CIM) has gained growing attention as it overcomes the analog CIM architecture bottlenecks such as limited computing accuracy due to process variation, non-linearity, power and area-hungry Analog-to-Digital Converters (ADCs), etc. However, digital CIM architectures are highly dominated by power-hungry adder-trees, which can nullify the benefits of SRAM-based digital CIM. To address this issue, this paper proposes an adder free SRAM-based digital CIM, AFSRAM-CIM, for BNN acceleration. The proposed CIM architecture utilizes a multi-functional 10-T SRAM cell-based crossbar array and a new energy-efficient approach to perform the popcount operation. Simulation results using the MNIST dataset show that the proposed architecture maintains the state-of-the-art inference accuracy of 99.21% with only 11.86 fJ energy per operation. Moreover, AFSRAM-CIM achieves over$3\times$energy and$\approx 17\times$area savings when compared to the conventional digital CIM approaches. Asmae El Arrassi, Mohammad Amin Yaldagard, Xingjian Tao, Taha Shahroodi, Fouwad Jamil Mir, Yashvardhan Biyani, Manil Dev Gomony, Anteneh Gebregiorgis, Rajiv V. Joshi, Said Hamdioui |
VLSI-SoC | 9 |
| 2023 | (Invited) Predictive analytics for cryogenic CMOS in future quantum computing systemsabstractThis paper presents predictive techniques that suggest a path to accelerated analysis and optimization of the yield of ultra-low power analog/mixed signal designs operating at cryogenic temperatures for use in future quantum computing applications. Analysis of 6σ variation is accelerated using mixture importance sampling (MixIS) techniques, with key studied specifications being compliance to spurious tone requirements at the quantum state controller output and optimization of design power consumption, both in the context of the high levels of modeling and matching uncertainties associated with cryogenic circuit design. Rajiv V. Joshi, Sudipto Chakraborty |
DAC | 1 |
| 2023 | A Best Balance Ratio Ordered Feature Selection Methodology for Robust and Fast Statistical Analysis of Memory DesignsabstractRecently, machine learning yield models for integrated circuit (IC) have gained widespread prominence in the EDA community, and are very promising in terms of emulating memory design functionality and thereby speeding up circuit simulation-based variance reduction methods. A main challenge that arises in this area is a class imbalance that occurs naturally due to the high targeted manufacturing yield. Thus, the imbalanced nature of the sampled memory datasets can compromise the model performance. In this work, we attain deep insights into the memory classification problem for modeling rare fail events in the context of importance sampling-based yield analysis. We propose a comprehensive and computationally efficient method that addresses the joint considerations of the best combination of relevant features and class balance ratios, which are key for classifier generalization capability. The methodology relies on synthetic minority over-sampling techniques to enforce the minority class while probing for the best data balance ratio in conjunction with an iterative$L_{1}$-SVM-based approach that qualifies as an approximation to the$L_{0}$-norm regularization for the best feature subset selection. We compare the proposed methodology against standalone$L_{1}$-SVM solutions, unbalanced$L_{0}$-norm approximation as well as an algorithmic data balancing method in the context of yield estimation methodology. The methodology is shown to result in high fidelity classifiers as demonstrated when analyzing the yield of a 14-nm FinFET SRAM cross-section with speedup of$179\times $for the importance sampling simulations compared to pure circuit simulation-based approaches and an average error of$0.19 \sigma $. Lama Shaer, Rouwaida Kanj, Rajiv V. Joshi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Referencing-in-Array Scheme for RRAM-based CIM ArchitectureabstractResistive random access memory (RRAM) based computation-in-memory (CIM) architectures are attracting a lot of attention due to their potential in performing fast and energy-efficient computing. However, the RRAM variability and non-idealities limit the computing accuracy of such architectures, especially for multi-operand logic operations. This paper pro-poses a voltage-based differential referencing-in-array scheme that enables accurate two and multi-operand logic operations for RRAM-based CIM architecture. The scheme makes use of a 2T2R cell configuration to create a complementary bitcell structure that inherently acts also as a reference during the operation execution; this results in a high sensing margin. More-over, the variation-sensitive multi-operand (N)AND operation is implemented using complementary-input (N)OR operation to further improve its accuracy. Simulation results for a post-layout extracted 512x512 (256Kb) RRAM-based CIM array show that up to 56 operand (N)OR/(N)AND operation can be accurately and reliably performed as opposed to a maximum of 4 operands supported by state-of-the-art solutions, while offering up to 11.4X better energy-efficiency. Abhairaj Singh, Rajendra Bishnoi, Rajiv V. Joshi, Said Hamdioui |
DATE | 3 |
| 2022 | A Cryo-CMOS Transmon Qubit Controller and Verification with FPGA EmulationabstractFuture generations of quantum computers are expected to operate in a paradigm where multi-qubit devices will predominantly perform circuits to support quantum error correction. Highly integrated cryogenic electronics are a key enabling technology to support the control of the large numbers of physical qubits that will be required in this fault-tolerant, error-corrected regime. Here, we describe our perspectives on cryoelectronics-driven qubit control architectures, and will then describe an implementation of a scalable, low-power, cryogenic qubit state controller that includes a domain-specific processor and a SSB upconversion I/Q-mixer-based RF AWG. We will also describe an FPGA-based emulation platform that is able to closely reproduce the system intention, and which was used to verify different aspects of the ASIC system design in in situ transmon qubit control experiments. Kevin Tien, Ken Inoue, Scott Lekuch, David J. Frank, Sudipto Chakraborty, Pat Rosno, Thomas Fox, Mark Yeck, Joseph A. Glick, Raphael Robertazzi, Ray Richetta, John F. Bulzacchelli, Daniel Ramirez, Dereje Yilma, Andrew Davies, Rajiv V. Joshi, Devin Underwood, Dorothy Wisnieff, Christian W. Baks, Donald Bethune, John Timmerwilke, Blake R. Johnson, Brian P. Gaucher, Daniel J. Friedman |
DATE | 16 |
| 2022 | A 65nm Compute-In-Memory 7T SRAM Macro Supporting 4-bit Multiply and Accumulate Operation by Employing Charge SharingabstractIn this work, we propose an energy-efficient 64$\times $ 64 compute-in-memory (CIM) SRAM macro using a 7T bit-cell in 65nm CMOS UMC PDK. It supports 4-bit inputs, 4-bit weights & 4-bit outputs and performs 4-bit MAC operations. It also supports multiple row activations performing 1024 4b$\times $4b multiply and accumulate (MAC) operations in one clock cycle. Inputs are realized by the number of pulses on the read wordline (RWL), which discharges read bitline (RBL) according to bitwise multiplication of weights & inputs. Outputs of 4 columns storing 4-bit weights are then combined via charge sharing to perform a binary-weighted average representing MAC operation, further quantized by a flash analog to digital converter (ADC) giving 4-bit output. The proposed CIM macro achieves an energy efficiency of 28.9 TOPS/W and throughput of 212.9 GOPS operating at supply voltage 1 V with a 2 GHz clock frequency. Dinesh Kushwaha, Ritik Raj, Ashish Joshi, Jwalant Mishra, Rajat Kohli, Sandeep Miryala, Rajiv V. Joshi, Sudeb Dasgupta, Bulusu Anand |
ISCAS | 9 |
| 2022 | Accelerating RRAM Testing with a Low-cost Computation-in-Memory based DFTabstractEmerging non-volatile resistive RAM (RRAM) device technology has shown great potential to cultivate not only high-density memory storage, but also energy-efficient computing units. However, the unique challenges related to RRAM fabrication process render the traditional memory testing solutions inefficient and inadequate for high product quality. This paper presents low-cost design-for-testability (DFT) solutions that augment the testing process and improve the fault coverage. A computation-in-memory (CIM) based DFT is realized to expedite the detection and diagnosis of faults by developing logic designs involving multi-row activation. A novel addressing scheme is introduced to facilitate the diagnosis of faults. Reconfigurable logic designs are developed to detect unique RRAM faults that offer features such as programmable reference generations, period, and voltage of operation. DFT implementations are validated on a post-layout extracted platform and testing sequences are introduced by incorporating the proposed DFTs. Results show that more than 2.3× speedup and better coverage are achieved with 6× area reduction when compared with state-of-the-art solutions. Abhairaj Singh, Moritz Fieback, Rajendra Bishnoi, Filip Bradaric, Anteneh Gebregiorgis, Rajiv V. Joshi, Said Hamdioui |
ITC | 6 |
| 2022 | A Voltage-Controlled, Oscillation-Based ADC Design for Computation-in-Memory Architectures Using Emerging ReRAMsabstractConventional von Neumann architectures cannot successfully meet the demands of emerging computation and data-intensive applications. These shortcomings can be improved by embracing new architectural paradigms using emerging technologies. In particular, Computation-In-Memory (CiM) using emerging technologies such as Resistive Random Access Memory (ReRAM) is a promising approach to meet the computational demands of data-intensive applications such as neural networks and database queries. In CiM, computation is done in an analog manner; digitization of the results is costly in several aspects, such as area, energy, and performance, which hinders the potential of CiM. In this article, we propose an efficient Voltage-Controlled-Oscillator (VCO)–based analog-to-digital converter (ADC) design to improve the performance and energy efficiency of the CiM architecture. Due to its efficiency, the proposed ADC can be assigned in a per-column manner instead of sharing one ADC among multiple columns. This will boost the parallel execution and overall efficiency of the CiM crossbar array. The proposed ADC is evaluated using a Multiplication and Accumulation (MAC) operation implemented in ReRAM-based CiM crossbar arrays. Simulations results show that our proposed ADC can distinguish up to 32 levels within 10 ns while consuming less than 5.2 pJ of energy. In addition, our proposed ADC can tolerate ≈30% variability with a negligible impact on the performance of the ADC. Mahta Mayahinia, Abhairaj Singh, Christopher Bengel, Stefan Wiefels, Muath Abu Lebdeh, Stephan Menzel, Dirk J. Wouters, Anteneh Gebregiorgis, Rajendra Bishnoi, Rajiv V. Joshi, Said Hamdioui |
ACM J. Emerg. Technol. Comput. Syst. | 10 |
| 2022 | Exploring Model Stability of Deep Neural Networks for Reliable RRAM-Based In-Memory AccelerationabstractRRAM-based in-memory computing (IMC) effectively accelerates deep neural networks (DNNs). Furthermore, model compression techniques, such as quantization and pruning, are necessary to improve algorithm mapping and hardware performance. However, in the presence of RRAM device variations, low-precision and sparse DNNs suffer from severe post-mapping accuracy loss. To address this, in this work, we investigate a new metric,model stability, from the loss landscape to help shed light on accuracy loss under variations and model compression, which guides an algorithmic solution to maximize model stability and mitigate accuracy loss. Based on statistical data from a CMOS/RRAM 1T1R test chip at 65nm, we characterize wafer-level RRAM variations and develop a cross-layer benchmark tool that incorporates quantization, pruning, device variations, model stability, and IMC architecture parameters to assess post-mapping accuracy and hardware performance. Leveraging this tool, we show that a loss-landscape-based DNN model selection for stability effectively tolerates device variations and achieves a post-mapping accuracy higher than that with 50% lower RRAM variations. Moreover, we quantitatively interpret why model pruning increases the sensitivity to variations, while a lower-precision model has better tolerance to variations. Finally, we propose a novel variation-aware training method to improve model stability, in which there exists the most stable model for the best post-mapping accuracy of compressed DNNs. Experimental evaluation of the method shows up to 19%, 21%, and 11% post-mapping accuracy improvement for our 65nm RRAM device, across various precision and sparsity, on CIFAR-10, CIFAR-100, and SVHN datasets, respectively. Li Yang 0009, Jingbo Sun 0003, Jubin Hazra, Xiaocong Du, Maximilian Liehr, Zheng Li 0020, Karsten Beckmann, Rajiv V. Joshi, Nathaniel C. Cady, Deliang Fan, Yu Cao 0001 |
IEEE Trans. Computers | 9 |
| 2022 | Hybrid RRAM/SRAM in-Memory Computing for Robust DNN AccelerationabstractRRAM-based in-memory computing (IMC) effectively accelerates deep neural networks (DNNs) and other machine learning algorithms. On the other hand, in the presence of RRAM device variations and lower precision, the mapping of DNNs to RRAM-based IMC suffers from severe accuracy loss. In this work, we propose a novel hybrid IMC architecture that integrates an RRAM-based IMC macro with a digital SRAM macro using a programmable shifter to compensate for the RRAM variations and recover the accuracy. The digital SRAM macro consists of a small SRAM memory array and an array of multiply-and-accumulate (MAC) units. The nonideal output from the RRAM macro, due to device and circuit nonidealities, is compensated by adding the precise output from the SRAM macro. In addition, the programmable shifter allows for different scales of compensation by shifting the SRAM macro output relative to the RRAM macro output. On the algorithm side, we develop a framework for the training of DNNs to support the hybrid IMC architecture through ensemble learning. The proposed framework performs quantization (weights and activations), pruning, RRAM IMC-aware training, and employs ensemble learning through different compensation scales by utilizing the programmable shifter. Finally, we design a silicon prototype of the proposed hybrid IMC architecture in the 65-nm SUNY process to demonstrate its efficacy. Experimental evaluation of the hybrid IMC architecture shows that the SRAM compensation allows for a realistic IMC architecture with multilevel RRAM cells (MLCs) even though they suffer from high variations. The hybrid IMC architecture achieves up to 21.9%, 12.65%, and 6.52% improvement in post-mapping accuracy over state-of-the-art techniques, at minimal overhead, for ResNet-20 on CIFAR-10, VGG-16 on CIFAR-10, and ResNet-18 on ImageNet, respectively. Zhenyu Wang 0016, Injune Yeo, Li Yang 0009, Jian Meng, Maximilian Liehr, Rajiv V. Joshi, Nathaniel C. Cady, Deliang Fan, Jae-sun Seo, Yu Cao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2022 | Group LARS-Based Iterative Reweighted Least Squares Methodology for Efficient Statistical Modeling of Memory DesignsabstractRegularized logistic regression is a popular classification tool that can be employed to accurately model the binary nature of the memory cell fail mechanisms for purposes of the yield analysis of memory designs. The iterative reweighted least squares (IRLS) method has been employed along with the least angle regression (LARS) to efficiently solve the$L_{1}$regularized logistic regression problem. In this brief, we propose an efficient$L_{1}$regularized logistic regression methodology. At the core lies a Group LARS-based approach that benefits from Group LARS inherent ability to handle groups of variables and exploits the natural evolution of the solution to speed up the search for the critical features of the classifier. Thus, it tracks Newton’s step direction from one round of the solution to the next and employs weighted directions to efficiently solve for the underlying$L_{1}$constrained iterative least squares problem. We apply the methodology in the context of an importance sampling-based yield analysis framework targeting rare fail probability estimation. We study the yield of 14-nm FinFET SRAM designs with programmable and resonant boosting. Our results demonstrate up to$14\times $–$20\times $speedup for the Group LARS compared to the pure LARS-based approach, and we report 98.7% accuracy and 0.12$\sigma $average error compared to pure circuit-simulations approach for the resulting classifier. Lama Shaer, Rouwaida Kanj, Rajiv V. Joshi, Ali Chehab |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Phase Noise Analysis of Separately Driven Ring OscillatorsabstractIn this paper, for the first time, the phase noise analysis of a Multi-loop Skew based Single Ended Oscillator (MSSROs) is derived and validated. Compared to the three stages of conventional ring oscillators (CROs), SDROs provide an equivalent oscillation frequency with improved phase noise with increasing stages. The primary distinction between these two designs (SDRO and three-stage CROs) is the inherent skew offset between the PMOS/NMOS gates caused by the unique connection. This skew offset is the fundamental cause of delay cell noise suppression; the SDROs have loosely coupled oscillators that run concurrently, forming multiple 3-stages of separately driven Ring Oscillators. As a result, a shaping function is derived in terms of skew offset, and simulating these with varying skew offset results in suppressing behavior. Additionally, we derived phase noise for a skew-based design and validated it in PDKs of 180nm and 65 nm. We plotted the thermal (flicker) noise contribution and found that increasing the number of stages leads to an approximately 1-2 dB reduction in phase noise while maintaining the same NMOS/PMOS size ratio. Finally, a 2-3 dB reduction in phase noise is achieved in MSSROs by incorporating the shaping function into phase noise equations. Neeraj Mishra, Anchit Proch, Lomash Chandra Acharya, Jeffrey Prinzie, Sudipto Chakraborty, Rajiv V. Joshi, Sudeb Dasgupta, Bulusu Anand |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2021 | Low-Power Memristor-Based Computing for Edge-AI ApplicationsabstractWith the rise of the Internet of Things (IoT), a huge market for so-called smart edge-devices is foreseen for millions of applications, like personalized healthcare and smart robotics. These devices have to bring smart computing directly where the data is generated, while coping with the limited energy budget. Conventional von-Neumann architecture fail to meet these requirements due to e.g., memory-processor data transfer bottleneck. Memristor-based computation-in-memory (CIM) has the potential to realize smart local computing for highly parallel data-dominated AI applications by exploiting the inherent properties of the architecture and the physical characteristics of the memristors. This paper provides a broad overview of CIM architecture highlighting its potential and unique properties in enabling smart local computing. Moreover, it discusses design considerations of such architectures including both crossbar array as well as peripheral circuits; special attention is given to analog-to-digital converter (ADC), as it is the most critical unit of analog-based CIM operation e.g., vector-matrix multiplication (VMM). Finally, the paper outlines the potential future directions for CIM-based edge smart computing. Abhairaj Singh, Sumit Diware, Anteneh Gebregiorgis, Rajendra Bishnoi, Francky Catthoor, Rajiv V. Joshi, Said Hamdioui |
ISCAS | 6 |
| 2021 | SRIF: Scalable and Reliable Integrate and Fire Circuit ADC for Memristor-Based CIM ArchitecturesabstractEmerging computation-in-memory (CIM) paradigm offers processing and storage of data at the same physical location, thus alleviating critical memory-processor communication bottlenecks suffered by conventional von-Neumann architecture. Storage of data in a CIM architecture is analog in nature and therefore computation is performed in analog domain i.e. inputs and outputs are analog values. Since the outside computing environment is digital, analog-to-digital converters (ADC) are utilized to perform the output data conversion. However, ADC designs are bulky, power-hungry circuits that are prone to design variations and therefore, play an important role in determining the computing efficiency of CIM architectures. In this paper, we present a scalable and reliable integrate and fire circuit ADC (SRIF-ADC) design for CIM architectures, suitable for stringent power and area constraints. We devise a technique to stabilize the node receiving analog inputs that allows more rows to be activated at the same time, thereby increasing the operand size of input vectors. This allows better scalability in terms of higher parallelism of operations. We employ a self-timed variation-aware design approach and design measures to drastically reduce read disturb of memristor devices that address reliability issues related to the ADC design. In addition, we present a compact, built-in sample-and-hold circuit to replace the large-sized capacitance and built-in weighting technique to alleviate the need for post-processing. For multiply-and-accumulate (MAC) operation, our simulation results show that we can improve the computational parallelism by 3X as well as ADC conversion speed and energy efficiency are improved by 2X and 11.6X, respectively, compared to the state-of-the-art design. Abhairaj Singh, Muath Abu Lebdeh, Anteneh Gebregiorgis, Rajendra Bishnoi, Rajiv V. Joshi, Said Hamdioui |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2020 | Accurate Inference with Inaccurate RRAM Devices: Statistical Data, Model Transfer, and On-line AdaptationabstractResistive random-access memory (RRAM) is a promising technology for in-memory computing with high storage density, fast inference, and good compatibility with CMOS. However, the mapping of a pre-trained deep neural network (DNN) model on RRAM suffers from realistic device issues, especially the variation and quantization error, resulting in a significant reduction in inference accuracy. In this work, we first extract these statistical properties from 65 nm RRAM data on 300mm wafers. The RRAM data present 10-levels in quantization and 50% variance, resulting in an accuracy drop to 31.76% and 10.49% for MNIST and CIFAR-10 datasets, respectively. Based on the experimental data, we propose a combination of machine learning algorithms and on-line adaptation to recover the accuracy with the minimum overhead. The recipe first applies Knowledge Distillation (KD) to transfer an ideal model into a student model with statistical variations and 10 levels. Furthermore, an on-line sparse adaptation (OSA) method is applied to the DNN model mapped on to the RRAM array. Using importance sampling, OSA adds a small SRAM array that is sparsely connected to the main RRAM array; only this SRAM array is updated to recover the accuracy. As demonstrated on MNIST and CIFAR-10 datasets, a 7.86% area cost is sufficient to achieve baseline accuracy for the 65 nm RRAM devices. Gouranga Charan, Jubin Hazra, Karsten Beckmann, Xiaocong Du, Rajiv V. Joshi, Nathaniel C. Cady, Yu Cao 0001 |
DAC | 6 |
| 2019 | Resilient Low Voltage Accelerators for High Energy EfficiencyabstractLow voltage architecture and design are key enablers of high throughput per watt in heterogeneous, accelerator-rich many-core designs. However, such low voltage operation poses significant challenges because of difficulties in achieving reliable functionality of on-chip memories, particularly SRAMs at these design points. In this paper, we present a technique of low-voltage neural network acceleration, where the embedded SRAM architecture is equipped with a novel application-aware supply voltage boosting capability. This technique mitigates low-voltage induced failures, while enabling Very low voltage (VLV)1operation during most of the application run, resulting in substantial improvement in net energy efficiency. We present a framework to evaluate the impact of low-voltage SRAM errors on machine learning applications and characterize trade-offs between output inference accuracy and energy efficiency in our application-programmable supply boosted SRAM architecture. Using the proposed technique we push the limits on the minimum operable voltage (Vmin) for the desired output quality. As a proof of concept, we demonstrate these techniques on Dante, a Deep Neural Network (DNN) accelerator chip taped out in state-of-the art 14nm technology. Nandhini Chandramoorthy, Karthik Swaminathan, Martin Cochet, Arun Paidimarri, Schuyler Eldridge, Rajiv V. Joshi, Matthew M. Ziegler, Alper Buyuktosunoglu, Pradip Bose |
HPCA | 6 |
| 2019 | Data Imbalance Handling Approaches for Accurate Statistical Modeling and Yield Analysis of Memory DesignsabstractData imbalance can impact the fidelity of a classifier. We rely on advances in data imbalance handling techniques for machine learning applications to propose an enhanced fast statistical analysis methodology. Particularly, we employ data handling techniques in the context of a logistic regression based importance sampling methodology for accurate statistical modeling of rare fail events in memory designs. We demonstrate that for purposes of achieving conservative yield estimates, the synthetic minority oversampling technique outperforms other data handling methods and portrays the best model recall and precision rates. We report more than 70% reduction in the number of False Negatives compared to imbalanced data set based approaches. We also report on average a low 5% relative error rate in the yield estimate for the balanced data set-based modeling approaches compared to the pure circuit simulation based approach. This is compared to on average an 18% relative error rate obtained for the imbalanced data set-based approaches. These results were verified on state-of-the-art industrial FinFET SRAM designs. Lama Shaer, Rouwaida Kanj, Rajiv V. Joshi |
ISCAS | 3 |
| 2019 | Big vs little core for energy-efficient Hadoop computing
Maria Malik, Katayoun Neshatpour, Setareh Rafatirad, Rajiv V. Joshi, Tinoosh Mohsenin, Hassan Ghasemzadeh 0001, Houman Homayoun |
J. Parallel Distributed Comput. | 4 |
| 2019 | Editorial TVLSI Positioning - Continuing and Accelerating an Upward TrajectoryabstractI. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5]. Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 23 |
| 2018 | Sparse Regression Driven Mixture Importance Sampling for Memory DesignabstractIn this paper, we present a sparse regression (SpaRe) model-based yield analysis methodology and apply it to memory designs with state-of-the-art write-assist circuitry. At the core of its engine is a mixture importance sampling technique which consists of a uniform sampling stage and an importance sampling stage. The proposed methodology allows for fast and accurate statistical analysis of rare fail events. In our approach, a SpaRe model is built using the uniform sampling stage data points obtained via circuit simulation (CktSim). Along with the model, an optimal threshold value is determined for proper pass/fail predict capability. The model and the threshold value are then used to predict the response in the importance sampling stage. This alleviates the need for CktSims in the latter stage and introduces significant speedup compared to fully CktSim-based approaches. The SpaRe model-based yield analysis is tested on a 14-nm FinFET SRAM design, and the results corroborate well with that of full CktSim-based yield analysis. The methodology is used to compare multiple state-of-the-art SRAM designs including selective boost and write-assist designs. The operating Vmin ranges and trends corroborate well with hardware measurements. Maria Malik, Rajiv V. Joshi, Rouwaida Kanj, Shupeng Sun, Houman Homayoun |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Very Low Voltage (VLV) DesignabstractThis paper is a tutorial-style introduction to a special session on: Effective Voltage Scaling in the Late CMOS Era. It covers the fundamental challenges and associated solution strategies in pursuing very low voltage (VLV) designs. We discuss the performance and system reliability constraints that are key impediments to VLV. The associated trade-offs across power, performance and reliability are helpful in inferring the optimal operational voltage-frequency point. This work was performed under the auspices of an ongoing DARPA program (named PERFECT) that is focused on maximizing system-level energy efficiency. Ramon Bertran Monfort, Pradip Bose, David Brooks 0001, Jeff Burns, Alper Buyuktosunoglu, Nandhini Chandramoorthy, Eric Cheng, Martin Cochet, Schuyler Eldridge, Daniel J. Friedman, Hans M. Jacobson, Rajiv V. Joshi, Subhasish Mitra, Robert K. Montoye, Arun Paidimarri, Pritish Parida, Kevin Skadron, Mircea R. Stan, Karthik Swaminathan, Augusto Vega, Swagath Venkataramani, Christos Vezyrtzis, Gu-Yeon Wei, John-David Wellman, Matthew M. Ziegler |
ICCD | 12 |
| 2017 | Distributed In-Memory Computing on Binary RRAM CrossbarabstractThe recently emerging resistive random-access memory (RRAM) can provide nonvolatile memory storage but also intrinsic computing for matrix-vector multiplication, which is ideal for the low-power and high-throughput data analytics accelerator performed in memory. However, the existing RRAM crossbar--based computing is mainly assumed as a multilevel analog computing, whose result is sensitive to process nonuniformity as well as additional overhead from AD-conversion and I/O. In this article, we explore the matrix-vector multiplication accelerator on a binary RRAM crossbar with adaptive 1-bit-comparator--based parallel conversion. Moreover, a distributed in-memory computing architecture is also developed with the according control protocol. Both memory array and logic accelerator are implemented on the binary RRAM crossbar, where the logic-memory pair can be distributed with the control bus protocol. Experimental results have shown that compared to the analog RRAM crossbar, the proposed binary RRAM crossbar can achieve significant area savings with better calculation accuracy. Moreover, significant speedup can be achieved for matrix-vector multiplication in neural network--based machine learning such that the overall training and testing time can be both reduced. In addition, large energy savings can be also achieved when compared to the traditional CMOS-based out-of-memory computing architecture. Leibin Ni, Hantao Huang, Zichuan Liu, Rajiv V. Joshi, Hao Yu 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2017 | EditorialabstractAs I start my second two-year term (2017–2018) as the Editor-in-Chief (EIC) of the IEEE Transactions on Very Large Scale Integration Systems (TVLSI), I wish the TVLSI readership a very happy new year and continued professional success. It gives me great pleasure to report on the state of the journal and our performance metrics. Over the past two years, TVLSI has seen a healthy increase in the number of submissions—from 687 in 2014 to 770 in 2015, and at the time of writing of this editorial, we are at 760 submissions for 2016. We expect the number of submissions for 2016 to cross 800 before the end of the year. TVLSI, therefore, continues to be the premier archival journal for university researchers and industry practitioners in the broad area of VLSI system design. Krishnendu Chakrabarty, Massimo Alioto, Bevan M. Baas, Chirn Chye Boon, Meng-Fan Chang, Naehyuck Chang, Yao-Wen Chang, Chip-Hong Chang, Shih-Chieh Chang 0001, Poki Chen, Masud H. Chowdhury, Pasquale Corsonello, Ibrahim M. Elfadel, Said Hamdioui, Masanori Hashimoto, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Rajiv V. Joshi, Tanay Karnik, Mehran Mozaffari Kermani, Chulwoo Kim, Jaydeep P. Kulkarni, Eren Kursun, Erik Larsson, Hai Li 0001, Huawei Li 0001, Patrick P. Mercier, Prabhat Mishra 0001, Makoto Nagata, Arun Natarajan 0001, Koji Nii, Partha Pratim Pande, Ioannis Savidis, Mingoo Seok, Sheldon X.-D. Tan, Mark Tehranipoor, Aida Todri, Miroslav N. Velev, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 19 |
| 2017 | EditorialabstractWe are pleased to announce this year’s winners of the IEEE Transactions on Very Large Scale integrated (VLSI) Systems (TVLSI) Circuits and Systems (CAS) Society Best Reviewer and Associate Editor Awards. We had a number of qualified candidates for each award, but after a thorough evaluation process including input from the Associate Editors and Selection Committee, these five individuals stood out among all the candidates: Krishnendu Chakrabarty, Massimo Alioto, Rajiv V. Joshi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Big biomedical image processing hardware acceleration: A case study for K-means and image filteringabstractMost hospitals today are dealing with the big data problem, as they generate and store petabytes of patient records most of which in form of medical imaging, such as pathological images, CT scans and X-rays in their datacenters. Analyzing such large amounts of biomedical imaging data to enable discovery and guide physicians in personalized care is becoming an important focus of data mining and machine learning algorithms developed for biomedical Informatics (BMI). Algorithms that are developed for BMI heavily rely on complex and computationally intensive machine learning and data mining methods to learn from large data. The high processing demand of big biomedical imaging data has given rise to their implementation in high-end server platforms running software ecosystems that are optimized for dealing with large amount of data including Apache Hadoop and Apache Spark. However, efficient processing of such large amount of imaging data running computational intensive learning methods is becoming a challenging problem using state-of-the-art high performance computing server architectures. To address this challenge, in this paper, we introduce a scalable and efficient hardware acceleration method using low cost commodity FPGAs that is interfaced with a server architecture through a high speed interface. In this work we present a full end-to-end implementation of big data image processing and machine learning applications in a heterogeneous CPU+FPGA architecture. We develop the MapReduce implementation of K-means and Laplacian Filtering in Hadoop Streaming environment that allows developing mapper functions in non-Java based languages suited for interfacing with FPGA-based hardware accelerating environment. We accelerate the mapper functions through hardware+software (HW+SW) co-design. We do a full implementation of the HW+SW mappers on the Zynq FPGA platform. The results show promising kernel speedup of up to 27× for large image data sets. This translate to 7.8× and 1.8× speedup in an end-to-end Hadoop MapReduce implementation of K-mean s and Laplacian Filtering algorithm, respectively. Katayoun Neshatpour, Arezou Koohi, Farnoud Farahmand, Rajiv V. Joshi, Setareh Rafatirad, Avesta Sasan, Houman Homayoun |
ISCAS | 4 |
| 2016 | Characterizing Hadoop applications on microservers for performance and energy efficiency optimizationsabstractThe traditional low-power embedded processors such as Atom and ARM are entering the high-performance server market. At the same time, as the size of data grows, emerging Big Data applications require more and more server computational power that yields challenges to process data energy-efficiently using current high performance server architectures. Furthermore, physical design constraints, such as power and density have become the dominant limiting factor for scaling out servers. Numerous big data applications rely on using the Hadoop MapReduce framework to perform their analysis on large-scale datasets. Since Hadoop configuration parameters as well as architecture parameters directly affect the MapReduce job performance and energy-efficiency, system and architecture level parameters tuning is vital to maximize the energy efficiency. In this work, through methodical investigation of performance and power measurements, we demonstrate how the interplay among various Hadoop configurations and system and architecture level parameters affect the performance and energy-efficiency across various Hadoop applications. Maria Malik, Avesta Sasan, Rajiv V. Joshi, Setareh Rafatirad, Houman Homayoun |
ISPASS | 3 |
| 2016 | A Universal Hardware-Driven PVT and Layout-Aware Predictive Failure Analytics for SRAMabstractThe impact of device variability, temperature, and technology CAD-based layout parasitics on low-voltage static random access memory (SRAM) yield is explored using a novel variability-aware statistical methodology. Threshold voltage, Vt, mismatches for planar 22- and 14-nm FinFET SRAM transistors are characterized based on unique array-like structures for capturing process voltage and temperature (PVT) impact on variability. In general, the mismatches are shown to be a consistent and unique function of Vdd, doping, and temperature across the two technologies. Stronger Vt mismatch impact is observed as a function of Vddand doping in the 22-nm technology, with higher mismatch recorded at lower temperatures. In the 14-nm technology, doping is found to have the strongest impact on Vtmismatch, and the mismatch increases with Vdddespite the reduced drain induced barrier lowering effects. Similar to the 22-nm technology, the mismatch increases at lower temperatures. Front-end-of-the line capacitance effects are found to be more significant than back-end-of-the-line effects in 14-nm technologies, as opposed to planar technologies. Accurate parasitic capacitance modeling along with PVT-aware variability process variations for different 22-/14-nm cell arrangements are incorporated into a physics based statistical analysis methodology for accurate Vmin analysis. The yield analysis results are corroborated with hardware yield using 4-16-Mb inline SRAM macro monitors. The methodology is unique in the industry, gives insight into the technology-circuit interactions, and is able to effectively predict the SRAM yield bounds. Rajiv V. Joshi, Sudesh Saroop, Rouwaida Kanj, Carl Radens, Karthik Yogendra |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | Corrections to "Super Fast Physics-Based Methodology for Accurate Memory Yield Prediction"abstractOn[1, p. 536], Section II-A, second paragraph, lines 5–8, [2] should be used instead of[1]and the correct statement is as follows. “Model-to-hardware corroboration shows an excellent matching between importance-sampling-based methods yield estimation [2] and the true hardware yield. We therefore adopt the methodology in [2] as the core statistical engine for our TfM methodology.” Rajiv V. Joshi, Rouwaida Kanj |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | Super Fast Physics-Based Methodology for Accurate Memory Yield PredictionabstractWe propose an efficient physics-based mixed-mode statistical simulation methodology for nanoscale devices and circuits. Here, 3-D Technology Computer Aided Design models pose a barrier for efficient simulation of variability as they generally involve millions of nodes in their mesh representations. The proposed methodology, which has been implemented for FinFET/tri-gate static random access memory (SRAM) design, overcomes this barrier by leveraging advanced physics-based 2-D (P2-D) devices with optimized meshes that are derived from 3-D FinFET models with tuned device parasitics. This enables physics-based simulation as well as physics-based variability input parameters. To improve accuracy, an embedded automated flow enables extraction of all external nodal parasitics, directly from a 3-D FinFET circuit layout representation. The circuits consisting of advanced P2-D devices are then back annotated with the nodal parasitics to enable fast and accurate SRAM dynamic margin mixed-mode simulations. Results demonstrate up to 200× speedup compared with traditional 3-D device simulations, and around five orders of magnitude wall clock time improvement on account of fast statistical methodologies, which are superior in comparison with traditional Monte Carlo analysis. This makes it feasible to supplant often inaccurate compact model-based simulations by true mixed-mode device simulations in statistical engines. The proposed physics-based methodology is also shown to corroborate well with hardware measurements. Rajiv V. Joshi, Keunwoo Kim, Rouwaida Kanj, Ajay N. Bhoj, Matthew M. Ziegler, Phil Oldiges, Pranita Kerber, Robert Wong, Terence Hook, Sudesh Saroop, Carl Radens, Chun-Chen Yeh |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | A New Field-assisted Access Scheme of STT-RAM with Self-reference CapabilityabstractSpin-transfer torque random access memory (STT-RAM) has demonstrated great potentials in embedded and stand-alone applications. However, process variations and thermal fluctuations greatly influence the operation reliability of STT-RAM and limit its scalability. In this work, we propose a new field-assisted access scheme to improve the read/write reliability and performance of STT-RAM. During read operations, an external magnetic field is applied to a magnetic tunneling junction (MTJ) device, generating a resistive sense signal without referring to other devices. Such a self-reference scheme offers a very promising alternative approach to overcome the severe cell-to-cell variations at highly scaled technology node. Furthermore, the external magnetic field can be also used to assist the MTJ switching during write operations without introducing extra hardware overhead. Simulation results show that compared to the existing self-reference scheme, our proposed design can improve the read sense margin by more than 200% and reduce the write error rate down to 2.29×10--9. Enes Eken, Yaojun Zhang, Wujie Wen, Rajiv V. Joshi, Hai Li 0001, Yiran Chen 0001 |
DAC | 4 |
| 2013 | Efficient Methodologies for 3-D TCAD Modeling of Emerging Devices and CircuitsabstractOver the past decade, 3-D process simulation, which is central to the 3-D Technology Computer-Aided Design (3-D TCAD) approach, has severely limited the scope and applicability of TCAD to circuits with a small number of field-effect transistors, owing to its prohibitively high computational costs for large layouts. Due to rapidly changing process recipes and shorter production cycles in the industry, design-time optimization and iterative layout-3-D TCAD exploration for yield-critical or yield-characterizing circuits, such as static random-access memories (SRAMs), ring oscillators, and others, is currently impossible in a practical time frame. In this paper, we architect a novel layout/process/device-independent TCAD methodology in the Sentaurus tool suite to overcome the process simulation barrier for accurate 3-D TCAD structure generation. We adopt an automated structure synthesis (SS) approach, thereby bypassing the need for repetitive 3-D process simulations for different layouts or different versions of the same layout. Results for 32-nm bulk process simulations versus SS and 32-nm silicon-on-insulator (SOI) hardware measurements versus corresponding synthesized structures indicate that the method is an excellent substitute to 3-D process simulation of large layouts, with extremely favorable time and memory scaling behavior. Finally, the robustness and scalability of the proposed abstractions are highlighted through the synthesis of 22-nm SOI 6T FinFET SRAMs and ring oscillator structures. Ajay N. Bhoj, Rajiv V. Joshi, Niraj K. Jha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2013 | 3-D-TCAD-Based Parasitic Capacitance Extraction for Emerging Multigate Devices and CircuitsabstractIn recent years, the multigate field-effect transistor (FET) has emerged as the most viable contender for technology scaling down to the sub-10-nm nodes. The nonplanar nature of multigate devices, along with rapidly shrinking front-end-of-line (FEOL) and back-end-of-line (BEOL) features, has compounded the problem of parasitics extraction in future technology nodes. In this paper, for the first time, we address the above problem through a holistic 3-D-technology CAD (3-D-TCAD) flow for the extraction of FEOL/(FEOL+BEOL) capacitances in generic multigate circuit layouts, using a transport analysis-based approach. We investigate device-level parasitic capacitances in 3-D-process-simulated bulk and silicon-on-insulator FinFETs, and uncover capacitance scaling trends for candidate single/multifin multigate FETs along the 22-nm/14-nm/10-nm technology nodes. Leveraging automated structure synthesis algorithms, we synthesize 3-D multigate 6T SRAM structures using the process-simulated devices, and examine the effects of fin pitch, gate pitch, and fin count on circuit-level parasitics. Thereafter, we show that traditional segregated FEOL/BEOL modeling approaches fail to provide accurate estimates, by back-annotating 3-D-TCAD-extracted capacitances into mixed-mode write simulations of a 6T FinFET SRAM bitcell. Finally, using FinFET NAND2 logic gate delay simulations, we establish the fact that capturing parasitics accurately is as important as modeling device transport accurately, and that performance/dynamic behavior in multigate circuits is highly sensitive to both factors. Ajay N. Bhoj, Rajiv V. Joshi, Niraj K. Jha |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | Yield estimation via multi-conesabstractWe propose a new yield estimation algorithm which estimates the acceptability region as the union of spherical cones. The algorithm works by dividing the input parameter space into approximately equi-probable cones, efficiently estimating the refined weight contributions for each cone, then combining the results to get the total yield. The algorithm is broadly similar to the worst-case-distances method, but is more generally applicable for cases with -for example- multiple failure regions. The algorithm is quite accurate, and offers several orders (>100x) of magnitude of speedup compared to traditional Monte Carlo. The paper includes example applications to difficult high-yield circuits like SRAM. Rouwaida Kanj, Rajiv V. Joshi, Zhuo Li 0001, Jerry Hayes, Sani R. Nassif |
DAC | 2 |
| 2012 | A thermal and process variation aware MTJ switching model and its applications in soft error analysisabstractSpin-transfer torque random access memory (STT-RAM) has recently gained increased attentions from circuit design and architecture societies. Although STT-RAM offers a good combination of small cell size, nanosecond access time and non-volatility for embedded memory applications, the reliability of STT-RAM is severely impacted by device variations and environmental disturbances. In this paper, we develop a compact switching model for magnetic tunneling junction (MTJ), which is the data storage device in STT-RAM cells. By leveraging the capability to simulate the impacts of thermal and process variations on MTJ switching, our model is able to analyze the diverse mechanisms of STT-RAM write operation failures. Besides the impacts of thermal and process variation, the soft error induced by radiation striking on the access transistor is another important threat to the MTJ reliability. It can also be analyzed by using our model. The incurred computation cost of our model is much less than the conventional macro-magnetic model, and hence, enabling its applications in comprehensive STT-RAM reliability analysis and design optimizations. Peiyuan Wang, Wei Zhang 0012, Rajiv V. Joshi, Rouwaida Kanj, Yiran Chen 0001 |
ICCAD | 3 |
| 2011 | Universal statistical cure for predicting memory lossabstractNovel nonvolatile memory (NVM) technologies are gaining significant attention from semiconductor industry in the competition of universal memory development. However, as nanoscale devices, these emerging NVMs suffer from the intrinsic technology challenges such as large process variations. The importance of effective statistical approaches for yield estimation and robust design arises in the commercialization of the emerging nonvolatile memory technologies. In this paper, we used Spin-Transfer Torque Random Access Memory (STT-RAM) as an example to explain some new memory failures mechanisms we have to face in the emerging memory technologies. Then, we applied a mixture importance sampling methodology to enable yield-driven design and extended its application beyond memories to peripheral circuits and logic blocks. The goal of these discussions is to propose a universal statistical methodology to predict memory loss and enable robust design practices. Rajiv V. Joshi, Rouwaida Kanj, Peiyuan Wang, Hai Li 0001 |
ICCAD | 1 |
| 2011 | Accelerated statistical simulation via on-demand Hermite spline interpolationsabstractWe propose an efficient Hermite spline-based SPICE simulation methodology for accurate statistical yield analysis. Unlike conventional methods, the spline-based transistor tables are built on-demand specific to the transient simulation requirements of the statistical experiments. Compared with traditional MOSFET table models, on-demand spline table models use ~500X less memory. This makes Hermite spline-based table models practical for use in simulations for process variation modeling. Furthermore, we propose an efficient gate voltage offset approach to model transistor threshold voltage variation. In this scenario, evaluations of the transistor model rely on a single reference table and require one set of spline function evaluations per VTsample point as opposed to two or more sets for VTinterpolation. This method is comprehensive and the results are in excellent agreement with traditional BSIM-based simulations. Around 4X improvement in speed, which includes the table generation cost, could be further improved by employing other fast-SPICE techniques or parallelism. To the best of our knowledge, this is the first time such a methodology has been coupled with importance sampling techniques to study the yield of memory designs. Rouwaida Kanj, Rajiv V. Joshi, Kanak Agarwal 0001, Ali Sadigh, David Winston, Sani R. Nassif |
ICCAD | 3 |
| 2011 | A Novel Column-Decoupled 8T Cell for Low-Power Differential and Domino-Based SRAM DesignabstractWe present a novel half-select disturb free transistor SRAM cell. The cell is 6T based and utilizes decoupling logic. It employs gated inverter SRAM cells to decouple the column select read disturb scenario in half-selected columns which is one of the impediments to lowering cell voltage. Furthermore, “false read” before write operation, common to conventional 6T designs due to bit-select and wordline timing mismatch, is eliminated using this design. Two design styles are studied to account for the emerging needs of technology scaling as designs migrate from 90 to 65 nm PD/SOI technology nodes. Namely we focus on a 90 nm PD/SOI sense Amp based and 65 nm PD/SOI domino read based designs. For the sense Amp based design, read disturbs to the fully-selected cell can be further minimized by relying on a read-assist array architecture which enables discharging the bit-line (BL) capacitance to GND during a read operation. This together with the elimination of half-select disturbs enhance the overall array low voltage operability and hence reduce power consumption by 20%-30%. The domino read based SRAM design also exploits the proposed cell to enhance cell stability while reducing the overall power consumption more than 30% by relying on a dynamic dual supply technique in combination of cell design and peripheral circuitry. Because half-selected columns/cells are inherently protected by the proposed scheme, the dynamic supply “High” voltage is only applied to read selected columns/cells, while dynamic supply “Low” is employed in all other situations, thereby reducing the overall design power. A short bitline loading of 16 cells/BL is adopted to achieve high-performance low-power operation and lower bitline capacitance to improve stability. A newly developed fast Monte Carlo based statistical method is used to analyze such a unique cell, and 65 nm design simulations are carried out at 5 GHz. The feasibility of the cell and sensitivity to sense Amp timing has been proved by fabricating a 32 kb array in a 90-nm PD/SOI technology. Hardware experiments and simulation results show improvements of cell Vddminover traditional 6T cells by more than 150 mV for 90 nm PD/SOI technology. Also experimental results based on fabricated 65 nm PD/SOI (1.6 kb/site × 80 sites) hardware also asserts half-select disturb elimination and hence the ability to enable significant power savings. The performance and speed are shown to be comparable with the conventional 6T design. Rajiv V. Joshi, Rouwaida Kanj, Vinod Ramadurai |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2010 | Statistical leakage modeling for accurate yield analysis: the CDF matching method and its alternativesabstractWe study the impact of statistical leakage modeling on the yield of memory designs. We critically evaluate different closed form models from a rare fail event perspective and propose CDF matching as a comprehensive and effective approach for accurate statistical leakage modeling. While Schwartz-Yeh method is found to match the body and left tail of the distribution, the Fenton-Wilkinson method aims more at matching the right tail of the distribution. The latter is more critical for purposes of yield estimation in the presence of leaky bitlines devices, as the right tail region is more crucial. However, for practical applications, it is shown that even Fenton-Wilkinson method leads to reduced accuracy compared to the CDF matching method. The error in estimating the probability of a false-read is shown to range from 10x-147x and is expected to increase with technology scaling. Rouwaida Kanj, Rajiv V. Joshi, Sani R. Nassif |
ISLPED | 2 |
| 2009 | An elegant hardware-corroborated statistical repair and test methodology for conquering aging effectsabstractWe propose a new and efficient statistical-simulation-based test methodology for optimally selecting repair elements at beginning-of-life (BOL) to improve the end-of-life (EOL) functionality of memory designs. This is achieved by identifying the best BOL test/repair corner that maximizes EOL yield, thereby exploiting redundancy to optimize EOL operability with minimal BOL yield loss. The statistical approach makes it possible to identify such corners with tremendous savings in terms of test time and hardware. To estimate yields and search for the best repair corner the approach relies on fast conditional importance sampling statistical simulations. The methodology is versatile and can handle complex aging effects with asymmetrical distributions. Results are demonstrated on state-of-the-art dual-supply memory designs subject to statistical negative bias temperature instability (NBTI) effects, and hardware results are shown to match predicted model trends. Rouwaida Kanj, Rajiv V. Joshi, Chad Adams, James D. Warnock, Sani R. Nassif |
ICCAD | 2 |
| 2008 | SRAM methodology for yield and power efficiency: per-element selectable supplies and memory reconfiguration schemesabstractWe present a novel power-aware yield enhancement design methodology and reconfiguration scheme for deep submicron SRAM designs. We show that with the continued trend of raising array supply to counter process variations, it is more effective to use a per-element selectable virtual power-supply scenario as opposed to single array supply with traditional redundancy schemes. The element can be a bank, a sub-array, or an independent row/column, and the element's virtual supply value is determined based on fail bitmaps. The technique can also be used in conjunction with traditional redundancy schemes to further improve the efficiency. The supply and redundancy assignments can be obtained by relying on memory reconfiguration algorithms. For this, we propose a greedy yet accurate algorithm that runs in O(nlogn) as opposed to average case O(n2) traditional algorithms. The methodology leads to significant power savings ranging from 20% to 50% for 65nm technology. We expect the savings to increase in future technologies as leakage powers dominate. To the best of our knowledge, this is the first time such a methodology is applied to SRAM designs. Rouwaida Kanj, Rajiv V. Joshi, Zhuo Li 0001, Jente B. Kuang, Hung C. Ngo, Nancy Y. Zhou, Weiping Shi, Sani R. Nassif |
ISLPED | 2 |
| 2007 | A floating-body dynamic supply boosting technique for low-voltage sram in nanoscale PD/SOI CMOS technologiesabstractThis paper presents a novel dynamic supply boosting technique for low voltage SRAMs at/beyond 65 nm PD/SOI technologies. For the first time the technique exploits the capacitive coupling effect in a floating-body PD/SOI device to dynamically boost the virtual array supply voltage during Read operation, thus improving the Read performance, Read/half-select stability, and Vmin. This enables significant reduction of the standby cell power and circuit active power in a single supply methodology. The performance and parametric yield improvements in the presence of variability are analyzed/validated using precise and fast Monte Carlo statistical circuit simulations with mixture importance sampling. Fabricated column-based 65nm PD/SOI SRAM circuits are confirmed with simulations and physical analysis and are shown to operate at 0.4 V. to 0.5V. Rajiv V. Joshi, Rouwaida Kanj, Keunwoo Kim, Richard Q. Williams, Ching-Te Chuang |
ISLPED | 1 |
| 2006 | A novel variation-aware low-power keeper architecture for wide fan-in dynamic gatesabstractSubstantial increase in leakage current and threshold voltage fluctuations are making design of robust wide fan-in dynamic gates a challenging task. Traditionally, a PMOS keeper transistor has been employed to compensate for leakage current of pull down (NMOS) network. However, to maintain acceptable noise margin level in sub-100 nm technologies, large PMOS is necessary, which results in substantial contention (during pull down) and severe loss of performance. In this paper, a novel keeper architecture is proposed which is capable of significantly reducing the contention and improving the performance and power consumption. Using circuit simulations, superior characteristics of the proposed keeper is demonstrated in comparison to those of the traditional as well as state-of-the-art keepers. It is shown that for an 8-input OR gate, in presence of 15% Vth fluctuations, the proposed architecture can lead to 20%, 15%, and more than 40% reduction in power consumption, mean delay, and standard deviation of delay, respectively, when compared to traditional keeper circuit. Hamed F. Dadgour, Rajiv V. Joshi, Kaustav Banerjee |
DAC | 2 |
| 2006 | Mixture importance sampling and its application to the analysis of SRAM designs in the presence of rare failure eventsabstractIn this paper, we propose a novel methodology for statistical SRAM design and analysis. It relies on an efficient form of importance sampling, mixture importance sampling. The method is comprehensive, computationally efficient and the results are in excellent agreement with those obtained via standard Monte Carlo techniques. All this comes at significant gains in speed and accuracy, with speedup of more than 100X compared to regular Monte Carlo. To the best of our knowledge, this is the first time such a methodology is applied to the analysis of SRAM designs. Rouwaida Kanj, Rajiv V. Joshi, Sani R. Nassif |
DAC | 2 |
| 2005 | A Low-Overhead Virtual Rail Technique for SRAM Leakage Power ReductionabstractWe propose a virtual supply rail control technique that reduces SRAM leakage. This method encompasses a cell-based image, serial tiling, pitch matching, small drive device overhead, and controlled power-on currents while incurring small circuit overhead. A virtual rail cell contains both the sleep transistor fingers and input/output drive transistors. The usual overhead associated with the drive circuit that controls the sleep transistors is significantly reduced due to reduced wire load and improved drive efficiency. This technique provides gradual power-on characteristics and good signal slews while effectively mitigating leakage current, maintaining read/write speed and achieving power-on latency compatible with high-performance designs. Jente B. Kuang, Hung C. Ngo, Kevin J. Nowka, Jethro C. Law, Rajiv V. Joshi |
ICCD | 5 |
| 2004 | Nanoscale CMOS circuit leakage power reduction by double-gate deviceabstractLeakage power for extremely scaled (Leff = 25 nm) double-gate devices is examined. Numerical two-dimensional simulation results for double-gate CMOS device/circuit power are presented from physics principle, identifying that double-gate technology is an ideal candidate for low-power applications. Unique double-gate device features resulting from gate-gate coupling are discussed and effectively exploited for optimal low-leakage device design. Design trade-offs for double-gate CMOS power and performance are suggested for low-power and high-performance applications. Total power consumptions of static and dynamic circuits and latches for double-gate device are analyzed considering state dependency, showing that leakage current is reduced by a factor of over 10X, compared with conventional bulk-Si counterpart. Keunwoo Kim, Koushik K. Das, Rajiv V. Joshi, Ching-Te Chuang |
ISLPED | 3 |
| 2003 | Design and CAD Challenges in sub-90nm CMOS Technologies
Kerry Bernstein, Ching-Te Chuang, Rajiv V. Joshi, Ruchir Puri |
ICCAD | 3 |
| 2003 | New optimal design strategies and analysis of ultra-low leakage circuits for nano-scale SOI technologyabstractThis paper proposes new SOI circuit strategies for simultaneous reduction of standby gate and sub-threshold leakages. Various enhanced MTCMOS design alternatives are analyzed. A new method for assigning the V/sub TH/ and sizes of header and footer transistors is proposed, and stacking of headers/footers is analyzed. The optimum stacking height and tapering/sizing ratio under various design constraints are determined. Our strategies reduce MTCMOS standby leakage further by as much as 20/spl times/ and reduce virtual supply noise by 15%. Koushik K. Das, Rajiv V. Joshi, Ching-Te Chuang, Peter W. Cook, Richard B. Brown |
ISLPED | 2 |
| 2003 | Strained-si devices and circuits for low-power applicationsabstractStatic and dynamic power for strained-Si device is analyzed and compared with conventional bulk-Si technology. Optimum device design points are suggested with controlling physical/structural device parameters. Strained-Si CMOS circuits are studied, showing substantially-reduced power consumptions due to unique advantageous features of strained-Si device. The trade-offs for power and performance in strained-Si devices/circuits are discussed. Further, analysis and low-power design points are applied and extended to strained Si on SOI substrate (SSOI) CMOS technology. Keunwoo Kim, Rajiv V. Joshi, Ching-Te Chuang |
ISLPED | 2 |
| 2003 | PD/SOI SRAM performance in presence of gate-to-body tunneling currentabstractThis paper presents a detailed study on the effects of gate-to-body tunneling current on partially depleted silicon-on-insulator (PD/SOI) CMOS SRAM. It is shown that the presence of gate-to-body tunneling current changes the strength of individual cell transistor in the quiescent (standby) state, thus affecting subsequent write/read operations. The degradation in the "write" performance is shown to be more significant than the degradation in the "read" performance, and the effect is more pronounced at lowered temperature. For the beneficial side, the presence of the gate-to-body tunneling current reduces the initial cycle parasitic bipolar disturb from unselected cells on the same bitline during write/read operation. Rajiv V. Joshi, Ching-Te Chuang, Samuel K. H. Fung, Fari Assaderaghi, Melanie Sherony, Ghavam V. Shahidi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2001 | SOI for asynchronous dynamic circuitsabstractArticle Share on SOI for asynchronous dynamic circuits Authors: R. V. Joshi IBM Research Division, T. J. Watson Research Center, Yorktown Heights, NY IBM Research Division, T. J. Watson Research Center, Yorktown Heights, NYView Profile , W. Huang IBM Research Division, T. J. Watson Research Center, Yorktown Heights, NY IBM Research Division, T. J. Watson Research Center, Yorktown Heights, NYView Profile , C. T. Chuang IBM Research Division, T. J. Watson Research Center, Yorktown Heights, NY IBM Research Division, T. J. Watson Research Center, Yorktown Heights, NYView Profile Authors Info & Claims GLSVLSI '01: Proceedings of the 11th Great Lakes symposium on VLSIMarch 2001 Pages 37–42https://doi.org/10.1145/368122.368734Online:01 March 2001Publication History 0citation158DownloadsMetricsTotal Citations0Total Downloads158Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Rajiv V. Joshi, Wei Hwang, Ching-Te Chuang |
ACM Great Lakes Symposium on VLSI | 1 |
| 2001 | Mixed multi-threshold differential cascode voltage switch (MT-DCVS) circuit styles and strategies for low power VLSI designabstractArticle Share on Mixed multi-threshold differential cascode voltage switch (MT-DCVS) circuit styles and strategies for low power VLSI design Authors: W. Chen IBM T. J. Watson Research Center, Yorktown Heights, NY IBM T. J. Watson Research Center, Yorktown Heights, NYView Profile , W. Hwang IBM T. J. Watson Research Center, Yorktown Heights, NY IBM T. J. Watson Research Center, Yorktown Heights, NYView Profile , P. Kudva IBM T. J. Watson Research Center, Yorktown Heights, NY IBM T. J. Watson Research Center, Yorktown Heights, NYView Profile , G. Gristede IBM T. J. Watson Research Center, Yorktown Heights, NY IBM T. J. Watson Research Center, Yorktown Heights, NYView Profile , S. Kosonocky IBM T. J. Watson Research Center, Yorktown Heights, NY IBM T. J. Watson Research Center, Yorktown Heights, NYView Profile , R. Joshi IBM T. J. Watson Research Center, Yorktown Heights, NY IBM T. J. Watson Research Center, Yorktown Heights, NYView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 263–266https://doi.org/10.1145/383082.383156Online:06 August 2001Publication History 6citation4,358DownloadsMetricsTotal Citations6Total Downloads4,358Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Wei Hwang, Prabhakar Kudva, George Gristede, Stephen V. Kosonocky, Rajiv V. Joshi |
ISLPED | 6 |
| 2000 | "Cool low power" 1GHz multi-port register file and dynamic latch in 1.8 V, 0.25 mum SOI and bulk technology (poster session)abstractThis paper describes power analysis at sub-zero temperatures for a high performance dynamic multiport register file (6 Read and 2 Write ports, 32 wordlines x 64 bitlines) fabricated in 0.25 μm Silicon on Insulator (SOI) and bulk technologies. Based on the hardware it is shown that the performance of both register file and latch improves by 2-3.5% per 10§ C reduction in temperature. The standby power for SOI reduces by 1.5% to 3 per 10§ C temperature drop down to -30§ C. The SOI chip is shown to have more significant performance improvement at low temperatures compared to bulk chip due to the floating body effect which partially offsets the increase in the threshold voltages (Vt). The low temperature performance gain is attributed to reduction in capacitance (around 7-8%) and rest is due to dynamic threshold voltages. At 30§ C the register file is capable of functioning close to 1.02 GHz for read and write operations in a single cycle. Rajiv V. Joshi, Wei Hwang, S. C. Wilson, Ching-Te Chuang |
ISLPED | 1 |
| 1998 | Designing a Testable System on a ChipabstractA "system on a chip" is described, which integrates 16 Mbits of DRAM, digital logic, SRAM, three PLLs, and a triple video digital-to-analog converter in a 0.5 micron CMOS DRAM process. Application specific integrated circuit (ASIC) techniques are employed, using multiple DRAM macros with built-in self test (BIST), full level-sensitive scan design (LSSD) logic, and externally accessible analog circuitry. Issues regarding functional debugging, DRAM macro isolation and low cost manufacturing test using only a logic tester are described. Stephen V. Kosonocky, Arthur A. Bright, Kevin W. Warren, Ruud A. Haring, Steve Klepner, Sameh W. Asaad, S. Basavaiah, Bob Havreluk, David F. Heidel, Michael Immediato, Keith A. Jenkins, Rajiv V. Joshi, Benjamin D. Parker, T. V. Rajeevakumar, Kevin Stawiasz |
VTS | 12 |
| 1997 | A Pulse-To-Static Conversion Latch with a Self-Timed Control CircuitabstractThe design and experimental demonstration of a low-power pulse-to-static conversion latch circuit is described. The circuit includes self-timed control and a 64-bit latch array, both designed utilizing self-resetting CMOS (SRCMOS) circuit techniques. The self-timed feature of the control requires only one system clock input. The evaluation, reset and write-enable controls are all generated within a control macro. The latch is level sensitive scan design (LSSD) compatible and complies with SRCMOS test modes. Use of these latches facilitates the synchronization, pipelined operation, power-management, and testing of advanced digital systems employing a mix of static and dynamic circuits to achieve high performance. An experimental 64-bit latch array and self-timed control macro, designed for 2.5 V-0.5 /spl mu/m CMOS technology, has been successfully fabricated and tested. The full circuit occupies an area of 1.704 mm/spl times/0.07 mm, and the size of latch bit cell is 21.6 /spl mu/m/spl times/70 /spl mu/m. Experimental results have shown the conversion latch to function properly, capturing 1.2 ns output pulses from an SRCMOS register file, and properly converting them to static levels. The measured delay from global clock to static output was 725 ps. Wei Hwang, Rajiv V. Joshi, Walter H. Henkels |
ICCD | 2 |
| 1997 | Development of a High Bandwidth Merged Logic/DRAM Multimedia ChipabstractThis paper describes the design methodology and the implementation of a merged logic/DRAM multimedia chip. The design is based on 0.25 micron DRAM-based CMOS technology with 4-layers of metal with device performance enhancement. Details of the architecture and system design of the multi-media was described in Katayama et. al. (1996). The present chip consists of 64 Mb of synchronous DRAM which is organized in two banks of eight 8 Mb SDRAM macros, a gate-array memory control and bus control unit, a custom-designed 8/spl times/32-bit parallel graphic processor, a 64-bit parallel ports for data transfer to/from the host processor bus, a 32-bit serial port for video display, and on-chip PLL. The multi-media co-processor chip provides high-density unified memory, high bus bandwidth (4.3 GB/s peak) and 1+GB/s BITBLT processing functions for an external host processor. The current design, floorplan and layout are structured in a way that, we believe, will provide a general framework for other merged logic/DRAM, ASIC+DRAM design for system scale integration. W. K. Luk, Yasunao Katayama, Wei Hwang, Matthew R. Wordeman, Toshiaki Kirihata, Akashi Satoh, Seiji Munetoh, Hing Wong, B. El-Kareh, P. Xiao, Rajiv V. Joshi |
ICCD | 11 |