EDBT 2026 Demo / reviewers in the wild / expert
Rouwaida Kanj
dblp:13/2746
· DBLP profile ↗
31ranked-venue papers
8as first author
10since 2021 · last 2024
0000-0002-3519-2917ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 8 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorComputer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Reconfigurable Precision SRAM-based Analog In-memory-compute Macro DesignabstractIn-memory computing (IMC) is a promising approach for accelerating multiply and accumulate (MAC) operations, which are the primary calculations used in artificial intelligence (AI). The demand for flexible architectures supporting different bit precisions in MAC computations becomes evident. This flexibility balances adapting to specific model requirements and optimizing design performance efficiency. As such, in this paper, we propose a reconfigurable IMC macro design, utilizing 8T static random-access memory (SRAM) bit-cells in 65nm technology, to efficiently perform MAC operations while supporting three bit precisions: 2, 3, and 4 bits for each of the input, weight, and output. The proposed 64×180 macro achieves a normalized peak throughput of 13.82 TOPS, a normalized peak energy efficiency of 291.66 TOPS/W, and a normalized peak area efficiency of 165.98 TOPS/mm2. Jinane Bazzi, Rachid Jamil, Dana El Hajj, Rouwaida Kanj, Mohamed E. Fouda, Ahmed M. Eltawil |
ISCAS | 4 |
| 2023 | High-Density FeFET-based CAM Cell Design Via Multi-Dimensional EncodingabstractContent addressable memory is one of the most frequently used technologies in Data-centric applications due to its exceptional search parallelism capability. SRAM cells were initially used to implement CAM designs. Recent innovations proposed using compact nonvolatile memories instead. FeFETs emerged as a multi-level NVM device with promising potential and 2T FeFET CAM designs were studied. In this paper, a new potential is discussed for increasing the density efficiency of FeFET CAM architectures by adapting higher-dimensional encoding using 3T and 4T CAM designs. We propose a scalable greedy search algorithm for maximizing encoding capabilities. We compare the density, latency, accuracy, and energy consumption of our designs to standard 2T architecture demonstrating a 4x and 8x decrease in fail probability with up to 16% and 26.5% increase in memory density (bits/unit-area) in the 3T and 4T designs respectively. Hadi Noureddine, Omar Bekdache, Mohamad Al Tawil, Rouwaida Kanj, Ali Chehab, Mohamed E. Fouda, Ahmed M. Eltawil |
ACM Great Lakes Symposium on VLSI | 4 |
| 2023 | Hardware Acceleration of DNA Pattern Matching with Binary MemristorsabstractDNA pattern matching is a key technique applied in many bioinformatics applications. Recently, this technique has become very popular and is widely used for genetic disease diagnosis, where finding the number of consecutive repeats of a specific DNA pattern indicates the type and intensity of the patient's disorder. However, the remarkable growth of DNA data exacerbates the latency and power consumption required to perform DNA pattern matching. In this work, we propose a hardware accelerator design to detect the presence of different diseases efficiently using DNA pattern matching. We propose a novel CAM cell using binary memristors for reliable and robust data encoding. The proposed architecture consists of two main building blocks the Content-addressable memory (CAM) and pattern detector circuits in addition to the needed peripheral circuits for CAM read, write and match operation. CMOS PTM 45nm technology was used to design and simulate the full architecture. The evaluation of the proposed design shows$\sim 2\times$improvement in energy-delay-area product compared to the state-of-art work in the literature, in addition to robustness against noise and process variations. Jinane Bazzi, Mohamed E. Fouda, Rouwaida Kanj, Ahmed M. Eltawil |
ISCAS | 3 |
| 2023 | Scalable Complementary FeFET CAM DesignabstractCAMs are frequently employed for data-centric applications. They offer excellent parallelism. Traditionally, they were implemented using the area-consuming SRAM. Recent advancements suggest using compact nonvolatile memories (NVMs) to create CAM cells to reduce area. The ferroelectric field effect transistor (FeFET) has therefore emerged as an NVM device showing great potential in these memory architectures. In this work, we propose a novel multi-bit CAM architecture that utilizes p-type FeFETs – a topic yet to be explored in the literature – and we compare the latency, accuracy, and energy consumption of our design to other FeFET-based architectures demonstrating a 3-30× reduction in fail probability. Omar Bekdache, Hadi Noureddine, Mohamad Al Tawil, Rouwaida Kanj, Mohamed E. Fouda, Ahmed M. Eltawil |
ISCAS | 4 |
| 2023 | Hardware Implementation and Evaluation of an Information Processing FactoryabstractThe Information Processing Factory (IPF) utilizes factory management principles to tackle the complexities of integrated embedded systems, ensuring continuous safe operation and optimization at runtime. This paper presents a hardware implementation of IPF that enables dynamic task migration across system resources, ensuring reliability in the face of internal or external failures. We demonstrate the effectiveness of IPF through the efficient migration of tasks in multiprocessor SoCs using a safety-critical pacemaker application as a case study. Despite the additional software and hardware requirements, implementing IPF in a pacemaker results in comparable reliability to dual modular redundancy (DMR) with faster service resumption and improved resource utilization. Walaa Amer, Mariam Rakka, Rachid Karami, Minjun Seo, Mazen A. R. Saghir, Rouwaida Kanj, Fadi J. Kurdahi |
VLSI-SoC | 6 |
| 2023 | A Best Balance Ratio Ordered Feature Selection Methodology for Robust and Fast Statistical Analysis of Memory DesignsabstractRecently, machine learning yield models for integrated circuit (IC) have gained widespread prominence in the EDA community, and are very promising in terms of emulating memory design functionality and thereby speeding up circuit simulation-based variance reduction methods. A main challenge that arises in this area is a class imbalance that occurs naturally due to the high targeted manufacturing yield. Thus, the imbalanced nature of the sampled memory datasets can compromise the model performance. In this work, we attain deep insights into the memory classification problem for modeling rare fail events in the context of importance sampling-based yield analysis. We propose a comprehensive and computationally efficient method that addresses the joint considerations of the best combination of relevant features and class balance ratios, which are key for classifier generalization capability. The methodology relies on synthetic minority over-sampling techniques to enforce the minority class while probing for the best data balance ratio in conjunction with an iterative$L_{1}$-SVM-based approach that qualifies as an approximation to the$L_{0}$-norm regularization for the best feature subset selection. We compare the proposed methodology against standalone$L_{1}$-SVM solutions, unbalanced$L_{0}$-norm approximation as well as an algorithmic data balancing method in the context of yield estimation methodology. The methodology is shown to result in high fidelity classifiers as demonstrated when analyzing the yield of a 14-nm FinFET SRAM cross-section with speedup of$179\times $for the importance sampling simulations compared to pure circuit simulation-based approaches and an average error of$0.19 \sigma $. Lama Shaer, Rouwaida Kanj, Rajiv V. Joshi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | 1T1R In-Memory Compute for Winner Takes All Application in Kohonen Neural NetworksabstractIn-memory computing is a promising candidate for overcoming the von Neumann memory wall and accelerating data processing in neural network applications. We propose a ITIR-based approach to carry out an efficient in-memory computation of the minimum logic function to be executed in the form of a series of 1-bit search operations similar to modern associative processors. We study the proposed design in the context of Kohonen Neural Networks that are also known as self-organizing maps. They are based on competitive learning where the minimum-search operation is a crucial part in the learning process. We study the trade off of different low-to-high resistance state windows associated with different endurance levels in the presence of process variations. We demonstrate ~0-4% reduction in the MNIST test dataset accuracy for the different scenarios. We report a maximum of 6fJ for the 1-bit search operation. Aya Mouallem, Hussein Fadlallah, Lina Bacha, Dana El Hajj, Rachid Jamil, Dana Bazazo, Rouwaida Kanj |
ISCAS | 7 |
| 2022 | Group LARS-Based Iterative Reweighted Least Squares Methodology for Efficient Statistical Modeling of Memory DesignsabstractRegularized logistic regression is a popular classification tool that can be employed to accurately model the binary nature of the memory cell fail mechanisms for purposes of the yield analysis of memory designs. The iterative reweighted least squares (IRLS) method has been employed along with the least angle regression (LARS) to efficiently solve the$L_{1}$regularized logistic regression problem. In this brief, we propose an efficient$L_{1}$regularized logistic regression methodology. At the core lies a Group LARS-based approach that benefits from Group LARS inherent ability to handle groups of variables and exploits the natural evolution of the solution to speed up the search for the critical features of the classifier. Thus, it tracks Newton’s step direction from one round of the solution to the next and employs weighted directions to efficiently solve for the underlying$L_{1}$constrained iterative least squares problem. We apply the methodology in the context of an importance sampling-based yield analysis framework targeting rare fail probability estimation. We study the yield of 14-nm FinFET SRAM designs with programmable and resonant boosting. Our results demonstrate up to$14\times $–$20\times $speedup for the Group LARS compared to the pure LARS-based approach, and we report 98.7% accuracy and 0.12$\sigma $average error compared to pure circuit-simulations approach for the resulting classifier. Lama Shaer, Rouwaida Kanj, Rajiv V. Joshi, Ali Chehab |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Importance Splitting Sample Point Reuse for Efficient Memory Yield EstimationabstractIn this paper, we propose and evaluate efficient sample point reuse methodologies for Importance Splitting- based yield estimation of memory designs to assess the impact of manufacturing variability induced design center shifts. The proposed methodologies enable yield estimation at almost no additional simulation cost to that incurred due to studying the yield at the nominal design center. In order to unbias the Importance Splitting conditional probabilities with respect to the manufacturing variability centers, we evaluate three techniques that rely on Importance Sampling, center projections and geometric ratios. The geometric-based approach is shown to be most accurate for center shifts up-to three sigma away from the nominal. The proposed reuse methodologies achieve up-to 4 orders of magnitude speedup for the analysis of a 132-D 16nm SRAM design compared to traditional Importance Splitting. Mariam Rakka, Rouwaida Kanj |
ISCAS | 2 |
| 2021 | Secure MIMO D2D communication based on a lightweight and robust PLS cipher scheme
Hassan N. Noura, Reem Melki, Rouwaida Kanj, Ali Chehab |
Wirel. Networks | 3 |
| 2020 | Hybrid Importance Splitting Importance Sampling Methodology for Fast Yield Analysis of Memory DesignsabstractRare fail event estimation methodologies suffer from inefficiency when dealing with high dimensional design space problems. Importance splitting overcomes this complexity by recursively computing the rare fail probability as a product of larger conditional probabilities in the 1-D performance metric space. Its efficiency, however, drops as the events become rarer. In this work, we propose a novel hybrid Importance Sampling Importance Splitting methodology for purposes of rare fail event estimation of high-dimensional memory designs. In this context, we propose and evaluate two methods for unbiasing the estimate, a geometric ratio-based and an Importance Sampling-based methodology. We demonstrate 3-5X reduction in runtime for both theoretical and 16nm SRAM design applications compared to traditional Importance Splitting approaches. Mariam Rakka, Rouwaida Kanj, Ragheb Raad |
ISCAS | 2 |
| 2019 | Data Imbalance Handling Approaches for Accurate Statistical Modeling and Yield Analysis of Memory DesignsabstractData imbalance can impact the fidelity of a classifier. We rely on advances in data imbalance handling techniques for machine learning applications to propose an enhanced fast statistical analysis methodology. Particularly, we employ data handling techniques in the context of a logistic regression based importance sampling methodology for accurate statistical modeling of rare fail events in memory designs. We demonstrate that for purposes of achieving conservative yield estimates, the synthetic minority oversampling technique outperforms other data handling methods and portrays the best model recall and precision rates. We report more than 70% reduction in the number of False Negatives compared to imbalanced data set based approaches. We also report on average a low 5% relative error rate in the yield estimate for the balanced data set-based modeling approaches compared to the pure circuit simulation based approach. This is compared to on average an 18% relative error rate obtained for the imbalanced data set-based approaches. These results were verified on state-of-the-art industrial FinFET SRAM designs. Lama Shaer, Rouwaida Kanj, Rajiv V. Joshi |
ISCAS | 2 |
| 2019 | Verification at RTL Using Separation of Design ConcernsabstractDesign-for-test, logic built-in self-test, memory technology mapping, and clocking concerns require team-months of verification time as they traditionally happen at gate-level. We present a novel concern-oriented methodology that enables automatic insertion of these concerns at the register-transfer-level where verification is easier. The methodology involves three main phases: 1) flipflop inference and instantiation algorithms that handle parametric register transfer level (RTL) modules; 2) transformations that take entry RTL and produce RTL modules where memory elements are separated from functionality; and 3) a concern weaving tool that automatically inserts memory related design concerns implemented in recipe files into the RTL modules. The transformation is sound as proven and validated by equivalence checking using formal verification. We implemented the methodology in a tool that is currently used in an industrial setting wherein it reduced design verification time by more than 40%. The methodology is also effective with open source embedded system frameworks. Maya H. Safieddine, Fadi A. Zaraket, Rouwaida Kanj, Ali El-Zein, Wolfgang Roesner |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Sparse Regression Driven Mixture Importance Sampling for Memory DesignabstractIn this paper, we present a sparse regression (SpaRe) model-based yield analysis methodology and apply it to memory designs with state-of-the-art write-assist circuitry. At the core of its engine is a mixture importance sampling technique which consists of a uniform sampling stage and an importance sampling stage. The proposed methodology allows for fast and accurate statistical analysis of rare fail events. In our approach, a SpaRe model is built using the uniform sampling stage data points obtained via circuit simulation (CktSim). Along with the model, an optimal threshold value is determined for proper pass/fail predict capability. The model and the threshold value are then used to predict the response in the importance sampling stage. This alleviates the need for CktSims in the latter stage and introduces significant speedup compared to fully CktSim-based approaches. The SpaRe model-based yield analysis is tested on a 14-nm FinFET SRAM design, and the results corroborate well with that of full CktSim-based yield analysis. The methodology is used to compare multiple state-of-the-art SRAM designs including selective boost and write-assist designs. The operating Vmin ranges and trends corroborate well with hardware measurements. Maria Malik, Rajiv V. Joshi, Rouwaida Kanj, Shupeng Sun, Houman Homayoun |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Bayesian Model Fusion: Large-Scale Performance Modeling of Analog and Mixed-Signal Circuits by Reusing Early-Stage DataabstractEfficient performance modeling of today's analog and mixed-signal circuits is an important yet challenging task, due to the high-dimensional variation space and expensive circuit simulation. In this paper, we propose a novel performance modeling algorithm that is referred to as Bayesian model fusion (BMF) to address this challenge. The key idea of BMF is to borrow the information collected from an early stage (e.g., schematic level) to facilitate efficient performance modeling at a late stage (e.g., post layout). Such a goal is achieved by statistically modeling the performance correlation between early and late stages through Bayesian inference. Furthermore, to make the proposed BMF method of practical utility, four implementation issues, including: 1) prior mapping; 2) missing prior knowledge; 3) fast solver; and 4) prior and hyper-parameter selection, are carefully considered in this paper. Two circuit examples designed in a commercial 32 nm CMOS silicon on insulator process demonstrate that the proposed BMF method achieves up to 9× runtime speed-up over the traditional modeling technique without surrendering any accuracy. Fa Wang, Paolo Cachecho, Wangyang Zhang, Shupeng Sun, Xin Li 0001, Rouwaida Kanj, Chenjie Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2016 | A Universal Hardware-Driven PVT and Layout-Aware Predictive Failure Analytics for SRAMabstractThe impact of device variability, temperature, and technology CAD-based layout parasitics on low-voltage static random access memory (SRAM) yield is explored using a novel variability-aware statistical methodology. Threshold voltage, Vt, mismatches for planar 22- and 14-nm FinFET SRAM transistors are characterized based on unique array-like structures for capturing process voltage and temperature (PVT) impact on variability. In general, the mismatches are shown to be a consistent and unique function of Vdd, doping, and temperature across the two technologies. Stronger Vt mismatch impact is observed as a function of Vddand doping in the 22-nm technology, with higher mismatch recorded at lower temperatures. In the 14-nm technology, doping is found to have the strongest impact on Vtmismatch, and the mismatch increases with Vdddespite the reduced drain induced barrier lowering effects. Similar to the 22-nm technology, the mismatch increases at lower temperatures. Front-end-of-the line capacitance effects are found to be more significant than back-end-of-the-line effects in 14-nm technologies, as opposed to planar technologies. Accurate parasitic capacitance modeling along with PVT-aware variability process variations for different 22-/14-nm cell arrangements are incorporated into a physics based statistical analysis methodology for accurate Vmin analysis. The yield analysis results are corroborated with hardware yield using 4-16-Mb inline SRAM macro monitors. The methodology is unique in the industry, gives insight into the technology-circuit interactions, and is able to effectively predict the SRAM yield bounds. Rajiv V. Joshi, Sudesh Saroop, Rouwaida Kanj, Carl Radens, Karthik Yogendra |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Corrections to "Super Fast Physics-Based Methodology for Accurate Memory Yield Prediction"abstractOn[1, p. 536], Section II-A, second paragraph, lines 5–8, [2] should be used instead of[1]and the correct statement is as follows. “Model-to-hardware corroboration shows an excellent matching between importance-sampling-based methods yield estimation [2] and the true hardware yield. We therefore adopt the methodology in [2] as the core statistical engine for our TfM methodology.” Rajiv V. Joshi, Rouwaida Kanj |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Super Fast Physics-Based Methodology for Accurate Memory Yield PredictionabstractWe propose an efficient physics-based mixed-mode statistical simulation methodology for nanoscale devices and circuits. Here, 3-D Technology Computer Aided Design models pose a barrier for efficient simulation of variability as they generally involve millions of nodes in their mesh representations. The proposed methodology, which has been implemented for FinFET/tri-gate static random access memory (SRAM) design, overcomes this barrier by leveraging advanced physics-based 2-D (P2-D) devices with optimized meshes that are derived from 3-D FinFET models with tuned device parasitics. This enables physics-based simulation as well as physics-based variability input parameters. To improve accuracy, an embedded automated flow enables extraction of all external nodal parasitics, directly from a 3-D FinFET circuit layout representation. The circuits consisting of advanced P2-D devices are then back annotated with the nodal parasitics to enable fast and accurate SRAM dynamic margin mixed-mode simulations. Results demonstrate up to 200× speedup compared with traditional 3-D device simulations, and around five orders of magnitude wall clock time improvement on account of fast statistical methodologies, which are superior in comparison with traditional Monte Carlo analysis. This makes it feasible to supplant often inaccurate compact model-based simulations by true mixed-mode device simulations in statistical engines. The proposed physics-based methodology is also shown to corroborate well with hardware measurements. Rajiv V. Joshi, Keunwoo Kim, Rouwaida Kanj, Ajay N. Bhoj, Matthew M. Ziegler, Phil Oldiges, Pranita Kerber, Robert Wong, Terence Hook, Sudesh Saroop, Carl Radens, Chun-Chen Yeh |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | Yield estimation via multi-conesabstractWe propose a new yield estimation algorithm which estimates the acceptability region as the union of spherical cones. The algorithm works by dividing the input parameter space into approximately equi-probable cones, efficiently estimating the refined weight contributions for each cone, then combining the results to get the total yield. The algorithm is broadly similar to the worst-case-distances method, but is more generally applicable for cases with -for example- multiple failure regions. The algorithm is quite accurate, and offers several orders (>100x) of magnitude of speedup compared to traditional Monte Carlo. The paper includes example applications to difficult high-yield circuits like SRAM. Rouwaida Kanj, Rajiv V. Joshi, Zhuo Li 0001, Jerry Hayes, Sani R. Nassif |
DAC | 1 |
| 2012 | A thermal and process variation aware MTJ switching model and its applications in soft error analysisabstractSpin-transfer torque random access memory (STT-RAM) has recently gained increased attentions from circuit design and architecture societies. Although STT-RAM offers a good combination of small cell size, nanosecond access time and non-volatility for embedded memory applications, the reliability of STT-RAM is severely impacted by device variations and environmental disturbances. In this paper, we develop a compact switching model for magnetic tunneling junction (MTJ), which is the data storage device in STT-RAM cells. By leveraging the capability to simulate the impacts of thermal and process variations on MTJ switching, our model is able to analyze the diverse mechanisms of STT-RAM write operation failures. Besides the impacts of thermal and process variation, the soft error induced by radiation striking on the access transistor is another important threat to the MTJ reliability. It can also be analyzed by using our model. The incurred computation cost of our model is much less than the conventional macro-magnetic model, and hence, enabling its applications in comprehensive STT-RAM reliability analysis and design optimizations. Peiyuan Wang, Wei Zhang 0012, Rajiv V. Joshi, Rouwaida Kanj, Yiran Chen 0001 |
ICCAD | 4 |
| 2011 | Universal statistical cure for predicting memory lossabstractNovel nonvolatile memory (NVM) technologies are gaining significant attention from semiconductor industry in the competition of universal memory development. However, as nanoscale devices, these emerging NVMs suffer from the intrinsic technology challenges such as large process variations. The importance of effective statistical approaches for yield estimation and robust design arises in the commercialization of the emerging nonvolatile memory technologies. In this paper, we used Spin-Transfer Torque Random Access Memory (STT-RAM) as an example to explain some new memory failures mechanisms we have to face in the emerging memory technologies. Then, we applied a mixture importance sampling methodology to enable yield-driven design and extended its application beyond memories to peripheral circuits and logic blocks. The goal of these discussions is to propose a universal statistical methodology to predict memory loss and enable robust design practices. Rajiv V. Joshi, Rouwaida Kanj, Peiyuan Wang, Hai Li 0001 |
ICCAD | 2 |
| 2011 | Accelerated statistical simulation via on-demand Hermite spline interpolationsabstractWe propose an efficient Hermite spline-based SPICE simulation methodology for accurate statistical yield analysis. Unlike conventional methods, the spline-based transistor tables are built on-demand specific to the transient simulation requirements of the statistical experiments. Compared with traditional MOSFET table models, on-demand spline table models use ~500X less memory. This makes Hermite spline-based table models practical for use in simulations for process variation modeling. Furthermore, we propose an efficient gate voltage offset approach to model transistor threshold voltage variation. In this scenario, evaluations of the transistor model rely on a single reference table and require one set of spline function evaluations per VTsample point as opposed to two or more sets for VTinterpolation. This method is comprehensive and the results are in excellent agreement with traditional BSIM-based simulations. Around 4X improvement in speed, which includes the table generation cost, could be further improved by employing other fast-SPICE techniques or parallelism. To the best of our knowledge, this is the first time such a methodology has been coupled with importance sampling techniques to study the yield of memory designs. Rouwaida Kanj, Rajiv V. Joshi, Kanak Agarwal 0001, Ali Sadigh, David Winston, Sani R. Nassif |
ICCAD | 1 |
| 2011 | A Novel Column-Decoupled 8T Cell for Low-Power Differential and Domino-Based SRAM DesignabstractWe present a novel half-select disturb free transistor SRAM cell. The cell is 6T based and utilizes decoupling logic. It employs gated inverter SRAM cells to decouple the column select read disturb scenario in half-selected columns which is one of the impediments to lowering cell voltage. Furthermore, “false read” before write operation, common to conventional 6T designs due to bit-select and wordline timing mismatch, is eliminated using this design. Two design styles are studied to account for the emerging needs of technology scaling as designs migrate from 90 to 65 nm PD/SOI technology nodes. Namely we focus on a 90 nm PD/SOI sense Amp based and 65 nm PD/SOI domino read based designs. For the sense Amp based design, read disturbs to the fully-selected cell can be further minimized by relying on a read-assist array architecture which enables discharging the bit-line (BL) capacitance to GND during a read operation. This together with the elimination of half-select disturbs enhance the overall array low voltage operability and hence reduce power consumption by 20%-30%. The domino read based SRAM design also exploits the proposed cell to enhance cell stability while reducing the overall power consumption more than 30% by relying on a dynamic dual supply technique in combination of cell design and peripheral circuitry. Because half-selected columns/cells are inherently protected by the proposed scheme, the dynamic supply “High” voltage is only applied to read selected columns/cells, while dynamic supply “Low” is employed in all other situations, thereby reducing the overall design power. A short bitline loading of 16 cells/BL is adopted to achieve high-performance low-power operation and lower bitline capacitance to improve stability. A newly developed fast Monte Carlo based statistical method is used to analyze such a unique cell, and 65 nm design simulations are carried out at 5 GHz. The feasibility of the cell and sensitivity to sense Amp timing has been proved by fabricating a 32 kb array in a 90-nm PD/SOI technology. Hardware experiments and simulation results show improvements of cell Vddminover traditional 6T cells by more than 150 mV for 90 nm PD/SOI technology. Also experimental results based on fabricated 65 nm PD/SOI (1.6 kb/site × 80 sites) hardware also asserts half-select disturb elimination and hence the ability to enable significant power savings. The performance and speed are shown to be comparable with the conventional 6T design. Rajiv V. Joshi, Rouwaida Kanj, Vinod Ramadurai |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Statistical leakage modeling for accurate yield analysis: the CDF matching method and its alternativesabstractWe study the impact of statistical leakage modeling on the yield of memory designs. We critically evaluate different closed form models from a rare fail event perspective and propose CDF matching as a comprehensive and effective approach for accurate statistical leakage modeling. While Schwartz-Yeh method is found to match the body and left tail of the distribution, the Fenton-Wilkinson method aims more at matching the right tail of the distribution. The latter is more critical for purposes of yield estimation in the presence of leaky bitlines devices, as the right tail region is more crucial. However, for practical applications, it is shown that even Fenton-Wilkinson method leads to reduced accuracy compared to the CDF matching method. The error in estimating the probability of a false-read is shown to range from 10x-147x and is expected to increase with technology scaling. Rouwaida Kanj, Rajiv V. Joshi, Sani R. Nassif |
ISLPED | 1 |
| 2009 | Yield estimation of SRAM circuits using "Virtual SRAM Fab"abstractStatic Random Access Memories (SRAMs) are key components of modern VLSI designs and a major bottleneck to technology scaling as they use the smallest size devices with high sensitivity to manufacturing details. Analysis performed at the "schematic" level can be deceiving as it ignores the interdependence between the implementation layout and the resulting electrical performance. We present a computational framework, referred to as "Virtual SRAM Fab", for analyzing and estimating pre-Si SRAM array manufacturing yield considering both lithographic and electrical variations. The framework is being demonstrated for SRAM design/optimization in 45nm nodes and currently being used for both 32nm and 22nm technology nodes. The application and merit of the framework are illustrated using two different SRAM cells in a 45nm PD/SOI technology, which have been designed for similar stability/performance, but exhibit different parametric yields due to layout/lithographic variations. We also demonstrate the application of Virtual SRAM Fab for prediction of layout-induced imbalance in an 8T cell, which is a popular candidate for SRAM implementation in 32-22nm technology nodes. Aditya Bansal, Rama N. Singh, Rouwaida Kanj, Saibal Mukhopadhyay, Jin-Fuw Lee, Emrah Acar, Amith Singhee, Keunwoo Kim, Ching-Te Chuang, Sani R. Nassif, Fook-Luen Heng, Koushik K. Das |
ICCAD | 3 |
| 2009 | An elegant hardware-corroborated statistical repair and test methodology for conquering aging effectsabstractWe propose a new and efficient statistical-simulation-based test methodology for optimally selecting repair elements at beginning-of-life (BOL) to improve the end-of-life (EOL) functionality of memory designs. This is achieved by identifying the best BOL test/repair corner that maximizes EOL yield, thereby exploiting redundancy to optimize EOL operability with minimal BOL yield loss. The statistical approach makes it possible to identify such corners with tremendous savings in terms of test time and hardware. To estimate yields and search for the best repair corner the approach relies on fast conditional importance sampling statistical simulations. The methodology is versatile and can handle complex aging effects with asymmetrical distributions. Results are demonstrated on state-of-the-art dual-supply memory designs subject to statistical negative bias temperature instability (NBTI) effects, and hardware results are shown to match predicted model trends. Rouwaida Kanj, Rajiv V. Joshi, Chad Adams, James D. Warnock, Sani R. Nassif |
ICCAD | 1 |
| 2008 | SRAM methodology for yield and power efficiency: per-element selectable supplies and memory reconfiguration schemesabstractWe present a novel power-aware yield enhancement design methodology and reconfiguration scheme for deep submicron SRAM designs. We show that with the continued trend of raising array supply to counter process variations, it is more effective to use a per-element selectable virtual power-supply scenario as opposed to single array supply with traditional redundancy schemes. The element can be a bank, a sub-array, or an independent row/column, and the element's virtual supply value is determined based on fail bitmaps. The technique can also be used in conjunction with traditional redundancy schemes to further improve the efficiency. The supply and redundancy assignments can be obtained by relying on memory reconfiguration algorithms. For this, we propose a greedy yet accurate algorithm that runs in O(nlogn) as opposed to average case O(n2) traditional algorithms. The methodology leads to significant power savings ranging from 20% to 50% for 65nm technology. We expect the savings to increase in future technologies as leakage powers dominate. To the best of our knowledge, this is the first time such a methodology is applied to SRAM designs. Rouwaida Kanj, Rajiv V. Joshi, Zhuo Li 0001, Jente B. Kuang, Hung C. Ngo, Nancy Y. Zhou, Weiping Shi, Sani R. Nassif |
ISLPED | 1 |
| 2007 | A floating-body dynamic supply boosting technique for low-voltage sram in nanoscale PD/SOI CMOS technologiesabstractThis paper presents a novel dynamic supply boosting technique for low voltage SRAMs at/beyond 65 nm PD/SOI technologies. For the first time the technique exploits the capacitive coupling effect in a floating-body PD/SOI device to dynamically boost the virtual array supply voltage during Read operation, thus improving the Read performance, Read/half-select stability, and Vmin. This enables significant reduction of the standby cell power and circuit active power in a single supply methodology. The performance and parametric yield improvements in the presence of variability are analyzed/validated using precise and fast Monte Carlo statistical circuit simulations with mixture importance sampling. Fabricated column-based 65nm PD/SOI SRAM circuits are confirmed with simulations and physical analysis and are shown to operate at 0.4 V. to 0.5V. Rajiv V. Joshi, Rouwaida Kanj, Keunwoo Kim, Richard Q. Williams, Ching-Te Chuang |
ISLPED | 2 |
| 2006 | Mixture importance sampling and its application to the analysis of SRAM designs in the presence of rare failure eventsabstractIn this paper, we propose a novel methodology for statistical SRAM design and analysis. It relies on an efficient form of importance sampling, mixture importance sampling. The method is comprehensive, computationally efficient and the results are in excellent agreement with those obtained via standard Monte Carlo techniques. All this comes at significant gains in speed and accuracy, with speedup of more than 100X compared to regular Monte Carlo. To the best of our knowledge, this is the first time such a methodology is applied to the analysis of SRAM designs. Rouwaida Kanj, Rajiv V. Joshi, Sani R. Nassif |
DAC | 1 |
| 2004 | Noise characterization of static CMOS gatesabstractWe present new macromodeling techniques for capturing the response of a CMOS logic gate to noise pulses at the input. Two approaches are presented. The first one is a robust mathematical model which enables the hierarchical generation of noise abstracts for circuits composed of the precharacterized cells. The second is a circuit equivalent model which generates accurate noise waveforms for arbitrarily shaped and timed multiple-input glitches, arbitrary loads, and external noise coupling. Rouwaida Kanj, Timothy Lehner, Bhavna Agrawal, Elyse Rosenbaum |
DAC | 1 |
| 2004 | Critical evaluation of SOI design guidelinesabstractDesign guidelines for static and domino silicon-on-insulator (SOI) CMOS circuits are evaluated. Restructuring the logic to eliminate gates with large fan-ins is almost as beneficial for SOI as for bulk-silicon. Most published design fixes for eliminating parasitic bipolar induced upset are shown to aggravate the charge sharing problem. A new and improved predischarge method for enhancing the noise tolerance of SOI domino circuits is thus proposed . The topic of multiple output domino logic in SOI technology is addressed for the first time. Multiple output domino logic is shown to be more prone to bipolar leakage induced upset than regular domino. Many of the design practices used to alleviate bipolar leakage in regular domino are no longer valid due to the multiple output domino logic's inherent design requirements. A novel SOI-specific multiple output domino logic, particularly suitable for adder designs, is introduced to minimize the bipolar leakage risk. Rouwaida Kanj, Elyse Rosenbaum |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |