Hamid Mahmoodi

dblp:m/HamidMahmoodi · also Hamid Mahmoodi-Meimand · DBLP profile ↗
← Back
54ranked-venue papers
2as first author
2since 2021 · last 2022
0000-0003-4237-3086ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 52 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 4Applied, interdisciplinary, general and emerging computing · 4Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
2 papers
Hardware security and side channels · 100%
Computer architecture, parallel and distributed computing, and storage systems
7 papers
Integrated circuit design · 41% Electronic design automation · 35% Reconfigurable computing and FPGAs · 15%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware security and side channels
intellectual property protection
0.612022
Silicon validation of LUT-based logic-locked IP cores · DAC 2022
Hardware security and side channels › hardware obfuscation › logic obfuscation
logic locking
0.612022
Silicon validation of LUT-based logic-locked IP cores · DAC 2022
Hardware security and side channels › hardware obfuscation › logic obfuscation › logic locking
SAT-resistant logic locking
0.612022
Silicon validation of LUT-based logic-locked IP cores · DAC 2022
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable logic
0.212016
Hybrid STT-CMOS designs for reverse-engineering prevention · DAC 2016
Electronic design automation
logic synthesis
0.232016
Hybrid STT-CMOS designs for reverse-engineering prevention · DAC 2016
Modeling and Circuit Synthesis for Independently Controlled Double Gate FinFET Devices · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
A novel synthesis approach for active leakage power reduction using dynamic supply gating · DAC 2005
Electronic design automation › hardware verification and test
silicon validation
0.212022
Silicon validation of LUT-based logic-locked IP cores · DAC 2022
Integrated circuit design › memory circuit design
SRAM design
0.122008
Reduction of Parametric Failures in Sub-100-nm SRAM Array Using Body Bias · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008
Modeling of failure probability and statistical design of SRAM array for yield enhancement in nanoscaled CMOS · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Electronic design automation › design for manufacturability › design for yield
yield enhancement
0.122008
Reduction of Parametric Failures in Sub-100-nm SRAM Array Using Body Bias · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008
Modeling of failure probability and statistical design of SRAM array for yield enhancement in nanoscaled CMOS · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Integrated circuit design
low-power circuit design
0.132007
Modeling and Circuit Synthesis for Independently Controlled Double Gate FinFET Devices · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Leakage current mechanisms and leakage reduction techniques in deep-submicrometer CMOS circuits · Proc. IEEE 2003
A novel synthesis approach for active leakage power reduction using dynamic supply gating · DAC 2005
Energy-efficient computing
leakage power reduction
0.122005
A novel synthesis approach for active leakage power reduction using dynamic supply gating · DAC 2005
Leakage current mechanisms and leakage reduction techniques in deep-submicrometer CMOS circuits · Proc. IEEE 2003
Integrated circuit design › variation-aware design
post-silicon tuning
0.112008
Reduction of Parametric Failures in Sub-100-nm SRAM Array Using Body Bias · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008
Electronic design automation
circuit synthesis
0.112007
Modeling and Circuit Synthesis for Independently Controlled Double Gate FinFET Devices · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Integrated circuit design › semiconductor device fabrication
CMOS technology
0.012003
Leakage current mechanisms and leakage reduction techniques in deep-submicrometer CMOS circuits · Proc. IEEE 2003
Memory systems › on-chip memory
SRAM array
0.022008
Reduction of Parametric Failures in Sub-100-nm SRAM Array Using Body Bias · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008
Modeling of failure probability and statistical design of SRAM array for yield enhancement in nanoscaled CMOS · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Integrated circuit design › semiconductor devices
short-channel effects
0.012003
Leakage current mechanisms and leakage reduction techniques in deep-submicrometer CMOS circuits · Proc. IEEE 2003

Methods — techniques the papers use, named apart from their topics

machine learning for SAT runtime validation · 1.1selection algorithms · 0.2selection algorithm · 0.2body bias · 0.1adaptive tuning · 0.1semi-analytical modeling · 0.1independent gate control · 0.1synthesis flow · 0.1statistical design · 0.1shannon expansion · 0.1process variation modeling · 0.1
YearPublicationVenuePosition
2022 Silicon validation of LUT-based logic-locked IP cores
abstract
Modern semiconductor manufacturing often leverages a fabless model in which design and fabrication are partitioned. This has led to a large body of work attempting to secure designs sent to an untrusted third party through obfuscation methods. On the other hand, efficient de-obfuscation attacks have been proposed, such as Boolean Satisfiability attacks (SAT attacks). However, there is a lack of frameworks to validate the security and functionality of obfuscated designs. Additionally, unconventional obfuscated design flows, which vary from one obfuscation to another, have been key impending factors in realizing logic locking as a mainstream approach for securing designs. In this work, we address these two issues for Lookup Table-based obfuscation. We study both Volatile and Non-volatile versions of LUT-based obfuscation and develop a framework to validate SAT runtime using machine learning. We can achieve unparallel SAT-resiliency using LUT-based obfuscation while incurring 7% area and less than 1% power overheads. Following this, we discuss and implement a validation flow for obfuscated designs. We then fabricate a chip consisting of several benchmark designs and a RISC-V CPU in TSMC 65nm for post functionality validation. We show that the design flow and SAT-runtime validation can easily integrate LUT-based obfuscation into existing CAD tools while adding minimal verification overhead. Finally, we justify SAT-resilient LUT-based obfuscation as a promising candidate for securing designs.
Gaurav Kolhe, Tyler David Sheaves, Kevin Immanuel Gubbi, Tejas Kadale, Setareh Rafatirad, Sai Manoj Pudukotai Dinakarrao, Avesta Sasan, Hamid Mahmoodi, Houman Homayoun
DAC8
2022 Breaking the Design and Security Trade-off of Look-up-table-based Obfuscation
abstract
Logic locking and Integrated Circuit (IC) camouflaging are the most prevalent protection schemes that can thwart most hardware security threats. However, the state-of-the-art attacks, including Boolean Satisfiability (SAT) and approximation-based attacks, question the efficacy of the existing defense schemes. Recent obfuscation schemes have employed reconfigurable logic to secure designs against various hardware security threats. However, they have focused on specific design elements such as SAT hardness. Despite meeting the focused criterion such as security, obfuscation incurs additional overheads, which are not evaluated in the present works. This work provides an extensive analysis of Look-up-table (LUT)–based obfuscation by exploring several factors such as LUT technology, size, number of LUTs, and replacement strategy as they have a substantial influence on Power-Performance-Area (PPA) and Security (PPA/S) of the design. We show that using large LUT makes LUT-based obfuscation resilient to hardware security threats. However, it also results in enormous design overheads beyond practical limits. To make the reconfigurable logic obfuscation efficient in terms of design overheads, this work proposes a novel LUT architecture where the security provided by the proposed primitive is superior to that of the traditional LUT-based obfuscation. Moreover, we leverage the security-driven design flow, which uses off-the-shelf industrial EDA tools to mitigate the design overheads further while being non-disruptive to the current industrial physical design flow. We empirically evaluate the security of the LUTs against state-of-the-art obfuscation techniques in terms of design overheads and SAT-attack resiliency. Our findings show that the proposed primitive significantly reduces both area and power by a factor of 8 \( \times \) and 2 \( \times \) , respectively, without compromising security.
Gaurav Kolhe, Tyler David Sheaves, Sai Manoj Pudukotai Dinakarrao, Hamid Mahmoodi, Setareh Rafatirad, Avesta Sasan, Houman Homayoun
ACM Trans. Design Autom. Electr. Syst.4
2019 On Custom LUT-based Obfuscation
abstract
Logic obfuscation yields hardware security against various threats, such as Intellectual Property (IP) piracy and reverse engineering. Evolving Boolean satisfiability (SAT) attacks have challenged the hardware security assurance rendered by various obfuscation methods. Recent works have centered on using re-configurable components such as Look-Up-Tables (LUTs) to enhance resiliency against attacks. Resiliency against SAT-attack is guaranteed when the size of LUT (number of inputs) is large. However, this incurs significant power, area and performance overheads. To address this challenge, this work proposes logic encryption based on customized LUT to make this practical. We propose two variants of the customized LUT based obfuscation: LUT+MUX based obfuscation, securing the design through routing obfuscation by MUX(multiplexer) and logic obfuscation of LUTs; and LUT+LUT based obfuscation, benefiting from LUT based obfuscation reinforced with additional logic/routing obfuscation. We evaluate the hardware security and overheads of the proposed two variants of customized LUT-based obfuscation on various benchmarks. Proposedcustomized LUT-based obfuscation breaks the security, power, and area trade-offs. The proposed solution is shown to be robust against SAT-attacks and power analysis-based side-channel attacks with8×reduced area and 3×reduced power on an average compared tostate-of-the-art LUT-based obfuscation.
Gaurav Kolhe, Sai Manoj Pudukotai Dinakarrao, Setareh Rafatirad, Hamid Mahmoodi, Avesta Sasan, Houman Homayoun
ACM Great Lakes Symposium on VLSI4
2019 Deep RNN-Oriented Paradigm Shift through BOCANet: Broken Obfuscated Circuit Attack
abstract
Logic encryption obfuscation has been used for thwarting counterfeiting, overproduction, and reverse engineering but vulnerable to attacks. However, it was recently shown that satisfiability - checking (SAT) can potentially compromise hardware obfuscation circuits. In this paper, we develop a novel attack called BOCANet that can be beneficial from deep learning architecture to compromise hardware obfuscation circuits's key. Our approach involves exploiting deep recurrent neural network (D-RNN) model, and developing attack model to compromise the obfuscated hardware at least an order-of magnitude more efficiently and under resource-constrained scenarios. In our experiments, the BOCANet approach achieves an average success rate of 100% for 32 bit key size, 93.4% for 64 bit key size, 92.2% and 91.7% for 128 and 256 bit key size, respectively.
Sara Tehranipoor, Nima Karimian, Mehran Mozaffari Kermani, Hamid Mahmoodi
ACM Great Lakes Symposium on VLSI4
2019 Security and Complexity Analysis of LUT-based Obfuscation: From Blueprint to Reality
abstract
Recent obfuscation schemes have leveraged reconfigurable logics to alleviate various hardware security threats. However, existing reconfigurable logic-based obfuscation schemes focus on specific design factors such as gate replacement strategy or an optimization metric such as SAT-hardness. Despite meeting the focused metrics such as security, the obfuscation also incurs overheads, which are not well analyzed in the existing works. In this work, we provide a comprehensive analysis on reconfigurable logic obfuscation schemes i.e., LUT-based obfuscation by investigating 3-key design factors such as (1) LUT size, (2) number of LUTs, and (3) replacement strategy as they have a considerable impact on design criteria, i.e., Power-Performance-Area (PPA) and Security (PPA/S). Our results show that among the studied parameters the size of LUT has the most prominent impact on improving the resiliency of LUT-based obfuscation against the SAT and removal attacks. However, using large size LUTs incur significant PPA overheads, making such solutions unfeasible and unpractical. To address this challenge, this work proposes a pragmatic solution based on a customized LUT, where the security provided by each LUT is superior to that of traditional LUT-based obfuscation. The proposed solution primarily benefits from LUT-based obfuscation reinforced with additional logic/routing obfuscation that is implemented using small 2-input LUTs. We evaluate the hardware security and overhead of the proposed customized LUT-based obfuscation on various benchmarks to prove that the customized LUT-based obfuscation breaks the PPA tradeoffs while exhibiting robustness against the SAT and removal attacks. The customized LUT-based obfuscation comes with 8× reduced area and 2× reduced power on an average compared to state-of-the-art LUT-based obfuscation without compromising security.
Gaurav Kolhe, Hadi Mardani Kamali, Miklesh Naicker, Tyler David Sheaves, Hamid Mahmoodi, Sai Manoj Pudukotai Dinakarrao, Houman Homayoun, Setareh Rafatirad, Avesta Sasan
ICCAD5
2018 Static Design of Spin Transfer Torques Magnetic Look Up Tables for ASIC Designs
abstract
In this paper, we propose a static approach to the design of Spin Transfer Torque Look Up Tables (STT-LUT) for integration in ASIC and investigate the sensing reliability in the proposed design in detail. The proposed design style utilizes STT-Latches that their sensing reliability is key in determining the overall reliability of the proposed static STT-LUT. The simulation results in a 10nm FinFET CMOS technology shows that the proposed static STT-LUT design exhibits up to 26% read delay reduction compared to the best dynamic STT-LUT design, and more than 2.5X reduction in sensing failure rate.
Aliyar Attaran, Tyler David Sheaves, Praveen Kumar Mugula, Hamid Mahmoodi
ACM Great Lakes Symposium on VLSI4
2018 Programmable Gates Using Hybrid CMOS-STT Design to Prevent IC Reverse Engineering
abstract
This article presents a rigorous step towards design-for-assurance by introducing a new class of logically reconfigurable design resilient to design reverse engineering. Based on the non-volatile spin transfer torque (STT) magnetic technology, we introduce a basic set of non-volatile reconfigurable Look-Up-Table (LUT) logic components (NV-STT-based LUTs). An STT-based LUT with a significantly different set of characteristics compared to CMOS provides new opportunities to enhance design security yet makes it challenging to remain highly competitive with custom CMOS or even SRAM-based LUT in terms of power, performance, and area. To address these challenges, we propose several algorithms to select and replace custom CMOS gates with reconfigurable STT-based LUTs during design implementation such that the functionality of STT-based components and therefore the entire design cannot be determined in any manageable time, rendering any design reverse engineering attack ineffective. Our study, conducted on a large number of standard circuit benchmarks, concludes significant resiliency of hybrid STT-CMOS circuits against various types of attacks. Furthermore, the selection algorithms on average have a small impact on the performance of the circuit. We also tested these techniques against satisfiability attacks developed recently and show that these techniques also render more advanced reverse-engineering techniques computationally infeasible.
Theodore Winograd, Gaurav Shenoy, Hassan Salmani, Hamid Mahmoodi, Setareh Rafatirad, Houman Homayoun
ACM Trans. Design Autom. Electr. Syst.4
2016 Hybrid STT-CMOS designs for reverse-engineering prevention
abstract
This paper presents a rigorous step towards design-for-assurance by introducing a new class of logically reconfigurable design resilient to design reverse engineering. Based on the non-volatile spin transfer torque (STT) magnetic technology, we introduce a basic set of non-volatile reconfigurable Look-Up-Table (LUT) logic components (NV-STT-based LUTs). STT-based LUT with significantly different set of characteristics compared to CMOS provides new opportunities to enhance design security yet makes it challenging to remain highly competitive with custom CMOS or even SRAM-based LUT in terms of power, performance and area. To address these challenges, we propose several algorithms to select and replace custom CMOS gates with reconfigurable STT-based LUTs during design implementation such that the functionality of STT-based components and therefore the entire design cannot be determined in any manageable time, rendering any design reverse engineering attack ineffective. Our study conducted on a large number of standard circuit benchmarks concludes significant resiliency of hybrid STT-CMOS circuits against various types of attacks. Furthermore, the selection algorithms on average have a small impact of less than 3%, 8%, and 3% on design parametric constraints including performance, power and area, respectively.
Theodore Winograd, Hassan Salmani, Hamid Mahmoodi, Kris Gaj, Houman Homayoun
DAC3
2016 Dynamic single and Dual Rail spin transfer torque look up tables with enhanced robustness under CMOS and MTJ process variations
abstract
In this paper, we investigate the limitation of existing STT-LUT designs and propose two new circuit styles of designing STT-LUTs that offer higher performance and robustness compared to the conventional STT-LUT design. The proposed styles include a Dynamic Single Rail (DSR) and a Dynamic Dual Rail (DDR) STT-LUT. The simulation results in a 16nm bulk CMOS technology shows that the proposed designs exhibits up to 3.3× read delay reduction, 2.4× active power reduction, and 441× sensing failure rate reduction compared to the best conventional STT-LUT design. The proposed DDR scheme offers the best overall performance even when considering the state of the art Separated Precharge Sensing Amplifier and Separated Decoding schemes.
Aliyar Attaran, Hassan Salmani, Houman Homayoun, Hamid Mahmoodi
ICCD4
2016 Comparative analysis of robustness of spin transfer torque based look up tables under process variations
abstract
Spin Transfer Torque (STT) switching realized using a Magnetic Tunnel Junction (MTJ) device has shown great potential for low power and non-volatile storage. A prime application of MTJs is in building non-volatile Look Up Tables (LUT) used in reconfigurable logic. Such LUTs use a hybrid integration of CMOS transistors and MTJ devices. This paper discusses the reliability of STT based LUTs under transistor and MTJ variations in nano-scale. The sources of process variations include both the CMOS device related variations and the MTJ variations. A key part of the STT based LUTs is the sense amplifier needed for reading out the MTJ state. We compare the voltage and current based sensing schemes in terms of the power, performance, and reliability metrics. Based on our simulation results in a 16nm CMOS, for the same total device area, the voltage mode sensing scheme offers 75% lower failure rates under threshold voltage (Vth) variations, 4.9X higher tolerance to MTJ resistance variations, 19% less delay, and 64% lower active power compared to the current sensing scheme.
Ragh Kuttappa, Houman Homayoun, Hassan Salmani, Hamid Mahmoodi
ISCAS4
2016 Power and energy reduction of racetrack-based caches by exploiting shared shift operations
abstract
In this paper, we propose a technique for reducing the power and energy consumptions of the racetrack-based caches. The technique uses a mapping method from the logical cache lines to the physical domains of the nanowires. The mapping method exploits the fact that, in a nanowire with several access heads, the shift operations are shared by the heads on that nanowire. Utilizing this inherent sharing, fewer nanowires are shifted to make a cache line available for both the read and write accesses. By using this method, the cache sets are shifted separately, which results in increase in the number of average shift operations. Thanks to the sharing of the shift operations among multiple heads, the total power and energy consumption of the shift operations are reduced. The effectiveness of the proposed technique is studied using the PARSEC benchmark package. The study shows that the power, energy consumption, energy-delay-product, and energy-delay-squared-product of L2 caches are reduced, on average, by 53%, 44%, 32%, 17%, respectively, compared to the state-of-the-art mapping methods.
Seyed Saber Nabavi Larimi, Mehdi Kamal, Ali Afzali-Kusha, Hamid Mahmoodi
VLSI-SoC4
2014 Exploiting STT-NV technology for reconfigurable, high performance, low power, and low temperature functional unit design
abstract
Unavailability of functional units and their unequal activity makes performance bottlenecks and thermal hot spot units in general-purpose processors. We propose to use reconfigurable functional units to overcome these challenges. A selected set of complex functional units that might be underutilized, such as a multiplier and divider, are realized in a time-multiplexed fashion using a shared programmable Look Up Table (LUT) based fabric. This allows for run-time reconfiguration and migration of their activity. LUT based implementation also allows under-utilized functional units to be dynamically reconfigured to the functional units that have a performance bottleneck and hence improving performance. The programmable LUTs are realized using Spin Transfer Torque (STT) Magnetic technology (also called STT-NV) due to its zero leakage and CMOS compatibility. The results show significant performance improvement of 16% on average across standard benchmarks, when replacing CMOS multiplier and divider with reconfigurable STT-NV LUT counterpart. In addition, reconfiguration reduces the maximum temperature of functional units by up to 27°C and almost eliminates the thermal variation across them. This comes with small power overhead and no area impact.
Adarsh Reddy Ashammagari, Hamid Mahmoodi, Houman Homayoun
DATE2
2014 Reconfigurable STT-NV LUT-based functional units to improve performance in general-purpose processors
abstract
Unavailability of functional units is a major performance bottleneck in general-purpose processors (GPP). In a GPP with limited number of functional units while a functional unit may be heavily utilized at times, creating a performance bottleneck, the other functional units might be under-utilized. We propose a novel idea for adapting functional units in GPP architecture in order to overcome this challenge. For this purpose, a selected set of complex functional units that might be under-utilized such as multiplier and divider, are realized using a programmable look up table-based fabric. This allows for run-time adaptation of functional units to improving performance. The programmable look up tables are realized using magnetic tunnel junction (MTJ) based memories that dissipate near zero leakage and are CMOS compatible. We have applied this idea to a dual issue architecture. The results show that compared to a design with all CMOS functional units a performance improvement of 18%, on average is achieved for standard benchmarks. This comes with 4.1% power increase in integer benchmarks and 2.3% power decrease in floating point benchmarks, compared to a CMOS design.
Adarsh Reddy Ashammagari, Hamid Mahmoodi, Tinoosh Mohsenin, Houman Homayoun
ACM Great Lakes Symposium on VLSI2
2013 Domino logic designs for high-performance and leakage-tolerant applications
Farshad Moradi, Tuan Vu Cao, Elena I. Vatajelu, Ali Peiravi, Hamid Mahmoodi, Dag T. Wisland
Integr.5
2012 Impact of technology scaling on performance of domino logic in nano-scale CMOS
Abhishek Guar, Hamid Mahmoodi
VLSI-SoC2
2012 Reliability enhancement of power gating transistor under time dependent dielectric breakdown
Hamid Mahmoodi
VLSI-SoC1
2011 Comparative analysis of copper and CNT interconnects for H-tree clock distribution
abstract
Clock distribution network is an important part of digital integrated circuits. The clock signal carried by the distribution network has to reach every end node at the same time to ensure synchronized switching. Due to mismatches among different nodes of the H-tree, the clock transitions among the final nodes of the distribution tree show some time difference, the maximum of which is called clock skew. In modern CMOS technologies, copper interconnect is popular for high level interconnects such as clock and power routing. Carbon Nanotube (CNT) exhibits less resistivity than copper making it a better material for interconnect. This paper compares the impact on clock skew of H-tree clock distribution network by replacing the traditional copper interconnects with carbon nanotube interconnects. By applying temperature mismatch, threshold voltage mismatch, and process mismatch, our findings show that using carbon nanotube interconnects reduces the clock skew significantly compared to traditional copper interconnects.
Vish Ganti, Hamid Mahmoodi
ICCD2
2011 Multi-level wordline driver for low power SRAMs in nano-scale CMOS technology
abstract
In this paper, a multi-level wordline driver scheme is presented to improve SRAM read and write stability while lowering power consumption during hold operation. The proposed circuit applies a shaped wordline voltage pulse during read mode and a boosted wordline pulse during write mode. During read, the applied shaped pulse is tuned at nominal voltage for short period of time, whereas for the remaining access time, the wordline voltage is reduced to a lower level. This pulse results in improved read noise margin without any degradation in access time which is explained by examining the dynamic and nonlinear behavior of the SRAM cell. Furthermore, during hold mode, the wordline voltage starts from a negative value and reaches zero voltage, resulting in a lower leakage current compared to conventional SRAM. Our simulations using TSMC 65nm process show that the proposed wordline driver results in 2X improvement in static read noise margin while the write margin is improved by 3X. In addition, the total leakage of the proposed SRAM is reduced by 10% while the total power is improved by 12% in the worst case scenario of a single SRAM cell. The total area penalty is 10% for a 128Kb standard SRAM array.
Farshad Moradi, Georgios Panagopoulos, Georgios Karakonstantis, Dag T. Wisland, Hamid Mahmoodi, Jens Kargaard Madsen, Kaushik Roy 0001
ICCD5
2011 Analysis of reliability of flip-flops under transistor aging effects in nano-scale CMOS technology
abstract
The effect of aging has become an important reliability concern in modern CMOS technology. NBTI and PBTI are known to bring about an increase in threshold voltage of the PMOS and NMOS respectively. This paper studies the effect of NBTI and PBTI on different flip-flop circuits with key parameters such as setup time, hold time, clock to output delay and data to output delay. The results in a predictive 32 nm technology show an increase of 0.43 to 1.23 pico-seconds in data-to-output delay depending on the Flip-Flop type. Moreover, we propose a method to use dual threshold voltage assignment to mitigate the effect of transistor aging on pulse triggered Flip-Flops. Dual Vthresults show lower delay as well as 30% reduction in delay aging using the proposed dual threshold voltage method.
Vikram G. Rao, Hamid Mahmoodi
ICCD2
2010 Reliability analysis of power gated SRAM under combined effects of NBTI and PBTI in nano-scale CMOS
abstract
Transistor aging effects (NBTI and PBTI) impact the reliability of SRAM in nano-scale CMOS technologies. In this research, the combined effect of NBTI and PBTI on power gated SRAM is analyzed. Optimal source biasing in the standby mode is presented as an effective method for guard-banding against the aging effects. The simulations results in a predictive 32nm technology shows maximum of 1.6% reduction in standby SNM over 5 year lifetime. For optimum operation, by decreasing the standby source bias voltage by only 0.012 volts, the SNM is safely margined for 5 year life time. This guard-banding comes at an insignificant power overhead of 0.6% for applied worse case scenarios. Given the insignificant power overhead with such guard-banding, it is concluded that adaptive tuning of the source biasing voltage is not required, given the not so negligible complexity and overhead associated with adaptive techniques.
Anuj Pushkarna, Hamid Mahmoodi
ACM Great Lakes Symposium on VLSI2
2010 Low-overhead Fmax calibration at multiple operating points using delay-sensitivity-based path selection
abstract
Maximum operating frequency ( F max ) of a system often needs to be determined at multiple operating points, defined by voltage and temperatures. Such calibration is important for the speed binning process, where the voltage-frequency (V- F max ) relation needs to be accurately determined to sort chips into different bins that can be used for different applications. Moreover, adaptive systems typically require F max calibration at multiple operating points in order to dynamically change operating condition such as supply voltage or body bias for power, temperature, or throughput management. For example, a Dynamic Voltage and Frequency Scaling (DVFS) system requires accurate delay calibration at multiple operating voltages in order to apply the correct operating frequency corresponding to a scaled supply. In this article, we propose a low-overhead design technique that allows efficient characterization of F max at different operating voltages and temperatures. The proposed method selects a set of representative timing paths in a circuit based on their temperature and voltage sensitivities and dynamically configures them into a ring oscillator to compute the critical path delay. Compared to existing F max calibration approaches, the proposed approach provides the following two main advantages: (1) it introduces a delay sensitivity metric to isolate few representative timing paths; (2) it considers actual timing paths instead of critical path replicas, thereby accounting for local within-die delay variations. The all-digital calibration method is robust under process variations and achieves high delay estimation accuracy (> 4% error) at the cost of negligible design overhead (1.7% in delay, 0.3% in power, and 3.5% in die-area).
Somnath Paul, Hamid Mahmoodi, Swarup Bhunia
ACM Trans. Design Autom. Electr. Syst.2
2009 Ultra Low Power Full Adder Topologies
abstract
In this paper several low power full adder topologies are presented. The main idea of these circuits is based on the sense energy recovery full adder (SERF) design and the GDI (gate diffusion input) technique. These subthreshold circuits are employed for ultra low power applications. While the proposed circuits have some area overhead that is negligible, they have at least 62% less power dissipation when compared with existing designs. In this paper, 65 nm standard models are used for simulations.
Farshad Moradi, Dag T. Wisland, Hamid Mahmoodi, Snorre Aunet, Tuan Vu Cao, Ali Peiravi
ISCAS3
2009 Ultra Low-Power Clocking Scheme Using Energy Recovery and Clock Gating
abstract
A significant fraction of the total power in highly synchronous systems is dissipated over clock networks. Hence, low-power clocking schemes are promising approaches for low-power design. We propose four novel energy recovery clocked flip-flops that enable energy recovery from the clock network, resulting in significant energy savings. The proposed flip-flops operate with a single-phase sinusoidal clock, which can be generated with high efficiency. In the TSMC 0.25-mum CMOS technology, we implemented 1024 proposed energy recovery clocked flip-flops through an H-tree clock network driven by a resonant clock-generator to generate a sinusoidal clock. Simulation results show a power reduction of 90% on the clock-tree and total power savings of up to 83% as compared to the same implementation using the conventional square-wave clocking scheme and flip-flops. Using a sinusoidal clock signal for energy recovery prevents application of existing clock gating solutions. In this paper, we also propose clock gating solutions for energy recovery clocking. Applying our clock gating to the energy recovery clocked flip-flops reduces their power by more than 1000times in the idle mode with negligible power and delay overhead in the active mode. Finally, a test chip containing two pipelined multipliers one designed with conventional square wave clocked flip-flops and the other one with the proposed energy recovery clocked flip-flops is fabricated and measured. Based on measurement results, the energy recovery clocking scheme and flip-flops show a power reduction of 71% on the clock-tree and 39% on flip-flops, resulting in an overall power savings of 25% for the multiplier chip.
Hamid Mahmoodi, Vishwanadh Tirumalashetty, Matthew Cooke, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2008 Arbitrary Two-Pattern Delay Testing Using a Low-Overhead Supply Gating Technique
Swarup Bhunia, Hamid Mahmoodi, Arijit Raychowdhury, Kaushik Roy 0001
J. Electron. Test.2
2008 Reduction of Parametric Failures in Sub-100-nm SRAM Array Using Body Bias
abstract
In this paper, we present a postsilicon-tuning technique to improve parametric yield of SRAM array using body bias (BB). First, we show that, although parametric failures in SRAM are due to local random intradie variations, the parametric failures increase at extreme interdie corners. Next, we show that proper BB can reduce different types of parametric failures. Finally, we show that adaptive application of BB to different dies, based on their interdie corners, reduces the total number of parametric failures in those dies. This helps to repair the faulty dies at different interdie corners, thereby improving SRAM yield. We show that postsilicon-tuning using BB can result in significant yield enhancement for SRAM (8%-25% in predictive 70-nm technology).
Saibal Mukhopadhyay, Hamid Mahmoodi, Kaushik Roy 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 Low-overhead design technique for calibration of maximum frequency at multiple operating points
abstract
Determination of maximum operating frequencies (Fmax) during manufacturing test at different operating voltages is required to: (a) to ensure that, for a Dynamic Voltage and Frequency Scaling (DVFS) system, the adaptation hardware actually applies the correct operating frequency corresponding to a scaled supply and (b) to sort chips in different voltage- frequency (V-Fmax)bins, so that chips at different bins can be used for different applications. Existing speed binning approach requires extensive delay testing at all operating points with all possible frequencies, which increases test cost and test time significantly. In this paper, we propose a low-overhead solution for characterizing Fmaxof a circuit at different operating voltages that can eliminate the complex and expensive Fmaxcalibration at multiple voltage points. The basic idea is to choose a small set of representative paths in a circuit based on their voltage sensitivity and dynamically configuring them into ring oscillator to compute the Fmax. The proposed calibration mechanism is all-digital, robust to process variations, reasonably accurate (average 2.8% error) and incorporates minimal hardware overhead (average 1.7% delay, 3.5% area and 0.28% power overhead).
Somnath Paul, Sivasubramaniam Krishnamurthy, Hamid Mahmoodi, Swarup Bhunia
ICCAD3
2007 Clock Gating and Negative Edge Triggering for Energy Recovery Clock
abstract
Energy recovery clocking has been demonstrated as an effective method for reducing the clock power. In this method the conventional square wave clock signal is replaced by a sinusoidal clock generated by a resonant circuit. Such a modification in clock signal prevents application of existing clock gating solutions. In this paper, we propose a clock gating solution for energy recovery clocking by gating the flip-flops. Applying our clock gating to the energy recovery clocked flip-flops reduces their power by 1000times in the idle mode with negligible power and delay overhead in the active mode. Applying the proposed clock gating technique to a system of 1000 flip-flops with idle mode probability and data switching activity of 50%, reduces the total power by 47%. We also propose a negative edge triggering solution for the energy recovery clocked flip-flops.
Vishwanadh Tirumalashetty, Hamid Mahmoodi
ISCAS2
2007 A low-power SRAM using bit-line charge-recycling technique
abstract
We propose a new low-power SRAM using bit-line Charge Recycling (CR-SRAM) for the write operation. In the proposed write scheme, differential voltage swing of a bit-line is obtained by recycled charge from its adjacent bit-line capacitance. In order to improve the data retention capability of un-selected cells during write, the power supply lines of memory cells in one column are connected to each other and separated from the power lines of other columns. A test-chip is fabricated in 0.13μm CMOS and measurement results show 88% reduction in total power compared to the conventional SRAM (CON-SRAM) at VDD=1.5V and f=100MHz.
Keejong Kim, Hamid Mahmoodi, Kaushik Roy 0001
ISLPED2
2007 Modeling and Circuit Synthesis for Independently Controlled Double Gate FinFET Devices
abstract
Independent control of front and back gate in double gate (DG) devices can be used to merge parallel transistors in noncritical paths. This reduces the effective switching capacitance and, hence, the dynamic power dissipation of a circuit. However, efficient design of large-scale circuits with DG devices is not well explored due to lack of proper modeling and large-scale design simulation tools. In this paper, we propose several low-power circuit options using independent gate FinFETs. We developed semianalytical models for different FinFET logic gates to predict their performance. An efficient circuit synthesis methodology comprised of proposed low-power logic options in FinFET design library has been developed. Results show about 8.5% area savings and 18% power savings over conventional FinFET technology for ISCAS85 benchmark circuits in 45-nm technology with no performance penalty.
Animesh Datta, Ashish Goel, R. T. Cakici, Hamid Mahmoodi, Dheepa Lekshmanan, Kaushik Roy 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2006 Low-overhead design of soft-error-tolerant scan flip-flops with enhanced-scan capability
abstract
With technology scaling, soft error resilience is becoming a major concern in circuit design. This paper presents a class of low-overhead flip-flops suitable for soft error detection and correction. The proposed design reuses logic elements typically available in a standard-cell implementation of a flip-flop to reduce hardware overhead. We demonstrate that the proposed flip-flops are also suitable for enhanced scan based delay fault testing, which allows arbitrary two-pattern test application for the best combinational path testability. The proposed flip-flops show an average power reduction of 16% and area improvement of 17% compared to the best alternative techniques with no additional delay overhead
Ashish Goel, Swarup Bhunia, Hamid Mahmoodi, Kaushik Roy 0001
ASP-DAC3
2006 Low power synthesis of dynamic logic circuits using fine-grained clock gating
abstract
Clock power consumes a significant fraction of total power dissipation in high speed precharge/evaluate logic styles. In this paper, we present a novel low-cost design methodology for reducing clock power in the active mode for dynamic circuits with fine-grained clock gating. The proposed technique also improves switching power by preventing redundant computations. A logic synthesis approach for domino/skewed logic styles based on Shannon expansion is proposed, that dynamically identifies idle parts of logic and applies clock gating to them to reduce power in the active mode of operation. Results on a set of MCNC benchmark circuits in predictive 70nm process exhibit improvements of 15% to 64% in total power with minimal overhead in terms of delay and area compared to conventionally synthesized domino/skewed logic
Nilanjan Banerjee, Kaushik Roy 0001, Hamid Mahmoodi, Swarup Bhunia
DATE3
2006 Novel Low-Overhead Operand Isolation Techniques for Low-Power Datapath Synthesis
abstract
Power consumption in datapath modules due to redundant switching is an important design concern for high-performance applications. Operand isolation schemes that reduce this redundant switching incur considerable overhead in terms of delay, power, and area. This paper presents novel operand isolation techniques based on supply gating that reduce overheads associated with isolating circuitry. The proposed schemes also target leakage minimization and additional operand isolation at the internal logic of datapath to further reduce power consumption. We integrate the proposed techniques and power/delay models to develop a synthesis flow for low-power datapath synthesis. Simulation results show that the proposed operand isolation techniques achieve at least 40% reduction in power consumption compared to original circuit with minimal area overhead (5%) and delay penalty (0.15%)
Nilanjan Banerjee, Arijit Raychowdhury, Kaushik Roy 0001, Swarup Bhunia, Hamid Mahmoodi
IEEE Trans. Very Large Scale Integr. Syst.5
2006 A novel high-performance and robust sense amplifier using independent gate control in sub-50-nm double-gate MOSFET
abstract
Double-gate (DG) transistor has emerged as one of the most promising devices for nano-scale circuit design. In this paper, we propose a high-performance and robust sense-amplifier design using independent gate control in symmetric and asymmetric DG devices for sub-50-nm technologies. The proposed sense amplifier has better performance (30%-35% less sensing delay) and robustness (60%-80% less minimum input bit-differential for correct operation considering 10% worst case silicon thickness mismatch) compared to the connected gate design. Hence, the proposed design successfully demonstrates the benefit of using independent gate control in DG devices for efficient circuit design in sub-50-nm regime.
Saibal Mukhopadhyay, Hamid Mahmoodi, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2005 Leakage Current Based Stabilization Scheme for Robust Sense-Amplifier Design for Yield Enhancement in Nano-scale SRAM
abstract
In this paper, we develop a method to analyze the probability of access failure in SRAM array (due to random Vt variation in transistors) by jointly considering variations in cell and senseamplifiers. Our analysis shows that, improving robustness of senseamplifier is extremely important for reducing memory access failure probability and improving yield. We present a process variation tolerant sense amplifier suitable for SRAM array designed in sub- 100nm CMOS technologies. The proposed technique reduces the failure probability of sense amplifiers by more than 80% with negligible penalty in the sensing delay.
Saibal Mukhopadhyay, Arijit Raychowdhury, Hamid Mahmoodi, Kaushik Roy 0001
Asian Test Symposium3
2005 A novel synthesis approach for active leakage power reduction using dynamic supply gating
abstract
Due to exponential increase in subthreshold leakage with technology scaling and temperature increase, leakage power is becoming a major fraction of total power in the active mode. We present a novel low-cost design methodology with associated synthesis flow for reducing both switching and active leakage power using dynamic supply gating. A logic synthesis approach based on Shannon expansion is proposed that dynamically applies supply gating to idle parts of general logic circuits even when they are performing useful computation. Experimental results on a set of MCNC benchmark circuits in a predictive 70nm process exhibits improvements of 15% to 88% in total active power compared to the results obtained by a conventional optimization flow.
Swarup Bhunia, Nilanjan Banerjee, Qikai Chen, Hamid Mahmoodi, Kaushik Roy 0001
DAC4
2005 A Novel Low-overhead Delay Testing Technique for Arbitrary Two-Pattern Test Application
abstract
With increasing process fluctuations in nano-scale technology, testing for delay faults is becoming essential in manufacturing test to complement stuck-at-fault testing. Design-for-testability techniques, such as enhanced scan are typically associated with considerable overhead in die-area, circuit performance, and power during normal mode of operation. This paper presents a novel test technique, which can be used as an alternative to the enhanced scan based delay fault testing method, with significantly less design overhead. Instead of using an extra latch as in the enhanced scan method, we propose using supply gating at the first level of logic gates to hold the state of a combinational circuit. Experimental results on a set of ISCAS89 benchmarks show an average reduction of 33% in area overhead with an average improvement of 71% in delay overhead and 90% in power overhead during normal mode of operation, compared to the enhanced scan implementation.
Swarup Bhunia, Hamid Mahmoodi, Arijit Raychowdhury, Kaushik Roy 0001
DATE2
2005 Energy recovery clocked dynamic logic
abstract
Energy recovery clocking results in significant energy savings in clock distribution networks as compared to conventional square-wave clocking. However, since energy recovery clocks are sinusoidal in nature, standard dynamic logic styles do not work efficiently when used with energy recovery clocks. We propose novel dynamic logic styles that operate more efficiently with sinusoidal clocks, enabling energy recovery from their clock networks, and resulting in significant energy savings. Based on the simulation results using TSMC 0.25 μm CMOS process technology, at iso-performance, the proposed dynamic logic styles exhibit up to 53% power reduction.
Matthew Cooke, Hamid Mahmoodi, Qikai Chen, Kaushik Roy 0001
ACM Great Lakes Symposium on VLSI2
2005 A high speed and leakage-tolerant domino logic for high fan-in gates
abstract
Robustness of high fan-in domino circuits is degraded by technology scaling due to exponential increase in leakage. In this paper, we propose a new domino circuit for high fan-in and high-speed applications in ultra deep submicron technologies. The proposed circuit employs a footer transistor that is initially OFF in the evaluation phase to reduce leakage and then turned ON to complete the evaluation. According to simulations in a predictive 70nm process, the proposed circuit increases noise immunity by more than 26X for wide OR gates and shows performance improvement of up to 20% compared to conventional domino logic circuits. The proposed circuit reduces the contention between keeper transistor and NMOS evaluation transistors at the beginning of evaluation phase. This results in less power dissipation for the proposed technique.
Farshad Moradi, Hamid Mahmoodi, Ali Peiravi
ACM Great Lakes Symposium on VLSI2
2005 Double-gate SOI devices for low-power and high-performance applications
abstract
Double-gate (DG) transistors have emerged as promising devices for nano-scale circuits due to their better scalability compared to bulk CMOS. Among the various types of DG devices, quasi-planar SOI FinFETs are easier to manufacture compared to planar double-gate devices. DG devices with independent gates (separate contacts to back and front gates) have recently been developed. DG devices with symmetric and asymmetric gates have also been demonstrated. Such device options have direct implications at the circuit level. Independent control of front and back gate in DG devices can be effectively used to improve performance and reduce power in sub-50nm circuits. Independent gate control can be used to merge parallel transistors in noncritical paths. This results in reduction in the effective switching capacitance and hence power dissipation. We show a variety of circuits in logic and memory that can benefit from independent gate operation of DG devices. As examples, we show the benefit of independent gate operation in circuits such as dynamic logic circuits, Schmitt triggers, sense amplifiers, and SRAM cells. In addition to independent gate option, we also investigate the usefulness of asymmetric devices and the impact of width quantization and process variations on circuit design.
Kaushik Roy 0001, Hamid Mahmoodi, Saibal Mukhopadhyay, Hari Ananthan, Aditya Bansal, Tamer Cakici
ICCAD2
2005 Novel Low-Overhead Operand Isolation Techniques for Low-Power Datapath Synthesis
abstract
Power consumption in datapath modules due to redundant switching is an important design concern for high-performance applications. Operand isolation schemes are adopted to reduce redundant switching in datapaths. However, they incur considerable overhead in terms of delay, power, and area. This paper presents novel operand isolation techniques based on supply gating that reduce the overheads associated with isolating circuitry. The proposed schemes also target leakage minimization and application of operand isolation at the internal logic of datapath to further reduce power consumption. We integrate the proposed techniques and power/delay models to develop a complete flow for low-power datapath synthesis. Simulation results show that the proposed operand isolation techniques can achieve at least 40% reduction in power consumption compared to the original circuit with minimal area overhead (5%) and small delay penalty (0.15%).
Nilanjan Banerjee, Arijit Raychowdhury, Swarup Bhunia, Hamid Mahmoodi, Kaushik Roy 0001
ICCD4
2005 Process Variation Tolerant Online Current Monitor for Robust Systems
abstract
Large inter-die and intra-die process variations result in significant uncertainty in delay of circuits. Large delay variations may lead to parametric/functional failures. In this paper we propose a leakage-variation-tolerant online current monitor, namely leakage canceling current sensor, to detect completion of operations in logic blocks. The current monitor is applied to self timed logic to design process variation tolerant circuits. It is observed that, for self-timed circuits, the probability of functional failures can be reduced by 50% with no performance degradation and with same power consumption.
Qikai Chen, Saibal Mukhopadhyay, Hamid Mahmoodi, Kaushik Roy 0001
IOLTS3
2005 A leakage control system for thermal stability during burn-in test
abstract
Increase in leakage current with technology scaling has been a major problem for IC technology. This problem becomes more crucial during burn-in test where stressed voltage and temperature are applied. Due to presence of a positive feedback between major components of leakage and temperature in CMOS circuits, excessive leakage may lead to thermal runaway and yield loss during burn-in test. This paper describes a novel integrated leakage control system to ensure thermal stability during burn-in test for a wide range of ambient temperatures and process variations
Mesut Meterelliyoz, Hamid Mahmoodi, Kaushik Roy 0001
ITC2
2005 Reliable and self-repairing SRAM in nano-scale technologies using leakage and delay monitoring
abstract
The inter-die and intra-die variations in process parameters result in large number of failures in an SRAM array degrading the design yield. In this paper, we propose an adaptive repairing technique for SRAM based on leakage and delay monitoring. Leakage and delay monitoring is used to effectively separate dies with different inter-die Vts from each other. Using the leakage (or delay) monitoring and adaptive body bias, we propose a reliable and self-repairing SRAM which has reduced number of parametric failures under high inter-die and intra-die Vt variations. The proposed self-repairing SRAM improves the design yield by 5%-40% in predictive 70nm technology from BPTM.
Saibal Mukhopadhyay, Kunhyuk Kang, Hamid Mahmoodi, Kaushik Roy 0001
ITC3
2005 Modeling and Testing of SRAM for New Failure Mechanisms Due to Process Variations in Nanoscale CMOS
abstract
In this paper, we have made a complete analysis of the emerging SRAM failure mechanisms due to process variations and mapped them to fault models. We have proposed two efficient test solutions for the process variation related failures in SRAM: (a) modification of March sequence, and (b) a low-overhead DFT circuit to complement the March test for an overall test time reduction of 29%, compared to the existing test technique with similar fault coverage.
Qikai Chen, Hamid Mahmoodi, Swarup Bhunia, Kaushik Roy 0001
VTS2
2005 Modeling of failure probability and statistical design of SRAM array for yield enhancement in nanoscaled CMOS
abstract
In this paper, we have analyzed and modeled failure probabilities (access-time failure, read/write failure, and hold failure) of synchronous random-access memory (SRAM) cells due to process-parameter variations. A method to predict the yield of a memory chip based on the cell-failure probability is proposed. A methodology to statistically design the SRAM cell and the memory organization is proposed using the failure-probability and the yield-prediction models. The developed design strategy statistically sizes different transistors of the SRAM cell and optimizes the number of redundant columns to be used in the SRAM array, to minimize the failure probability of a memory chip under area and leakage constraints. The developed method can be used in an early stage of a design cycle to enhance memory yield in nanometer regime.
Saibal Mukhopadhyay, Hamid Mahmoodi, Kaushik Roy 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2005 A process-tolerant cache architecture for improved yield in nanoscale technologies
abstract
Process parameter variations are expected to be significantly high in a sub-50-nm technology regime, which can severely affect the yield, unless very conservative design techniques are employed. The parameter variations are random in nature and are expected to be more pronounced in minimum geometry transistors commonly used in memories such as SRAM. Consequently, a large number of cells in a memory are expected to be faulty due to variations in different process parameters. We analyze the impact of process variation on the different failure mechanisms in SRAM cells. We also propose a process-tolerant cache architecture suitable for high-performance memory. This technique dynamically detects and replaces faulty cells by dynamically resizing the cache. It surpasses all the contemporary fault tolerant schemes such as row/column redundancy and error-correcting code (ECC) in handling failures due to process variation. Experimental results on a 64-K direct map L1 cache show that the proposed technique can achieve 94% yield compared to its original 33% yield (standard cache) in a 45-nm predictive technology under /spl sigma//sub Vt-inter/=/spl sigma//sub Vt-intra/=30 mV.
Amit Agarwal 0001, Bipul Chandra Paul, Hamid Mahmoodi, Animesh Datta, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2005 Low-power scan design using first-level supply gating
abstract
Reduction in test power is important to improve battery lifetime in portable electronic devices employing periodic self-test, to increase reliability of testing, and to reduce test cost. In scan-based testing, a significant fraction of total test power is dissipated in the combinational block. In this paper, we present a novel circuit technique to virtually eliminate test power dissipation in combinational logic by masking signal transitions at the logic inputs during scan shifting. We implement the masking effect by inserting an extra supply gating transistor in the supply to ground path for the first-level gates at the outputs of the scan flip-flops. The supply gating transistor is turned off in the scan-in mode, essentially gating the supply. Adding an extra transistor in only one logic level renders significant advantages with respect to area, delay, and power overhead compared to existing methods, which use gating logic at the output of scan flip-flops. Moreover, the proposed gating technique allows a reduction in leakage power by input vector control during scan shifting. Simulation results on ISCAS89 benchmarks show an average improvement of 62% in area overhead, 101% in power overhead (in normal mode), and 94% in delay overhead, compared to the lowest cost existing method.
Swarup Bhunia, Hamid Mahmoodi, Debjyoti Ghosh, Saibal Mukhopadhyay, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2005 Efficient testing of SRAM with optimized march sequences and a novel DFT technique for emerging failures due to process variations
abstract
With increasing inter-die and intra-die parameter variations in sub-100-nm process technologies, new failure mechanisms are emerging in CMOS circuits. These failures lead to reduction in reliability of circuits, especially the area-constrained SRAM cells. In this paper, we have analyzed the emerging failure mechanisms in SRAM caches due to transistor V/sub t/ variations, which results from process variations. Also we have proposed solutions to detect those failures efficiently. In particular, in this work, SRAM failure mechanisms under transistor V/sub t/ variations are mapped to logic fault models. March test sequences have been optimized to address the emerging failure mechanisms with minimal overhead on test time. Moreover, we have proposed a design for test circuit to complement the March test sequence for at-speed testing of SRAMs. The proposed technique, referred as double sensing, can be used to test the stability of SRAM cells during read operations. Using the proposed March test sequence along with the double sensing technique, a test time reduction of 29% is achieved, compared to the existing test techniques with the same fault coverage. We have also demonstrated that double sensing can be used during SRAM normal operation for online detection and correction of any number of random read faults.
Qikai Chen, Hamid Mahmoodi, Swarup Bhunia, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2004 Hardware architecture and VLSI implementation of a low-power high-performance polyphase channelizer with applications to subband adaptive filtering
abstract
The polyphase channelizer is an important component of a subband adaptive filtering system. This paper presents an efficient hardware architecture and VLSI implementation of a low-power high-performance polyphase channelizer, integrating optimizations at algorithmic, architectural and circuit level. At the algorithm level, a computationally efficient structure is derived. Tradeoffs between hardware complexity and system performance are explored during the fixed-point modeling of the system. A computational complexity reduction technique is also employed to reduce the complexity of the hardware architecture. Circuit-level optimizations, including an efficient commutator implementation, dual-VDD scheme and novel level-converting flip-flops, are also integrated. Simulation results show that the design consumes 352 mW power with system throughput of 480 million samples per second (MSPS). A test chip has been submitted for fabrication to validate the proposed hardware architecture and VLSI design techniques.
Yongtao Wang, Hamid Mahmoodi, Lih-Yih Chiou, Hunsoo Choo, Jongsun Park 0001, Woopyo Jeong, Kaushik Roy 0001
ICASSP (5)2
2004 Statistical design and optimization of SRAM cell for yield enhancement
abstract
We have analyzed and modeled the failure probabilities of SRAM cells due to process parameter variations. A method to predict the yield of a memory chip based on the cell failure probability is proposed. The developed method is used in an early stage of a design cycle to minimize memory failure probability by statistically sizing of SRAM cell.
Saibal Mukhopadhyay, Hamid Mahmoodi, Kaushik Roy 0001
ICCAD2
2004 A Novel Low-Power Scan Design Technique Using Supply Gating
abstract
Reduction in test power is important to improve battery life in portable devices employing periodic self-test, to increase reliability of testing and to reduce test-cost. In scan-based testing, about 80% of total test power is dissipated in the combinational block. In this paper, we present a novel circuit technique to virtually eliminate test power dissipation in combinational logic by masking signal transition at the logic inputs during scan shifting. We realize the masking effect by inserting an extra supply gating transistor in the VDD to GND path for the first level cells at output of the scan flops. The supply gating transistor is turned off in the scan-in mode, essentially gating the supply. Adding an extra transistor in only one logic level renders significant advantage with respect to area, delay and power (in normal mode of operation) overhead compared to existing methods, which use gating logic at the output of scan flops. Simulation results on ISCAS89 benchmarks show up to 79% improvement in area, up to 32% in power (in normal mode) and up to 7% in delay compared to lowest-cost known alternative.
Swarup Bhunia, Hamid Mahmoodi, Saibal Mukhopadhyay, Debjyoti Ghosh, Kaushik Roy 0001
ICCD2
2003 Energy recovery clocking scheme and flip-flops for ultra low-energy applications
abstract
A significant fraction of the total power in highly synchronous systems is dissipated over clock networks. Hence, low-power clocking schemes would be promising approaches for future designs. We propose four novel energy recovery flip-flops that enable energy recovery from the clock network, resulting in significant energy savings. The proposed flip-flops operate with a single-phase sinusoidal clock, which can be generated with high efficiency. Based on the simulation results using TSMC 0.25mm CMOS process technology, at a frequency of 200MHz, the proposed flip-flops exhibit more than 80% delay reduction, power reduction of up to 46%, and area reduction of up to 77%, as compared to the conventional energy recovery flip-flop. We implemented 1024 proposed energy recovery flip-flops through an H-tree clock network driven by a resonant clock-generator that generates a sinusoidal clock. Results show a power reduction of 90% on the clock-tree and total power savings of up to 83% as compared to the same implementation using the conventional square-wave clocking scheme and flip-flops.
Matthew Cooke, Hamid Mahmoodi, Kaushik Roy 0001
ISLPED2
2003 Leakage current mechanisms and leakage reduction techniques in deep-submicrometer CMOS circuits
abstract
High leakage current in deep-submicrometer regimes is becoming a significant contributor to power dissipation of CMOS circuits as threshold voltage, channel length, and gate oxide thickness are reduced. Consequently, the identification and modeling of different leakage components is very important for estimation and reduction of leakage power, especially for low-power applications. This paper reviews various transistor intrinsic leakage mechanisms, including weak inversion, drain-induced barrier lowering, gate-induced drain leakage, and gate oxide tunneling. Channel engineering techniques including retrograde well and halo doping are explained as means to manage short-channel effects for continuous scaling of CMOS devices. Finally, the paper explores different circuit techniques to reduce the leakage power consumption.
Kauschick Roy, Saibal Mukhopadhyay, Hamid Mahmoodi
Proc. IEEE3
2002 High performance and low power FIR filter design based on sharing multiplication
abstract
We present a high performance and low power FIR filter design, which is based on computation sharing multiplier (CSHM). CSHM specifically targets computation re-use in vector-scalar products and is effectively used in our FIR filter design. Efficient circuit level techniques: a new carry select adder and conditional capture flip-flop (CCFF), are also used to further improve power and performance. The proposed FIR filter architecture was implemented in 0.25 μm technology. Experimental results on a 10 tap low pass CSHM FIR filter show speed and power improvement of 19% and 17%, respectively, with respect to an FIR filter based on Wallace tree multiplier.
Jongsun Park 0001, Woopyo Jeong, Hunsoo Choo, Hamid Mahmoodi, Yongtao Wang, Kaushik Roy 0001
ISLPED4