EDBT 2026 Demo / reviewers in the wild / expert
Yehea I. Ismail
dblp:i/YeheaIIsmail · also Yehea Ismail
· DBLP profile ↗
111ranked-venue papers
22as first author
3since 2021 · last 2025
0000-0003-3956-7533ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 110 · 22 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorSecurity and privacy · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A 0.4 V, 12.2 pW Leakage, 36.5 fJ/Step Switching Efficiency Data Retention Flip-Flop in 22 nm FDSOIabstractData-retention flip-flops (DR-FFs) efficiently maintain data during sleep mode, and retain state during transitions between active and sleep mode. This brief proposes an ultralow power DR-FF design with an improved autonomous data-retention (ADR) latch operating with a supply voltage range down to near/subthreshold, achieving a sleep mode leakage power of 12.2 pW,$1.4\times $–$3.8\times $less than the prior CMOS DR-FFs. Our proposed DR-FFs consume the lowest active mode switching efficiency of 36.5 fJ/step,$1.2\times $–$4\times $less than the prior works, and a comparable transition efficiency of 1.9 fJ/step. Furthermore, our proposed DR-FFs require minimal control signals, logic gates, and switches, significantly reducing design complexity, and avoiding the drawbacks of nonvolatile data retention FFs (NV-FFs). Yuxin Ji, Yuhang Zhang 0008, Changyan Chen, Jian Zhao 0004, Fakhrul Z. Rokhani, Yehea I. Ismail, Yongfu Li 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2022 | A Fast and Accurate Middle End of Line Parasitic Capacitance Extraction for MOSFET and FinFET Technologies Using Machine LearningabstractA novel machine learning modeling methodology for parasitic capacitance extraction of middle-end-of-line metal layers around FinFETs and MOSFETs is developed. Due to the increasing complexity and parasitic extraction accuracy requirements of middle-end-of-line patterns in advanced process nodes, most of the current parasitic extraction tools rely on field-solvers to extract middle-end-of-line parasitic capacitances. As a result, a lot of time, memory, and computational resources are consumed. The proposed modeling methodology overcomes these problems by providing compact models that predict middle-end-of-line parasitic capacitances efficiently. The compact models are pre-characterized and technology-dependent. Also, they can handle the increasing accuracy requirements in advanced process nodes. The proposed methodology scans layouts for devices, extracts geometrical features of each device using a novel geometry-based pattern representation, and uses the extracted features as inputs to the required machine learning models. Two machine learning methods are used including: support vector regressions and neural networks. The testing covered more than 40M devices in several different real designs that belong to 28nm and 7nm process technology nodes. The proposed methodology managed to provide outstanding results as compared to field-solvers with an average error < 0.2%, a standard deviation < 3%, and a speed up of 100X. Mohamed Saleh Abouelyazid, Sherif Hammouda, Yehea I. Ismail |
ASP-DAC | 3 |
| 2022 | Accuracy-Based Hybrid Parasitic Capacitance Extraction Using Rule-Based, Neural-Networks, and Field-Solver MethodsabstractAs process technologies scale down, the accuracy requirements of parasitic capacitance extractions for integrated circuits significantly increase. This work introduces a novel accuracy-based hybrid parasitic capacitance extraction flow, where the chip is subdivided into windows, and each window’s capacitances are calculated using one of three extraction methods: field-solver, rule-based, and novel deep-neural-networks-based methods. This hybrid methodology uses a density-map feature representation as an input to neural-networks classifiers to determine an extraction method for each window. As an intermediate method between rule-based and field-solver methods, a novel deep-neural-networks-based extraction method is introduced. This intermediate level of accuracy and speed is needed since using only rule-based and field-solver methods results in using the field-solver most of the time for any required high accuracy extraction. This method uses a novel hybrid density-voltage representation as an input to improve its accuracy and speed. The proposed hybrid flow identifies the accuracy limits of the three extraction methods and directs each window to the fastest method that meets the user predetermined accuracy level. The proposed flow is tested on different real designs and showed outstanding accuracy and runtime as compared to commercial field-solver and rule-based tools. The results show that the proposed deep-neural-networks extraction method extracts capacitances of complicated structures with high accuracy ($100\times $faster than field-solvers. However, few outliers have an error exceeding 5% in extracted capacitances. Furthermore, the proposed hybrid flow managed to meet the required accuracy ($70\times $faster than field-solvers. Mohamed Saleh Abouelyazid, Sherif Hammouda, Yehea I. Ismail |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | An accurate model of domain-wall-based spintronic memristor
Sherif F. Nafea, Ahmed A. S. Dessouki, S. El-Rabaie 0001, Basem E. Elnaghi, Yehea I. Ismail, Hassan Mostafa |
Integr. | 5 |
| 2018 | Fully Integrated Mixed Mode Interface Circuit in 65 nm CMOS for Leukemia Detection and ClassificationabstractLab on a Chip (LOC) applications are taking the world by storm. This paper proposes a mixed mode interface circuit for LOC applications for Bioimpedance measurement. The interface circuit consists of a low noise instrumentation amplifier and an analog to digital converter. The proposed interface circuit utilizes the benefits of both current mode and voltage mode circuits. The design of the instrumentation amplifier uses a second-generation current conveyor to provide the low noise and high gain required to acquire the weak signal from the electrodes. The IA provides a maximum gain of 55.6 dB with 2 MHz bandwidth. A second order Sigma Delta Modulator with a maximum sampling rate of 128 MS/s provides the output digital bit stream capable of reaching a resolution of 11 bits for an input sinusoidal with 50 kHz frequency. The use of a buffered comparator allows the relaxation on the slew rate requirement while maintaining a high sampling rate. The full system power consumption is 5.5 mW with a chip area of only 0.0153 mm2. Mohammed A. Eldeeb, Yehya H. Ghallab, Yehea I. Ismail, Hassan El-Ghitani |
ISCAS | 3 |
| 2018 | Electroporation Improvement of Leukemic Cells Using Dielectrophoresis TechniqueabstractElectroporation is a powerful technique used to improve the permeability of the cell membrane; pores are created in the cell membrane based on the application of a high voltage. The high voltage improves the number of membrane pores, however, it might kill the cells. Electroporation is achieved using a microelectrode placed in a microfluidic device. This paper focuses on the improvement of the electroporation of a cell based on levitation dielectrophoresis and an electrorotation force. Then, the change of the conductivity is detected at both the frequency and time domain, i.e., individual cells are trapped due to the effect of levitation and electrorotation force, and then the variation in the conductivity is detected. We use frequency and time domain spectroscopy to measure the conductivity of the white blood cell, i.e., B-cell, T-cell and leukemia K365. In this paper, we review studies of conductivity separation in a microfluidic device based on a 2-D electrode system. Also, a comparison between the conductivity of leukemia k365, B-cell, and T-cell due to the drag and electrorotaional forces are presented and discussed. Sameh Sherif, Yehya H. Ghallab, Yehea I. Ismail |
ISCAS | 3 |
| 2018 | Optimizing FPGA-based hard networks-on-chip by minimizing and sharing resources
Sameh Attia, Hossam A. H. Fahmy, Yehea I. Ismail, Hassan Mostafa |
Integr. | 3 |
| 2017 | Power efficient AES core for IoT constrained devices implemented in 130nm CMOSabstractThe Internet of Things (IoT) constrained devices show the urgent need for low power data security hardware cores. This paper presents a power efficient AES Core fabricated in UMC 130 nm CMOS technology by using Faraday standard cells library. The maximum throughput of the proposed AES Core is up to 2.6 Gb/s consuming about 0.2148 mW/MHz at 1.2V. The Dynamic Voltage and Frequency Scaling (DVFS) technique is applied to reduce the power consumption of the AES Core. The experimental measurements show about 3x reduction in power consumption, consuming about 0.0697 mW/MHz by scaling the supply voltage from 1.2 V to 0.7 V. Shady O. Agwa, Eslam Yahya, Yehea I. Ismail |
ISCAS | 3 |
| 2017 | 130nm Low power asynchronous AES coreabstractInternet of Things (IoT) devices are always having low power budget and high security demands. This paper describes the design and results of fabricated Advanced Encryption Standard (AES) chip in UMC 130 nm CMOS technology by using Faraday standard cells. The AES core is designed in fully QDI asynchronous circuit style. The core ciphers 128-bit data/key in 300 ns and consumes 5.47 mW. Nada El-meligy, Moustafa Amin, Eslam Yahya, Yehea I. Ismail |
ISCAS | 4 |
| 2017 | Design guidelines for the high-speed dynamic partial reconfiguration based software defined radio implementations on Xilinx Zynq FPGAabstractReconfigurability of Field Programmable Gate Array (FPGA) makes it one of the most promising approaches in the implementation of the Software Defined Radio (SDR). FPGA Dynamic Partial Reconfiguration (DPR) feature emphasizes that approach by allowing the implemented SDR system to switch between multiple communications standards in runtime reusing the same FPGA hardware resources. Reconfiguration time is a significant parameter in DPR designs especially when a fast switching is required in real time system like SDR. In this paper, different designs of Partial Reconfiguration (PR) controllers are studied and evaluated according to their impact to improve the reconfiguration time of DPR-based SDR implementation. A multi-standard convolutional encoder design is implemented using DPR with different PR controllers as a case study. The design is implemented and tested on Xilinx Zynq evaluation board “ZC702”. This comparative study provides important design insights and recommendations to the DPR-based SDR designers to help them select the best PR controller based on their system throughput requirement and power budget. Ahmed Kamaleldin, Ahmed M. Soliman, Ahmed Nagy, Youssef Gamal, Ahmed Shalash, Yehea I. Ismail, Hassan Mostafa |
ISCAS | 6 |
| 2017 | Dielectric analysis of changes in electric properties of leukemic cells through travelling and negative dielectrophoresis with 2-D electrodesabstractDielectrophoretic force has been used to manipulate biological microparticles, such as red blood cells, white blood cells, cancer cells, etc... This technique has been used for trapping, sorting and separating biological particles suspended in a buffer medium. The strength of dielectrophoretic force depends strongly on the particle's electrical properties, the geometry of the microfluidic device, dielectric properties of the medium, particle shape and size and the frequency of the electric field. Therefore, to design an effective microfluidic separation platform, it is important to understand the effect and the role of the aforementioned parameters on the particle's motion. In this paper, the dielectrophoretic separation in a microfluidic device based on 2D electrodes is studied. Also, a comparison between the velocity of B and T leukemia K365 cells due to the drag force is conducted and presented. Moreover, the conductance of the leukemia K365 cells before and after dragging is presented and discussed. Sameh Sherif, Yehya H. Ghallab, Hamdy Abd Elhamid, Yehea I. Ismail |
ISCAS | 4 |
| 2017 | PASSIOT: A Pareto-optimal multi-objective optimization approach for synthesis of analog circuits using Sobol' indices-based engine
Taher Kourany, Maged Ghoneima, Emad Hegazi, Yehea I. Ismail |
Integr. | 4 |
| 2016 | RC-In-RC-Out model order reduction via node mergingabstractIn this paper, we introduce a method for realizable reduction of extracted RC netlists by merging nodes. This method can achieve high reduction (reaching 96%) with high accuracy and can be used to complement existing techniques of realizable reduction such as TICER [6]. The method preserves sparsity; has controllable accuracy and can result in lossless reduction (exact reduction) for certain circuits. The node merging translates to a simple matrix operation and thus can be easily adopted commercially and realized in CAD tools. Manar Abdel-Galil, Hazem Hegazy, Yehea I. Ismail |
ISCAS | 3 |
| 2016 | Evaluation of multi-level buck converters for low-power applicationsabstractThis work investigates power losses in conventional and multi-level buck converters. Conduction and switching losses are modeled for conventional, 3-level and 5-level buck converters as a function of technology parameters, design variables and the operating point. It is shown that for a given set of technology parameters and optimized design variables, multi-level buck converters achieve higher efficiencies than the conventional buck converter at low output voltages and load currents. This shows that multi-level buck converters are more suitable for low power applications than the conventional buck converter. In addition, for low inductor quality factor (QF), multi-level buck converters achieve higher efficiencies than the conventional buck converter which makes them better suited for on-chip integration. The 65nm CMOS technology is employed. All models are validated against Spice simulations. Abdullah Abdulslam, S. H. Amer, Ahmed S. Emara, Yehea I. Ismail |
ISCAS | 4 |
| 2016 | Five-level hybrid DC-DC converter with enhanced light-load efficiencyabstractIn this paper, a 5-level hybrid converter is proposed. The circuit structure and the working principle are illustrated. The circuit is capable of providing five different voltage levels at the inductor input with the help of two flying capacitors. By reducing the voltage difference at the inductor input, the 5-level buck converter can use smaller inductors as compared to both 3-level and conventional buck converters which makes it more suitable for on-chip DC-DC conversion. Simulation results on TSMC 65nm technology using 0.5nH on-chip spiral inductor show that the 5-level hybrid converter achieves more than 15% improvement in efficiency over a 3-level buck converter at certain output voltage ranges. A test PCB is also implemented for verification of the functionality and experimental measurements show that the 5-level hybrid converter can achieve better efficiency as compared to conventional and 3-level Buck converters especially at low load currents. Abdullah Abdulslam, Farid El-Sehrawy, Yehea I. Ismail |
ISCAS | 3 |
| 2016 | TCG-SP: An improved floorplan representation based on an efficient hybrid of Transitive Closure Graph and Sequence PairabstractThis paper presents TCG-SP, a hybrid floorplan representation based on full and cost-effective integration of Transitive Closure Graph (TCG) and Sequence Pair (SP). TCG-SP efficiently reduces the complexity in the operations performed on placement modules, and in the transformation between a representation and its floorplan, compared to other approaches. TCG-SP proposes a fast O(n log n) time complexity scheme to construct SP from placement. In addition, by retaining a full knowledge of SP ordering sequence in TCG-SP representation, an O(n log log n) time complexity packing cost function calculation is attainable. Furthermore, TCG properties allow a dynamic update of geometric dependency among blocks in a floorplan during solution perturbation. Using this property along with SP ordering property, TCG-SP proposes a fast O(n) runtime cost function update to perturb a placement solution. TCG-SP based simulated annealing scheme for floorplan design is developed. Our floorplanner is simulated with real objective functions and proven to be competitive in terms of runtime and solution quality. Taher Kourany, Emad Hegazi, Yehea I. Ismail |
ISCAS | 3 |
| 2016 | Accuracy-improved coupling capacitance model for through-silicon via (TSV) arrays using dimensional analysisabstractIn this paper, we show that using the relation between the inductance matrix and the capacitance matrix in a homogeneous medium to extract the coupling capacitance in a through silicon via (TSV) array is inaccurate. This is because this relation assumes a lossless, homogeneous surrounding medium. We show that this model can cause an error up to 70% in coupling capacitance compared to Q3D extractor simulations. Instead of using high accuracy, time consuming numerical electromagnetic techniques, we suggest a correction methodology for the inverse inductance model so that it can account for the non-homogeneous nature of the TSVs surrounding medium and the lossy nature of the silicon substrate. Dimensional analysis is used to understand the correction function dependencies on TSV dimensions and to reduce the number of independent variables needed for regression analysis. Once the independent variables are reduced, multiple regression techniques are used to estimate the correction function for TSV arrays. Q3D extractor simulations are used to show that the corrected model reduces the coupling capacitance error significantly. Then, the correction function behavior vs. frequency is discussed. Tarek Ramadan, Eslam Yahya, Mohamed Dessouky, Yehea I. Ismail |
ISCAS | 4 |
| 2015 | A tunable multi-band/multi-standard receiver front-end supporting LTEabstractCurrent wireless communication devices demand multi-band/multi-standard receiver that can access all the available services specifications. This work introduces a tunable receiver front-end for multi-band multi-standard applications. The receiver adopts a down-conversion quadrature band-pass FIR charge sampling mixer tuned via its controlling clocks. A time varying impedance matching network provides further selectivity. The architecture is simulated over three different frequencies spanning two octaves (2G, 1G and 500MHz) targeting LTE specifications. The proposed design is tested across process corners and post layout. Simulations result in Noise Figure of 7.5 to 9 dB, out-of-band IIP3 of -1.9 to -7 dBm, in-band IIP3 of -1.5 to -10 dBm and S11 <;-10dB. The design is implemented using 65nm CMOS technology. Hoda Abdelsalam, Emad Hegazi, Hassan Mostafa, Yehea I. Ismail |
ISCAS | 4 |
| 2015 | Analysis and optimization for dynamic read stability in 28nm SRAM bitcellsabstractThe importance of the dynamic analysis for SRAM operation increases as a result of shrinking access cycle time, voltage scaling and increased process variations. In this paper, quantitative study of the dynamic read noise margin (DNM) is introduced showing the evolution from the static read noise margin (SNM) to DNM through cumulative dynamic effects in 28nm FDSOI. The impact of parasitic capacitances on the DNM is further analyzed. Finally, we show that by sizing for a 150-mV DNM instead of a 150-mV SNM and by inserting two 0.5fF extra caps in the bitcell allows reducing the pull-down NMOS width by a factor 3.5×. Ahmed T. Elthakeb, Thomas Haine, Denis Flandre, Yehea I. Ismail, Hamdy Abd Elhamid, David Bol |
ISCAS | 4 |
| 2015 | Design of adiabatic TSV, SWCNT TSV, and Air-Gap Coaxial TSVabstractHigh performance three-dimensional (3D) through silicon via (TSV) interconnects are important for reliability, choice of the filler material is also a critical issue as thermal incompatibility, electromigration, and high resistivity are still a bottleneck. In this paper, single wall carbon nanotube (SW-CNT) bundles as a prospective filler material for TSV are investigated compared to conventional filler materials like Cu, W, and poly-silicon. It is found that SW-CNT bundles exhibit unique electrical, thermal, and mechanical characteristics that can be used to fabricate better TSV interconnects. Moreover, performance comparison between Air-Gap Based Coaxial TSV and conventional circular TSV are presented. The comparison shows that the air-gap TSVs reduce the overall parasitic capacitance and the overall energy loss compared to the conventional circular TSV or conventional coaxial TSV. In addition, TSV-based ADIABATIC logic based on the adiabatic switching principle is presented and analyzed. ADIABATIC logic is a design technique for minimizing the energy dissipation. Its major limitation is the requirement for passive components, which cannot be efficiently integrated into current generation ICs. TSV-based 3D heterogeneous integration may enable efficient integration of these passive elements, which were not practically feasible in the past due to technology limitations. Khaled Salah 0001, Yehea I. Ismail |
ISCAS | 2 |
| 2014 | Low-power all-digital manchester-encoding-based high-speed serdes transceiver for on-chip networksabstractThis paper proposes a new architecture that multiplexes both data and clock on serial links, reduces Inter symbol interference (ISI) by using a resistive termination technique, and uses two-level Manchester encoding to solve the reduced swing problem and enable the use of power efficient circuitry. Using this signaling scheme makes the system insensitive to jitter accumulation along the transmission line, and avoids the need for a power hungry clock and data recovery (CDR) circuit. A self-calibrating digital-delay line is also implemented inside the decoder to enable the system to operate efficiently across process, voltage and temperature variations. The proposed scheme is implemented for a 3mm long on-chip transmission line in TSMC 65nm technology and simulation results are presented. Abdelrahman H. Elsayed, Ramy N. Tadros, Maged Ghoneima, Yehea I. Ismail |
ISCAS | 4 |
| 2014 | A novel dimensional analysis method for TSV modeling and analysis in three dimensional integrated circuitsabstractDimensional analysis is one of the most powerful modeling methods, where it is used to reduce the number of physical variables through combining two or more variables into a single dimension neutral one in order to simplify the process of describing a relationship among those variables. That makes curve fitting to obtain final equations simpler. In this paper, dimensional analysis is applied as a new design methodology for TSV modeling. By using the dimensional analysis and the curve-fitting technique, simple formulas of TSV parameters, such as resistance, capacitance, and inductance are obtained. The results are compared with previous work that model TSVs based on measurements, where excellent agreement is obtained. Khaled Salah 0001, Yehea I. Ismail |
ISCAS | 2 |
| 2014 | A variation tolerant driving technique for all-digital self-timed 3-level signaling high-speed SerDes transceivers for on-chip networksabstractThis paper presents a variation tolerant driving technique for all-digital self-timed 3-level signaling high-speed SerDes transceivers. The proposed design generates the 3-level signal without a ½VDD driver, thus removing all the overhead and hassle of an additional supply. Moreover, the proposed all-digital scheme uses half the clock frequency while maintaining the same data rate of the conventional scheme. As a result, the proposed design is much more robust across all possible variations. The transceiver is designed for a 3mm long lossy on-chip differential interconnect in TSMC 65nm CMOS technology. The transmitter serializes the parallel 1.9375Gbps 8-bit, then multiplexes it with the 15.5GHz clock to generate the three level signal on the differential TL, and a simple RX detects both the data and the clock from the signal. Ramy N. Tadros, Abdelrahman H. Elsayed, Maged Ghoneima, Yehea I. Ismail |
ISCAS | 4 |
| 2013 | TSV-based on-chip inductive coupling communicationsabstractThis paper presents a novel wireless through silicon via (TSV) communication structure based on near-field inductive-coupling. The proposed system uses a spiral inductor built using TSV technology. The electrical performance obtained from EM-based S-parameter simulations for different number of turns and configurations are presented. The proposed wireless TSV shows good coupling coefficient. A closed form expression for the coupling coefficient is presented. This coupling expression is verified against EM simulations and showed good agreement. The proposed wireless TSV system provides small area (30 μm)2for typical design values. This proposed communication system is enabling NoC. Khaled Salah 0001, Alaa B. El-Rouby, Hani F. Ragai, Yehea I. Ismail |
ISCAS | 4 |
| 2013 | Editorial Appointments for the 2013-2014 TermabstractYehea I. Ismail will serve another term as the Editor-in- Chief (EiC) of the IEEE Transactions on Very Large Scale Integration (VLSI) Systems. Previous EiCs of this journal have worked tirelessly to bring this journal to the very top. During his previous term the impact factor of TVLSI has moved up from 0.9 to 1.2, which is a significant improvement. Dr Ismail will work with the steering committee and the board to further improve the impact factor of TVLSI. He looks forward to the whole community helping with this goal by submitting quality work, accepting reviews, giving valuable comments, and making suggestions. TVLSI had a backlog in the acceptance to publication time because of a large number of submissions to the Journal. We have increased the number of pages of TVLSI from 1800 to 2400 and have not accepted any special issues resulting in a significant improvement in this backlog. We will continue this policy and find other ideas to completely eliminate the problem. This item also announces the appointment of Professor Massimo Alioto of NUS as the Associate EiC of TVLSI. Dr. Alioto has a long editorial record and is a well-known scholar in the VLSI area. In addition, the biographies and photographs of the Associate Editors on the new editorial board are provided. Yehea I. Ismail |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2012 | InMnAs magnetoresistive spin-diode logicabstractElectronic computing relies on systematically controlling the flow of electrons to perform logical functions. Various technologies and logic families are used in modern computing, each with its own tradeoffs. In particular, diode logic allows for the execution of logic with many fewer devices than complementary metal-oxide-semiconductor (CMOS) architectures, which implies the potential to be faster, cheaper, and dissipate less power. It has heretofore been impossible to fully utilize diode logic, however, as standard diodes lack the capability of performing signal inversion. Here we create a binary logic family based on high and low current states in which the InMnAs magnetoresistive semiconductor heterojunction diodes implement the first complete logic family based solely on diodes. The diodes are used as switches by manipulating the magnetoresistance with control currents that generate magnetic fields through the junction. With this device structure, we present basis logic elements and complex circuits consisting of as few as 10% of the devices required in their conventional CMOS counterparts. These circuits are evaluated based on InMnAs experimental data, and design techniques are discussed. As Si scaling reaches its inherent limits, this spin-diode logic family is an intriguing potential replacement for CMOS technology due to its material characteristics and compact circuits. Joseph S. Friedman, Nikhil Rangaraju, Yehea I. Ismail, Bruce W. Wessels |
ACM Great Lakes Symposium on VLSI | 3 |
| 2012 | A 16Gbps low power self-timed SerDes transceiver for multi-core communicationabstractThis paper presents a modified design for a self-timed SerDes transceiver that was recently published [1]. The new architecture overcomes the main problems that arise in [1], while offering the same advantages. Resistive termination is used instead of source matching to eliminate the need for Manchester coding in [1], this resistive termination increased the data rate to be 16Gbps compared to 12Gbps in [1]. Moreover, resistive termination removed the limitation on the minimum operating frequency that existed in [1], solving a lot of problems at the slow process corners. A single ended transmission line is used instead of the differential transmission line in [1]. A calibration circuit is implemented to control the switching threshold of the detector at the receiver side to account for voltage and process variations. The SerDes transceiver is implemented for a 3mm long on-chip transmission line in 65nm TSMC CMOS technology, which is the same as [1]. The total power consumed in the Tx/Rx pair with the transmission line is 18.1mWatt, compared to 15.5mWatt in [1]. The proposed architecture have the same advantages in [1] of being self timed, eliminating the need for complex power hungry blocks such as Clock and Data Recovery (CDR) at the receiver, and being insensitive to jitter accumulated during transmission. Ezz El-Din O. Hussein, Sally Safwat, Maged Ghoneima, Yehea I. Ismail |
ISCAS | 4 |
| 2012 | A 5-10GHz low power bang-bang all digital PLL based on programmable digital loop filterabstractThis paper presents the design and the implementation of a low power bang-bang all digital phase locked loop (BBADPLL). The design of the proposed architecture is based on the programmable coefficients of the digital loop filter (DLF) that manages the tradeoffs between stability and jitter of a closed loop. A proposed simple digital controlled oscillator (DCO) based on three stages ring oscillator provides a wide frequency range, and proven to be of lower area and power compared to arrayed DCO. The proposed design results in a significant reduction in the area and power compared to other time-to-digital converter (TDC) based ADPLL architectures. This reduction results from eliminating the need for complex, power, and area consuming TDC block, and arrayed DCO. A counter-based frequency acquisition loop using a binary search algorithm reduced the lock-in time significantly compared to similar work. The proposed BBADPLL architecture was implemented on TSMC CMOS 65nm technology with a frequency range 5–10GHz and a frequency resolution equals to 500MHz. The lock-in time is 2.4µs. The peak-to-peak period jitter and the RMS jitter at 10GHz are 1.49ps and 0.19ps, respectively. The total power consumed at 10GHz is only 2.7mWatt and the total area of the proposed ADPLL is 4372µm2, which is very small compared to other published architectures. Sally Safwat, Amr Lotfy, Maged Ghoneima, Yehea I. Ismail |
ISCAS | 4 |
| 2012 | A closed form expression for TSV-based on-chip spiral inductorabstractThis paper presents a new architecture for on-chip spiral inductors, based on through silicon via (TSV) technology. A highly accurate closed form expression for the proposed TSV-based spiral inductor equivalent inductance is presented. This closed form is the first in literature. Moreover, this form is verified against a large number of EM simulations for different setups and shows excellent agreement with less than 5% error. According to the simulation results, the TSV-based spiral inductor exhibits better quality factor than a planar on-chip spiral inductor (120% improvement) and a 3D via-based spiral inductor (76% improvement) for identical inductance. Moreover, the self-resonance frequency of the proposed TSV-based 3D spiral inductor is 38% higher than the conventional 2D spiral inductor and 3% higher than the 3D via-based inductor. The proposed inductor occupies only 15% of the area of the conventional 2D planar spiral inductor and 60% of the area of 3D via-based spiral inductor for the same inductance and higher quality factor. This small area of the proposed inductor leads to significant reducing the cost of the radio frequency system on chip (RF-SoC). Khaled Salah 0001, Alaa B. El-Rouby, Hani F. Ragai, Yehea I. Ismail |
ISCAS | 4 |
| 2012 | Switched-capacitor dc-dc converters with output inductive filterabstractAnalysis and optimization of switched-capacitor (SC) dc-dc converters with a series inductive filter are developed. The steady-state output impedance of such SC resonant converters is calculated for a 2:1 conversion ratio. In addition, the necessary conditions for proper application of the output inductive filter are derived. The proposed optimization methodology applies numerical optimization to evaluate different loss components in order to find the optimal design point of highest conversion efficiency. This optimization method is verified through SPICE simulations on a 2:1 SC power stage in 65-nm CMOS process. Loai G. Salem, Yehea I. Ismail |
ISCAS | 2 |
| 2011 | Equivalent lumped element models for various n-port Through Silicon Vias networksabstractThis paper proposes an equivalent lumped element model for various multi-TSV arrangements and introduces closed form expressions for the capacitive, resistive, and inductive coupling between those arrangements. The closed form expressions are in terms of physical dimensions and material properties and are driven based on the dimensional analysis method. The model's compactness and compatibility with SPICE simulators allows the electrical modeling of various TSV arrangements without the need for computationally expensive field-solvers and the fast investigation of a TSV impact on a 3-D circuit performance. The proposed model accuracy is tested versus a detailed electromagnetic simulation and showed less than 6% difference. Finally, the proposed model can be a possible solution to the industrial need for broadband electrical modeling of TSVs interconnections arising in 3-D integration. Also, our presented work provides valuable insight into creating guidelines for TSV macro-modeling. Khaled Salah 0001, Hani F. Ragai, Yehea I. Ismail, Alaa B. El-Rouby |
ASP-DAC | 3 |
| 2011 | A Comprehensive Tapered buffer optimization algorithm for unified design metricsabstractTapered buffers are widely used in CMOS integrated circuits to drive large capacitive loads. During the design of a tapered buffer, there are several design objectives to consider including delay, area, and power consumption. Existing methods produce suboptimal solutions considering multiple metrics, largely because they decouple different metrics during the design phase, which restricts the solution space for the combined metric. In this paper, a new algorithm is proposed to derive the optimal solution for a unified metric through comprehensive exploration of the solution space. Compared with existing methods, our method yields as high as 18.8% (9.0% on average) improvement in a unified design metric optimizing area, delay, and power simultaneously. The proposed algorithm also reduces buffer power consumption under delay constraints. The power reduction over existing alternatives is as much as 48.1% (28.7% on average). Seda Ogrenci Memik, Yehea I. Ismail |
ISCAS | 3 |
| 2011 | A dynamic calibration scheme for on-chip process and temperature variationsabstractA process and temperature variation calibration scheme is proposed in this paper. The proposed system uses the supply voltage and body bias to calibrate the device parameters to match those of a certain process corner that is determined by the system designer. This scheme is characterized by its ability to dynamically change the desired mapping target according to the computational load. Moreover, the proposed system provides the ability to detect and control the n- and p-type variations independently through the use of an all-n and all-p ring oscillators. The calibration system has been implemented and simulated in TSMC 90-nm technology. Simulation results show that the system was able to reduce frequency spread (sigma) from 75 MHz to an average of 10 MHz and frequency variations from 34% to 3.1%. The results also show the system's ability to compensate for dynamic load variations. Mina Raymond, Maged Ghoneima, Yehea I. Ismail |
ISCAS | 3 |
| 2011 | A 12Gbps all digital low power SerDes transceiver for on-chip networkingabstractIn this paper, a new self-timed signaling technique for reliable low-power on-chip SerDes (Serialization and DeSerialization) links is presented. The transmitter serializes 8 parallel bits at 1.5GHz, and multiplexes the 12Gbps serial data stream with a 24GHz clock on a single line using three level signaling. This new signaling technique enables the receiver to recover the clock from the data with a simple phase detector circuitry. Moreover, this technique is insensitive to jitter accumulated during signal propagation or at the receiver input because the clock signal is extracted from the multiplexed data stream. Hence, timing errors in the received signal reflects in both the data and the extracted clock, and the data will be sampled correctly. The SerDes transceiver was implemented for a 3mm long lossy on-chip differential transmission line in 65nm TSMC CMOS technology. A primary advantage of building an all digital SerDes transceiver is the ease of scaling with technology, and the power and area reduction. The total power consumed in the Tx/Rx pair with the transmission line is 15.5mWatt, which is very small as compared to similar published signaling architectures. Sally Safwat, Ezz El-Din O. Hussein, Maged Ghoneima, Yehea I. Ismail |
ISCAS | 4 |
| 2011 | Compact lumped element model for TSV in 3D-ICsabstractA wide-band lumped element model for a through silicon via (TSV) is proposed based on electromagnetic simulations. Closed form expressions for the TSV parasitics based on the dimensional analysis method are introduced. The proposed model enables direct extraction of the TSV resistance, self-inductance, oxide capacitance, and parasitic elements due to the finite substrate resistivity. The model's compactness and compatibility with SPICE simulations allows the fast investigation of a TSV impact on a 3-D circuit performance. The parameters' values of the proposed TSV model are fitted to the simulated S-parameters up to 10 GHz with an error less than 5%. It is shown that a TSV capacitance is highly dependent on the positions of ground contacts and has a value of tens of femto farads in a typical current technology. This value is much higher than a minimum device capacitance and requires special design methodologies such as cascaded buffers. Coupling between TSVs will be handled in another paper. Khaled Salah 0001, Alaa B. El-Rouby, Hani F. Ragai, Karim Amin, Yehea I. Ismail |
ISCAS | 5 |
| 2011 | Fast hysteretic control of on-chip multi-phase switched-capacitor dc-dc convertersabstractA novel double-bound hysteretic control of multi-phase switched-capacitor (SC) converters is presented. The technique adjusts the number of interleaved phases with the output load to significantly reduce the operating frequency of the control comparator, enabling the practical application of hysteretic control with large number of interleaved phases. Using the proposed technique, the maximum required speed of the hysteretic comparator is reduced from 7.3 Ghz to 1.8 Ghz in a 16-phase 2:1 SC converter, designed in 65-nm CMOS process. In addition, the achieved dynamic response with such control is much faster than any reported integrated converter. For a 1.2-V input voltage and 0.45-V output voltage, the regulator enables a 35-mVppoutput droop for a 50% load step, without using a decoupling capacitor. In addition, the output for a 100-mV input reference step settles in 2.8-ns. The SC converter's efficiency is not affected by such reconfigurable interleaving scheme and reaches 81%. Loai G. Salem, Yehea I. Ismail |
ISCAS | 2 |
| 2011 | A Novel Moment Based Framework for Accurate and Efficient Static Timing AnalysisabstractA novel methodology for accurate and efficient static timing analysis is presented in this paper. Our methodology uses the traditional cell library table structure with one modification. The cell library tables are filled with the gate output signal moments instead of the gate output 50% delay and output slew. Using only few moments gives much better accuracy and visibility for the gate output waveform than using the time domain information. Simple convolution of the gate output moments with the interconnect moments yields the signal moments at the stage output. The parameters of the gate input signal, which are used for the table access of the successive stage, are directly computed from the predecessor stage output moments using the closed form expressions without having to explicitly transform the frequency domain moments to time domain. Thus, the interconnects and the gates are treated in a unified moment-based homogeneous framework. The proposed approach inherits the classical cell library tables approach efficiency with even reduced computation complexities. As compared to the classical cell library table approach, the proposed approach accounts for the increasingly nonlinear and non-monotonic waveform shapes which are prohibitively difficult to represent in the classical approaches. In contrary to the classical approaches, increasing the accuracy in the novel approach is made flexible and can be achieved by simply using more moments. To illustrate the concept and prove its merits, multiple examples are presented with 2-3 moments which maintain accuracy within 1%-3% as compared to SPICE. Ahmed Shebaita, Debasish Das, Dusan Petranovic, Yehea I. Ismail |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2011 | FA-STAC: An Algorithmic Framework for Fast and Accurate Coupling Aware Static Timing AnalysisabstractThis paper presents an algorithmic framework for fast and accurate static timing analysis considering coupling. With technology scaling to smaller dimensions, the impact of coupling induced delay variations can no longer be ignored. Timing analysis considering coupling is iterative, and can have considerably larger run-times than a single pass approach. We propose two different classes of coupling delay models: heuristic-based coupling model and current source-based coupling model, and present techniques to increase the convergence rate of timing analysis when such coupling models are employed. Our proposed coupling model show promising accuracy improvements compared to SPICE. Experimental results on ISCAS85 benchmarks validates the effec tiveness of our efficient iteration scheme. Our iteration algorithm obtained speedups of up to 62.1 % using a heuristic coupling model while 2.7 x using a current-based coupling model in comparison to traditional approaches. Debasish Das, Ahmed Shebaita, Hai Zhou 0001, Yehea I. Ismail, Kip Killpack |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2011 | Editorial
Yehea I. Ismail |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2010 | A novel variation insensitive clock distribution methodologyabstractA new clock distribution technique is introduced in this paper. The technique avoids repeaters completely and distributes the clock directly on the passive interconnect network. The wires can be highly lossy, yet the clock is delivered with a very good shape and eye. The technique uses the characteristics of the interconnect to attenuate all frequency components equally. The resulting clock at the sinks does not depend on supply variations at all and only depends on the LC time constant of the wires. Interestingly, the technique works even better with higher clock frequencies. Signal equalization and boosting at the clock source is applied to further improve the clock shape at the receivers. Ezz El-Din O. Hussein, Yehea I. Ismail |
ISCAS | 2 |
| 2010 | SACTA: A Self-Adjusting Clock Tree Architecture for Adapting to Thermal-Induced Delay VariationabstractAggressive technology scaling down and low-power design techniques lead to uneven distributed power density, which translates into heat flow in the chips, causing significant temperature variations in both spatial and temporal terms. In order to mitigate the negative impacts of temperature variations on circuit timing, we propose SACTA, a self-adjusting clock tree architecture, which performs temperature-dependent dynamic clock skew scheduling to prevent timing violations in a pipelined circuit. The dynamic and adaptive features of SACTA are enabled by our proposed automatic temperature-adjustable skew buffers and temperature-insensitive skew buffers. These special delay elements are carefully tuned to ensure resilience of the entire circuit against temperature variation. To determine their configurations, we proposed an efficient and general clock tree design and optimization framework. Furthermore, we show that SACTA is applicable across a wide spectrum of circuits, including multi-${V}_{\rm dd}/{V}_{\rm th}$designs. Experimental results show that a pipeline supported by SACTA is able to prevent thermal-induced timing violations within a significantly larger range of operating temperatures (on average, the violation-free range can be enhanced by over 15$^{\circ}\hbox {C}$). Jieyi Long, Ja Chun Ku, Seda Ogrenci Memik, Yehea I. Ismail |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2009 | A Timing-Dependent Power Estimation Framework Considering CouplingabstractIn this paper, a timing-dependent dynamic power estimation framework that considers the impact of coupling in combinational circuits is proposed. Relative switching activities and delays of coupled interconnects significantly affect dynamic power dissipation in parasitic coupling capacitances (coupling power). To enable capturing the switching and timing dependence, detailed switching distributions and timing information are essential in accurate estimation of dynamic power consumption. An approach to efficiently represent and propagate switching and timing distributions through circuits is developed. Based on propagated switching and timing distributions, power consumption in coupling capacitances is accurately calculated. Experimental results using ISCAS'85 benchmarks demonstrate that ignoring timing dependence of coupling power consumption can cause up to 25% error in dynamic power estimation (corresponding to 59% error in coupling power estimation). DiaaEldin Khalil, Debjit Sinha, Hai Zhou 0001, Yehea I. Ismail |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2008 | Interconnect design and limitations in nanoscale technologiesabstractIn this paper, the limitations posed by metallic interconnect in nano-scale technologies are discussed as well as design methodologies to deal with non-ideal interconnect. It is shown that there is a limit on how ideal an on-chip interconnect can be made independent of the how wide the wire is made. This limit is a function of the height of the interconnect, its length, and some other physical and material constants. It is also shown that it is possible to design a repeater system that results in close to speed of light propagation on narrow nanoscale wires. In addition, wider wires can be used to effectively eliminate the need for repeaters in current technologies. Yehea I. Ismail |
ISCAS | 1 |
| 2008 | Power-supply-variation-aware timing analysis of synchronous systemsabstractWith state of the art technology scaling, the problem of delay variability due to power supply variations is becoming more and more critical. This paper addresses the problem of analyzing the speed degradation in synchronous systems caused by power supply IR-drop in deep submicron CMOS devices. Considering the impact of power supply variation on the clock skew value, violations of the timing constraints equations are presented. To satisfy the timing constraints over a range of 20% of VDDvariation, a 42% increase in the operational clock period has to be met with circuits operating at 2 GHz and implemented on 65 nm CMOS technology. Sami Kirolos, Yehia Massoud, Yehea I. Ismail |
ISCAS | 3 |
| 2008 | Accurate analytical delay modeling of CMOS clock buffers considering power supply variationsabstractIn this paper, we present an accurate method for analytical derivation of CMOS clock buffers delay under power supply variations. The method involves modeling of the pull-up and pull-down resistances using approximated drain saturation current device equations for the buffers together with lumped resistive capacitive elements for the interconnects. Compared to circuit simulation results, the analytical model provides more than four orders of magnitude speedup while maintaining an average error of 0.26% with 3.0% standard deviation over the entire range of power supply and circuit parameters variations, making it suitable for timing analysis and optimization. Sami Kirolos, Yehia Massoud, Yehea I. Ismail |
ISCAS | 3 |
| 2008 | Area Optimization for Leakage Reduction and Thermal Stability in Nanometer-Scale TechnologiesabstractTraditionally, the minimum possible area of a very large scale integration (VLSI) layout is considered to be the best for delay and power minimization due to decreased interconnect capacitance. This paper, however, shows that the use of minimum area does not result in minimum power and/or delay in nanometer-scale technologies due to thermal effects and, in some cases, may cause thermal runaway. A methodology using area as a design parameter to reduce the leakage power and prevent thermal runaway is presented. A 16-bit adder example in 70-nm technology shows total power savings of 17% with 15% increase in area and no increase in delay. The power savings using this technique are expected to increase in future technologies. Ja Chun Ku, Yehea I. Ismail |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2008 | Accurate Estimation of SRAM Dynamic StabilityabstractIn this paper, an accurate approach for estimating SRAM dynamic stability is proposed. The conventional methods of SRAM stability estimation suffer from two major drawbacks: 1) using static failure criteria, such as static noise margin (SNM), which does not capture the transient and dynamic behavior of SRAM operation and 2) using quasi-Monte Carlo simulation, which approximates the failure distribution, resulting in large errors at the tails where the desired failure probabilities exist. These drawbacks are eliminated by employing a new distribution-independent, most-probable-failure-point search technique for accurate probability calculation along with accurate simulation-based dynamic failure criteria. Compared to previously published techniques, the proposed technique offers orders of magnitude improvement in accuracy. Furthermore, the proposed technique enables the correct evaluation of stability in real operation conditions and for different dynamic circuit techniques, such as dynamic write-back, where the conventional methods are not applicable. D. E. Khalil, Muhammad M. Khellah, Nam-Sung Kim, Yehea I. Ismail, Tanay Karnik, Vivek De |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2007 | NostraXtalk: a predictive framework for accurate static timing analysis in udsm vlsi circuitsabstractThis paper presents a predictive framework for accurate static timing analysis in UDSM VLSI circuits. As technology scales to smaller dimensions, coupling capacitances are becoming a critical factor in signal integrity analysis. Coupling capacitances contribute to the noise and play a seminalrole in determining the timing windows of a circuit. Accuratean alysis of coupling effects is indispensable for meaning fulstatic timing and signal integrity analysis. Our proposed framework presents a Directed Search technique to calculate accurate coupling effects. We performed experiments on theISCAS'85 benchmarks and present the accuracy improvement up-to 45.5% compared to existing approaches. We also show that our framework decreased cell delay look-uptable accesses up-to 64.8%. Our results present the coupling effect on static timing analysis. Debasish Das, Ahmed Shebaita, Yehea I. Ismail, Hai Zhou 0001, Kip Killpack |
ACM Great Lakes Symposium on VLSI | 3 |
| 2007 | Multi-layer interconnect performance corners for variation-aware timing analysisabstractParasitic interconnect corner methods are known to be inaccurate. This paper explains the sources of their errors and shows that errors in excess of 22% can occur in the predicted corner delays of a multi-layer stage in the presence of process variations. It is shown that exhaustive corner search methods are infeasible in practice as they have an exponential complexity in terms of required SPICE simulations with respect to the number of layers a stage is routed through. This exponential complexity is reduced to a linear one with a new simulation-based search method with the aid of stage delay properties. The ideas behind the simulation-based methodology are shown to be expandable to an analytical-based multi-layer performance corner location methodology. The simulated best/worst case delays based on these analytical corners produce errors below 4% as compared to the exhaustive search simulation based method. Frank Huebbers, Ali Dasdan, Yehea I. Ismail |
ICCAD | 3 |
| 2007 | A self-adjusting clock tree architecture to cope with temperature variationsabstractEnsuring resilience against environmental variations is becoming one of the great challenges of chip design. In this paper, we propose a self adjusting clock tree architecture, SACTA, to improve chip performance and reliability in the presence of on-chip temperature variations. SACTA performs temperature dependent dynamic clock skew scheduling to prevent timing violations in a pipelined circuit. We present an automatic temperature adjustable skew buffer design, which enables the adaptive feature of SACTA. Furthermore, we propose an efficient and general optimization framework to determine the configuration of these special delay elements. Experimental results show that a pipeline supported by SACTA is able to prevent thermal induced timing violations within a significantly larger range of operating temperatures (enhancing the violation-free range by as much as 45°C). Jieyi Long, Ja Chun Ku, Seda Ogrenci Memik, Yehea I. Ismail |
ICCAD | 4 |
| 2007 | Including inductance in static timing analysisabstractIn this paper analytical expressions are derived for effective load capacitances of RLC interconnects to accurately estimate both the propagation delay and transition time at the output of a CMOS gate. The new effective capacitance calculation technique poses no extra complexity as compared to the RC based approaches but can accommodate inductance. These new expressions are derived based on a generalized driving point admittance. The generalized driving point admittance takes inductance into consideration and hence accounts for the inductive shielding that in some cases can even exceed the resistive shielding in current technologies. Another improvement in the new effective capacitance calculation method is the utilization of a more general waveform shape that accounts for the non-monotonic behavior due to inductance effects. It is shown throughout the paper that two effective capacitances are required for accurate estimation of the propagation delay and rise time with an RLC interconnect load. Simulation results show that the error in propagation delays and rise times when neglecting inductance can be over 60% as compared to an RLC model in realistic interconnects. On the other hand, simulations show that the propagation delay and rise time maximum errors associated with the proposed approach are less than 10% as compared to SPICE. Ahmed Shebaita, Dusan Petranovic, Yehea I. Ismail |
ICCAD | 3 |
| 2007 | Approximate Frequency Response Models for RLC Power GridsabstractLarge power grid RLC models are complex and computationally expensive. This work explores the efficient modeling of power grid at early stages of the design with reduced complexity. This paper presents a simplified model of the power grid circuit based on the assumption of uniform load current distribution and equipotential nodes approximation. In addition, based on the moment matching technique an approximate analytical model is derived. The presented models provide huge reduction in complexity without scarifying much accuracy in the trans-impedance frequency response compared to the full RLC power grid model. This is desirable in the early phases of design where only initial estimates of load currents are available. The analytical model presented can be easily used for power grid optimization and trade-offs. DiaaEldin Khalil, Yehea I. Ismail |
ISCAS | 2 |
| 2007 | Attaining Thermal Integrity in Nanometer ChipsabstractAs technology moves into the nanometer era, undesirable trends such as increasing power density, leakage power, and temperature variation within a chip have made thermal effects emerge as a major bottleneck for further technology scaling. Thermal effects are no longer just considered as a reliability issue, but it has also become a fundamentally important and comprehensive problem that includes timing and power issues as well. This paper first overviews the impact of thermal effects on power and performance. Two thermal-aware design techniques, area optimization and low-power cache design, are briefly described. The paper states that there is still plenty of room for further improvement in the area of thermal-aware design. Ja Chun Ku, Yehea I. Ismail |
ISCAS | 2 |
| 2007 | A Compact and Accurate Temperature-Dependent Model for CMOS Circuit DelayabstractWith ever increasing power density and temperature variations within chips, it is very important to correctly model temperature effects on the devices in a compact way. In this paper, it is first shown that the temperature dependencies of the mobility and the saturation velocity need to be treated separately in modeling the current with temperature effects. Then, a new compact temperature-dependent model is presented for the on-current and transient behavior of a CMOS inverter based on the alpha-power law. The proposed model is shown to have an excellent agreement with BSIM3. Ja Chun Ku, Yehea I. Ismail |
ISCAS | 2 |
| 2007 | Variable Threshold Voltage Design Scheme for CMOS Tapered BuffersabstractThis paper proposes a low power, low delay design for CMOS tapered buffers. A slight increase in the threshold voltage is shown to have an exponential effect in reducing the total power dissipation. The corresponding increase in the propagation delay is compensated for by increasing the number of buffer stages such that there is still an overall significant reduction in the total power dissipation. As compared to the constant threshold voltage design based on a cost function of PT2, the proposed scheme can lead to either a power dissipation reduction of about 70% while maintaining the same delay, or up to 30% in power dissipation with 10% propagation delay reduction, respectively. Closed form expressions that give the optimum threshold voltage and number of stages are presented Ahmed Shebaita, Yehea I. Ismail |
ISCAS | 2 |
| 2007 | Thermal-aware methodology for repeater insertion in low-power VLSI circuitsabstractIn this paper, the impact of thermal effects on low-power repeater insertion methodology is studied. An analytical methodology for thermal-aware repeater insertion that includes the electrothermal coupling between power, delay, and temperature is presented, and simulation results with global interconnect repeaters are discussed for 90nm and 65nm technology. Simulation results show that the proposed thermal-aware methodology can save 17.5% more power consumed by the repeaters compared to a thermal-unaware methodology for a given allowed delay penalty. In addition, the proposed methodology also results in a lower chip temperature, and thus, extra leakage power savings from other logic blocks. Ja Chun Ku, Yehea I. Ismail |
ISLPED | 2 |
| 2007 | Modeling and Characterizing Power Variability in Multicore ArchitecturesabstractParameter variation due to manufacturing error is an unavoidable consequence of technology scaling in future generations. The impact of random variation in physical factors such as gate length and interconnect spacing have a profound impact on not only performance of chips, but also their power behavior. While circuit-level techniques such as adaptive body-biasing can help to mitigate mal-fabricated chips, they cannot completely alleviate severe within die variations forecasted for near future designs. Despite the large impact that power variability have on future designs, there is a lack of published work that examines architectural implications of this phenomenon. In this work, we develop architecture level models that model power variability due to manufacturing error and examine its influence on multicore designs. We introduce VariPower, a tool for modeling power variability based on an microarchitectural description and floorplan of a chip. In particular, our models are based on layout level SPICE simulations and project power variability for different microarchitectural blocks using statistical analysis. Using VariPower: (1) we characterize power variability for multicore processors, (2) explore application sensitivity to power variability, and (3) examine clustering techniques that can appropriately classify groups of processors and chips that have similar variability characteristics Frank Huebbers, Russ Joseph, Yehea I. Ismail |
ISPASS | 4 |
| 2007 | Variable latency caches for nanoscale processorabstractVariability is one of the important issues in nanoscale processors. Due to increasing importance of interconnect structures in submicron technologies, the physical location and phenomena such as coupling have an increasing impact on the latency of operations. Therefore, traditional view of rigid access latencies to components wil result in suboptimal architectures. In this paper, we devise a cache architecture with variable access latency. Particularly, we a) develop a non-uniform access level 1 data-cache, b) study the impact of coupling and physical location on level 1 data cache access latencies, and c) develop and study an architecture where the variable latency cache can be accessed while the rest of the pipeline remains synchronous. To find the access latency with different input address transitions and environmental conditions, we first build a SPICE model at a 45nm technology for a cache similar to that of the level 1 data cache of the Intel Prescott architecture. Motivated by the large difference between the worst and best case latencies and the shape of the distribution curve, we change the cache architecture to allow variable latency accesses. Since the latency of the cache is not known at the time of instruction scheduling, we also modify the functional units with the addition of special queues that will temporarily store the dependent instructions and allow the data to be forwarded from the cache to the functional units correctly. Simulations based on SPEC2000 benchmarks show that our variable access latency cache structure can reduce the execution time by as much as 19.4% and 10.7% on average compared to a conventional cache architecture. Serkan Ozdemir, Arindam Mallik, Ja Chun Ku, Gokhan Memik, Yehea I. Ismail |
SC | 5 |
| 2007 | On the Scaling of Temperature-Dependent EffectsabstractWith ever increasing power density and temperature variations within chips, it is very important to correctly model temperature effects on the devices in a compact way and to predict their scaling. In this paper, it is first shown that the temperature dependences of the mobility and the saturation velocity need to be treated separately in modeling the current with the temperature effects. A new compact temperature- dependent model for the ON-current is presented based on the alpha-power law and is verified with BSIM3. Then, the scaling of the ON-current temperature dependence is discussed. It is also shown in this paper that the temperature effects will have an increasing impact on repeater-insertion methodology. Furthermore, the temperature-dependence scaling of the leakage current is analyzed. It is shown that its temperature dependence decreases with technology scaling, but temperature-aware power- reduction techniques will actually save larger fraction of the total power due to the increasing dominance of leakage power. Ja Chun Ku, Yehea I. Ismail |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | Thermal-Aware Methodology for Repeater Insertion in Low-Power VLSI CircuitsabstractIn this paper, the impact of thermal effects on low-power repeater insertion methodology is studied. An analytical methodology for thermal-aware repeater insertion that includes the electrothermal coupling between power, delay, and temperature is presented, and simulation results with global interconnect repeaters are discussed for 90- and 65-nm technology. Simulation results show that the proposed thermal-aware methodology can save 17.5% more power consumed by the repeaters compared to a thermal-unaware methodology for a given allowed delay penalty. In addition, the proposed methodology also results in a lower chip temperature, and thus, extra leakage power savings from other logic blocks. Ja Chun Ku, Yehea I. Ismail |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2007 | Thermal Management of On-Chip Caches Through Power Density MinimizationabstractVarious architectural power reduction techniques have been proposed for on-chip caches in the last decade. In this paper, we first show that these power reduction techniques can be suboptimal when thermal effects are considered. Then, we propose a thermal-aware cache power-down technique that minimizes the power density of the active parts by turning off alternating rows of memory cells instead of entire banks. The decrease in the power density lowers the temperature, which then exponentially reduces the leakage. Thus, leakage power of the active parts is reduced in addition to the power eliminated from the parts that are turned off. Simulations based on SPEC2000, NetBench, and MediaBench applications in a 70-nm technology show that the proposed thermal-aware architecture can reduce the total energy consumption by 53% compared to a conventional cache, and 14% compared to a cache architecture with thermal-unaware power reduction scheme. Second, we show a block permutation scheme that can be used during the design of the caches to maximize the distance between blocks with consecutive addresses. Because of spatial locality, blocks with consecutive addresses are likely to be accessed within a short time interval. By maximizing the distance between such blocks, we minimize the power density of the hot spots in the cache, and hence reduce the peak temperature. This, in return, results in an average leakage power reduction of 8.7% compared to a conventional cache without affecting the dynamic power and the latency. Overall, both of our architectures add no extra run-time penalty compared to the thermal-unaware power reduction schemes, yet they result in a significant reduction in the total energy consumption of a cache Ja Chun Ku, Serkan Ozdemir, Gokhan Memik, Yehea I. Ismail |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2006 | Area optimization for leakage reduction and thermal stability in nanometer scale technologiesabstractTraditionally, minimum possible area of a VLSI layout is considered the best for delay and power minimization due to decreased interconnect capacitance. This paper shows however that the use of minimum area does not result in the minimum power and/or delay in nanometer scale technologies due to thermal effects, and in some cases, may result in thermal runaway. A methodology using area as a design parameter to reduce the leakage power, and prevent thermal runaway is presented. A 16-bit adder example in a 70nm technology shows a total power savings of 17% with 15% increase in area, and no increase in delay. The power savings using this technique are expected to increase in future technologies. Ja Chun Ku, Yehea I. Ismail |
ASP-DAC | 2 |
| 2006 | Computation of accurate interconnect process parameter values for performance corners under process variationsabstractThis paper introduces a fast analytical model for determining accurate parasitic values for best- and worst-case delays of a stage under interconnect process variations. The inputs to the model are the nominal values for each interconnect and device parameter and the amount of variation in each interconnect parameter. The outputs of the model are the interconnect parameter dimensions within the range of process variation that yield the best- and worst-case delay of a stage. Simulations show that our model accurately predicts the performance corners of a stage while those predicted by traditional best/worst-case analysis methodologies can have an error of up to 28.42%. Frank Huebbers, Ali Dasdan, Yehea I. Ismail |
DAC | 3 |
| 2006 | Power density minimization for highly-associative caches in embedded processorsabstractCaches are essential components in embedded processors, taking up a significant fraction of the chip area and power. As a result of the relatively large size and infrequent activity, leakage power of caches is becoming an important problem. There exist a number of power density minimization schemes that distribute the activity evenly among computational entities, thereby lowering the temperature to reduce the leakage power. In this paper, we first present various power density minimization schemes for highly-associative caches in embedded processors via access distribution. It is then suggested that they should be used in conjunction with other power-down techniques to be more effective. We show that conventional power-down techniques for on-chip caches can be suboptimal if thermal effects are ignored, and propose a thermal-aware power-down technique that minimizes power density of the active parts. Simulations based on MediaBench, NetBench, and MiBench applications in a 70nm technology show that the proposed thermal-aware schemes can improve leakage power savings of a conventional power-down technique by 8.5% on average, and up to 23%. Ja Chun Ku, Serkan Ozdemir, Gokhan Memik, Yehea I. Ismail |
ACM Great Lakes Symposium on VLSI | 4 |
| 2006 | Importance of volume discretization of single and coupled interconnectsabstractThis paper presents figures of merit and error formulae to determine which interconnects require volume discretization in the GHZ range. Most of the previous work focused mainly on efficient modeling of volume discretized interconnects using several integration and reduction techniques. However, little work has been done to characterize when using the simple DC model has an impact on critical circuit metrics such as delay, impedance ...etc. Most of the previous work simply assumes that when skin depth becomes smaller than the wire cross section dimensions, volume discretization becomes essential. However, careful analysis in this paper shows that this assumption is invalid and a figure of merit is derived to characterize when volume discretization of single and coupled wires is required. This derived figure of merit is shown to depend solely on the interconnect dimensions and spacing and is independent of the type of the materials used or technology scaling. Ahmed Shebaita, Dusan Petranovic, Yehea I. Ismail |
ICCAD | 3 |
| 2006 | A timing dependent power estimation framework considering couplingabstractIn this paper, we propose a timing dependent dynamic power estimation framework that considers the impact of coupling and glitches. We show that relative switching activities and times of coupled nets significantly affect dynamic power consumption, and neither should be ignored during power estimation. To capture the timing dependence, an approach to efficient representation and propagation of switching-window distributions through a circuit, considering coupling induced delay variations, is developed. Based on the propagated switchingwindow distributions, power consumption in charging or discharging coupling capacitances is calculated, and accounted for in the total power. Experimental results for the ISCAS'85 benchmarks demonstrate that ignoring the impact of timing dependent coupling on power can cause up to 59% error in coupling power estimation (up to 25% error in total power estimation). Debjit Sinha, DiaaEldin Khalil, Yehea I. Ismail, Hai Zhou 0001 |
ICCAD | 3 |
| 2006 | FA-STAC: A Framework for Fast and Accurate Static Timing Analysis with CouplingabstractThis paper presents a framework for fast and accurate static timing analysis considering coupling. With technology scaling to smaller dimensions, the impact of coupling induced delay variations can no longer be ignored. Timing analysis considering coupling is iterative, and can have considerably larger run-times than a single pass approach. We propose a novel and accurate coupling delay model, and present techniques to increase the convergence rate of timing analysis when complex coupling models are employed. Experimental results obtained for the ISCAS benchmarks show promising accuracy improvements using our coupling model while an efficient iteration scheme shows significant speedup (up to 62.1%) in comparison to traditional approaches. Debasish Das, Ahmed Shebaita, Hai Zhou 0001, Yehea I. Ismail, Kip Killpack |
ICCD | 4 |
| 2006 | Reducing the data switching activity of serialized datastreamsabstractOn-chip serial link buses have been previously proposed as a strong solution to reduce the complexity and/or the energy dissipation of on-chip interconnect fabrics. However, it was noticed that serializing m-bits on a single interconnect (serial-link) increases the overall data switching activity. This paper presents a quantitative analysis of the switching activity of serial links, and provides closed form expressions for the average activity factors. Two transition encoding schemes, to reduce the activity factor of serial links, are discussed and analyzed. The impact of the encoding schemes on the MCF between neighboring interconnects is also discussed. The analysis shows that both of the schemes provide significant reduction in the average activity factor and energy dissipation reduction, but each in a different range of input activity factors. The two encoding bus schemes were modeled in a 70nm CMOS technology, and compared to an unencoded serial link bus and a parallel line bus. Simulation results show that the transition encoded bus schemes reduce the overall energy dissipation of the unencoded serial link bus by up to 96% Maged Ghoneima, Yehea I. Ismail, Muhammad M. Khellah, Vivek De |
ISCAS | 2 |
| 2006 | Optimum sizing of power grids for IR dropabstractThis paper explores the optimum design of the power grid. A simplified model of the power grid, based on the assumption of uniform current distribution and some approximations, is presented. This model is shown to give the IR drop at the grid center with error less than 0.1% compared to full simulation. In addition, the optimum sizing of power grid lines for minimum IR drop is derived, based on a general version of the model. Applying the optimum sizing, a uniform grid has the minimum IR drop, for uniform current profile. However, for real chip examples where current profiles are non-uniform, the optimum sizing results in 14% reduction in the IR drop at the center compared to uniform sizing DiaaEldin Khalil, Yehea I. Ismail |
ISCAS | 2 |
| 2006 | Time-borrowing multi-cycle on-chip interconnects for delay variation toleranceabstractInsertion of time-borrowing (TB) flip-flops in multi-cycle repeater-based on-chip interconnects enables significant improvements in mean performance and energy by averaging systematic and random within-die (WID) delay variations across multiple interconnect segments. A statistically-based analytical model is derived to design a TB N-cycle interconnect with optimal delay variation tolerance. The model elucidates the dependency of the transparency window required to achieve data delay averaging on the delay variation mismatch between interconnect segments. Statistical circuit simulations and analyses in a 65nm process technology demonstrate that TB multi-cycle interconnects enable a 4-6% mean maximum clock frequency (FMAX) improvement and a corresponding 10% average energy savings over optimally designed multi-cycle interconnects with conventional master-slave flip-flops. The maximum mean FMAX benefit ranges from 4.0-7.5%, corresponding to approximately a bin-split shift in the FMAX distribution. For 1.41X larger WID delay variations, the maximum mean FMAX gain rises to 5-10%. Keith A. Bowman, James W. Tschanz, Muhammad M. Khellah, Maged Ghoneima, Yehea I. Ismail, Vivek De |
ISLPED | 5 |
| 2006 | Formal derivation of optimal active shielding for low-power on-chip busesabstractPassive shielding has been used to reduce the capacitive coupling effects of adjacent bus lines by inserting passive ground or power lines (shields) between them. Active shielding is another shielding technique in which the shield is allowed to switch depending on the switching pattern of its adjacent bus lines. This paper formally derives the optimal active shielding logic function for minimum power dissipation. It is also shown that this optimal active shielding architecture depends on the ratio of coupling to ground capacitance (/spl gamma/=C/sub c//C/sub g/). Optimal active shielding is shown to provide up to 25% reduction in bus power dissipation compared to conventional passive shielding. A suboptimal active shielding architecture with simpler hardware is also proposed. Theoretically, using the suboptimal shielding architecture leads to less than 6% bus power penalty compared to the optimal active shielding logic circuit. However, due to the simpler shield encoding circuitry, simulation results show that the suboptimal active shielding architecture leads to higher overall energy savings compared to the optimal active shielding architectures. Maged Ghoneima, Yehea I. Ismail, Muhammad M. Khellah, James W. Tschanz, Vivek De |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | Realistic scalability of noise in dynamic circuitsabstractThe usage of noise-sensitive dynamic circuits has become commonplace due to speed and area requirements, making the noise issue even more prominent. This paper focuses on the trends of coupling and its effects on dynamic circuits. It presents closed form analytical solutions for noise, as well as noise tolerance metrics for dynamic circuits. These solutions are within 5% of dynamic simulations. It is shown that not all scaling trends are negative for noise, and that the scaling down of supply voltage and increasing frequency, help improve certain aspects of the noise immunity of dynamic circuits. Most of the works treated the noise immunity and the noise content separately. This paper introduces an analysis of noise scalability by looking at the noise immunity and the noise content simultaneously Masud H. Chowdhury, Yehea I. Ismail |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Piece-wise approximations of RLCK circuit responses using moment matchingabstractCapturing RLCK circuit responses accurately with existing model order reduction (MOR) techniques is very expensive. Direct metrics for fast analysis of RC circuits exist but there is no such technique for RLCK circuits. This paper introduces a new family of MOR techniques based on piece-wise functions to capture RLCK circuit responses accurately using only four or five moments. The time-domain response is approximated using a piece-wise function whose pieces are simple polynomials. The proposed method is fast and guaranteed stable and it avoids the calculation of poles and residues associated with existing model order reduction techniques. Results for many different industrial netlists indicate that delay and transition time can be captured within 5% error using only four moments. To the authors' knowledge, there is no existing method that can extract as much information about RLCK circuits with only four or five moments. Chirayu S. Amin, Yehea I. Ismail, Florentin Dartu |
DAC | 2 |
| 2005 | Statistical static timing analysis: how simple can we get?abstractWith an increasing trend in the variation of the primary parameters affecting circuit performance, the need for statistical static timing analysis (SSTA) has been firmly established in the last few years. While it is generally accepted that a timing analysis tool should handle parameter variations, the benefits of advanced SSTA algorithms are still questioned by the designer community because of their significant impact on complexity of STA flows. In this paper, we present convincing evidence that a path-based SSTA approach implemented as a post-processing step captures the effect of parameter variations on circuit performance fairly accurately. On a microprocessor block implemented in 90nm technology, the error in estimating the standard deviation of the timing margin at the inputs of sequential elements is at most 0.066 FO4 delays, which translates in to only 0.31% of worst case path delay. Chirayu S. Amin, Noel Menezes, Kip Killpack, Florentin Dartu, Umakanta Choudhury, Nagib Hakim, Yehea I. Ismail |
DAC | 7 |
| 2005 | Engineering Over-Clocking: Reliability-Performance Trade-Offs for High-Performance Register FilesabstractRegister files are in the critical path of most high-performance processors and their latency is one of the most important factors that limit their size. Our goal is to develop error correction mechanisms at the architecture level. Utilizing this increased robustness, the clock frequencies of the circuits are pushed beyond the point of allowing full voltage swing. This increases the errors observed due to noise and other external factors. The resulting errors are then corrected through the error correction mechanisms. We first develop a realistic model for error probability in register files for a given clock frequency. Then, we present the overall architecture, which allows the error detection computation to be overlapped with other computation in the pipeline. We develop novel techniques that utilize the fact that at a given instance many physical registers are not used in superscalar processors. These underutilized registers are used to store the values of active registers. Our simulation results show that for a fixed architecture the access times to the registers can be reduced by as much as 80% while increasing the number of execution cycles by 0.12%. On the other hand, by reducing the register file access pipeline stages by 75%, the average number of execution cycles of SPEC applications can be reduced by 11.5%. Gokhan Memik, Masud H. Chowdhury, Arindam Mallik, Yehea I. Ismail |
DSN | 4 |
| 2005 | Physical limitations on the bit-rate of on-chip interconnectsabstractIt is shown in this paper that contrary to the common conception, widening an on-chip interconnect wire will not cause the wire to behave as a lossless line. It is shown that the energy losses actually increase when widening a wire. The damping factor of the line, which determines the maximum bit rate that can be transmitted on a wire, also does not go to zero for wide wires as in lossless lines. This fact results in a serious physical limitation on the maximum bit rate that can be transmitted on a wire. It is shown that by increasing the cross-sectional width of an interconnect wire, the damping factor will saturate to a certain limit. An expression for this minimum damping factor for very wide wires is derived from fundamental physical constants and used to introduce an expression for the maximum bit rate that can be reached by widening an interconnect wire for a certain technology. It is shown that for current technology numbers, the maximum bit rate on a wire can be quite limiting for longer wires and that scaling trends will even decrease more this maximum bit rate, posing a serious limitation. Noha H. Mahmoud, Maged Ghoneima, Yehea I. Ismail |
ACM Great Lakes Symposium on VLSI | 3 |
| 2005 | Serial-link bus: a low-power on-chip bus architectureabstractAs technology scales, the shrinking wire width increases the interconnect resistivity, while the decreasing interconnect spacing significantly increases the coupling capacitance. This paper proposes reducing the number of bus lines of the conventional parallel-line bus CB architecture by multiplexing each m-bits onto a single line. This bus architecture, the serial-link bus SLB, transforms an n-bit conventional parallel-line bus into an n/m-line (serial-link) bus. The advantage of serial-link buses is that they have fewer lines, and if the bus width is kept the same, serial- link buses will have larger line width and spacing. Increasing the line width has a twofold reduction effect on the line resistance, as the resistivity of sub-100 nm wires significantly drops as the line width increases. Also, increasing the line width and spacing reduces the coupling capacitance between adjacent lines, but increases the line-to-ground capacitance. Thus, an optimum degree of multiplexing m exists that minimizes the bus energy dissipation and maximizes the bus throughput per-unit area. The optimum degree of multiplexing for maximum throughput-per- unit-area and for minimum energy dissipation for the 25-130 nm technologies was determined in this paper. HSPICE simulations show that; for the same throughput-per-unit-area as conventional parallel-line buses, the serial-link bus architecture reduces the energy dissipation by up to 31.42% for a 64-bit bus implemented in an intermediate metal layer of a 50 nm technology and a reduction of 52.7% is projected for the 25 nm technology. Maged Ghoneima, Yehea I. Ismail, Muhammad M. Khellah, James W. Tschanz, Vivek De |
ICCAD | 2 |
| 2005 | Expanding the frequency range of AWE via time shiftingabstractThe new technique of time shifted moment matching (TSMM) is introduced in this paper. The TSMM technique performs moment matching (for expansion around s = 0) on a time-shifted version of the original signal. As compared to other well-known techniques (such as AWE by Pillage and Rohrer, 1990), TSMM offers distinct advantages. The 50% delay and rise time are determined with much more accuracy for a given approximation order. Moreover, the solutions have significantly improved accuracy as compared to AWE, especially for moderate to highly inductive signals. TSMM is able to achieve the approximation capability of PVL (Feldmann and Freund, 1995) and PRIMA (Odabasioglu et al., 1998) with much lower approximation order. Ahmed M. Shebaita, Chirayu S. Amin, Florentin Dartu, Yehea I. Ismail |
ICCAD | 4 |
| 2005 | A Skewed Repeater Bus Architecture for On-Chip Energy Reduction in MicroprocessorsabstractThis paper proposes a bus architecture called skewed repeater bus (SRB) for reducing on-chip interconnect energy in microprocessors. By introducing relative delay between neighboring bus lines, SRB reduces both average and worst-case coupling capacitance between those lines. SRB is compared to previously published techniques like delayed data bus (DDB) and delayed clock bus (DCB). Simulation results in 65-nm process show that bus energy reduction of 18% is achieved when SRB is applied to a real microprocessor example, versus 11% and 7% only for DDB and DCB; respectively. Muhammad M. Khellah, Maged Ghoneima, James W. Tschanz, Yibin Ye, Nasser A. Kurd, Javed Barkatullah, Srikanth Nimmagadda, Yehea I. Ismail |
ICCD | 8 |
| 2005 | Thermal Management of On-Chip Caches Through Power Density MinimizationabstractVarious architectural power reduction techniques have been proposed for on-chip caches in the last decade. However, these techniques mostly ignore the effects of temperature on the power consumption. In this paper, first we show that these power reduction techniques can be suboptimal when thermal effects are considered. Particularly, we propose a thermal-aware cache power-down technique that minimizes the power density of the active parts by turning off alternating rows of memory cells instead of entire banks. The decrease in the power density lowers the temperature, which in return, reduces the leakage of the active parts. Simulations based on SPEC2000 benchmarks in a 70nm technology show that the proposed thermal-aware architecture can reduce the total energy consumption by 53% compared to a conventional cache, and 14% compared to a cache architecture with thermal-unaware power reduction scheme. Second, we show a block permutation scheme that can be used during the design of caches to maximize the distance between blocks with consecutive addresses. By maximizing the distance between consecutively accessed blocks, we minimize the power density of the hot spots in the cache, and hence reduce the peak temperature. This, in return, results in an average leakage power reduction of 8.7% compared to a conventional cache without affecting the dynamic power and the latency. Overall, both of our architectures add no extra run-time penalty compared to the thermal-unaware power reduction schemes, yet they reduce the total energy consumption of a conventional cache by 53% and 5.6% on average, respectively. Ja Chun Ku, Serkan Ozdemir, Gokhan Memik, Yehea I. Ismail |
MICRO | 4 |
| 2005 | Realizable reduction of interconnect circuits including self and mutual inductancesabstractReduction of an extracted netlist is an important preprocessing step for techniques such as model order reduction (MOR) in the design and analysis of very large scale integration circuits (VLSICs). This work describes a method for realizable reduction of extracted resistance-capacitance-inductance-mutual inductance netlists by node elimination. The method is much faster than MOR techniques and, hence, is appropriate as a preprocessing step. The proposed method eliminates nodes with time constants below a user-specified time constant. By giving the freedom to the user to select a critical point in the spectrum of nodal time constants, this method provides an option to make a tradeoff between accuracy and reduction. The proposed method preserves the dc characteristics and the first two moments at all nodes. It also recognizes and eliminates all the redundant inductances generated by the extraction tools. The proposed method naturally reduces to TICER (Sheehan, 1999) in the absence of any inductances. Chirayu S. Amin, Masud H. Chowdhury, Yehea I. Ismail |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2005 | Weibull-based analytical waveform modelabstractCurrent complimentary metal-oxide-semiconductor technologies are characterized by interconnect lines with increased relative resistance with respect to driver output resistance. Designs generate signal waveshapes that are very difficult to model using a single-parameter model such as the transition time. In this paper, we present a simple and robust two-parameter analytical expression for waveform modeling based on the Weibull cumulative distribution function. The Weibull model accurately captures the variety of waveshapes without introducing significant runtime overhead and produces results with less than 5% error. We also present a fast and simple algorithm to convert waveforms obtained by circuit simulation to the Weibull model. A methodology for characterizing gates for the new model is also presented. Simulation results for many single- and multiple-input gates show errors well below 5%. Our model can be used in a mixed environment where some signals may still be characterized by a single parameter. Chirayu S. Amin, Florentin Dartu, Yehea I. Ismail |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2005 | Optimum positioning of interleaved repeaters in bidirectional busesabstractIt is shown in this paper that the optimum position of interleaved repeaters for minimum delay and noise is not the midpoint as commonly practiced. A closed-form solution for the optimum position has been derived in this paper and verified by simulation. Bidirectional buses with the optimum interleaved repeater position are compared to commonly used bidirectional buses and shown to provide an improvement greater than 50% in the propagation delay and bit-rate per unit area. The area of the induced noise pulse on victim lines is shown to be zero indicating that the aggressor lines are virtually static with the optimum repeater position. The presented optimum repeater positioning also provides lower noise pulse amplitude as well as lower sensitivity of propagation delay and noise pulse peak to segment length variation, compared to commonly used midway repeater positioning. Maged Ghoneima, Yehea I. Ismail |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2004 | Modeling unbuffered latches for timing analysisabstractUnbuffered latches are often used in high-performance designs with custom timing flows. Adding these circuits to a standard library enables improved designs without blowing the library size. We observe a high potential frequency gain (up to 16%) for smaller power consumption. Accurate models for static timing analysis are required to reach a good point on the safety to performance trade-off. We are proposing a complete modeling methodology that can fit in a standard timing analysis flow. An accurate n-model is presented for the input impedance of an unbuffered latch with less than 2% error. We also present a new setup criteria required for these latches. We also show that more advanced waveform models are required to model the output. A Weibull waveform model proves to be effective in this case. Chirayu S. Amin, Florentin Dartu, Yehea I. Ismail |
ICCAD | 3 |
| 2004 | Formal derivation of optimal active shielding for low-power on-chip busesabstractPassive shielding has been used to reduce capacitive coupling effects of adjacent bus lines by inserting passive ground or power lines (shields) between the bus lines. Active shielding is another shielding technique, in which the shield is allowed to switch depending on the switching pattern of its adjacent bus lines. This work formally derives the optimal active shielding logic function for minimum power dissipation. It is also shown that this optimal active shielding architecture depends on the ratio of coupling to ground capacitance (/spl gamma/ = C/sub c//C/sub g/). Optimal active shielding is shown to provide up to 25% reduction in bus power dissipation compared to conventional passive shielding. A sub-optimal active shielding architecture with simpler hardware is also proposed. Simulation results show that using the sub-optimal shielding architecture leads to less than 6% bus power penalty compared to the optimal active shielding logic circuit. Maged Ghoneima, Yehea I. Ismail |
ICCAD | 2 |
| 2004 | Computation of signal threshold crossing times directly from higher order momentsabstractThis work introduces a simple method for calculating the times at which any signal crosses a pre-specified threshold voltage (e.g. 10%, 20%, 50%, etc.) directly from the moments. The method can use higher order moments to asymptotically improve the accuracy of the estimated crossing times. This technique bypasses the steps involved in calculating poles and residues to obtain time-domain information. Once q moments are calculated, only 2q multiplications and (q-I) additions are required to determine any threshold crossing time at a certain node. Moreover, this technique avoids other problems such as pole instability. Several orders of approximations are presented for different threshold crossing times depending on the number of moments involved. For example, the worst case error of a first to a seventh order (single to seven moments) approximation of 50% RC delay is 1650%, 192.26%, 11.31%, 3.37%, 2.57%, 2.56%, and 1.43%, respectively. If the whole waveform is required it can be easily determined by interpolation between different threshold crossing points. The presented technique works for RC circuits for both step and nonstep inputs, including piecewise linear waveforms. Yehea I. Ismail, Chirayu S. Amin |
ICCAD | 1 |
| 2004 | Delayed line bus scheme: a low-power bus scheme for coupled on-chip busesabstractThis paper presents a comprehensive qualitative and analytical analysis of the effect of relative delay on the dissipated energy of coupled lines. Closed form expressions modeling the effect of relative delay on the dissipated energy, and the Miller coupling factor, MCF, are also presented. Skewing the worst switching case is shown to provide up to 50% reduction in energy dissipation. This observation was implemented in a low-power bus scheme, DLBS, which leads to a power reduction of up to 25%. Maged Ghoneima, Yehea I. Ismail |
ISLPED | 2 |
| 2004 | Computation of signal-threshold crossing times directly from higher order momentsabstractThis paper introduces a simple method for calculating the times at which any signal crosses a prespecified threshold voltage (e.g., 10%, 20%, 50%, etc.) directly from the moments. The method can use higher order moments to asymptotically improve the accuracy of the estimated crossing times. This technique bypasses the steps involved in calculating poles and residues to obtain time-domain information. Once q moments are calculated, only 2q, multiplications and (q-1) additions are required to determine any threshold-crossing time at a vermin node. Moreover, this technique avoids other problem such as pole instability. The final outcome of this paper is a set of empirical expressions relating the moments to different threshold-crossing times in analogy to the t/sub d/=-0.693m/sub 1/ formula. The presented methodology can also be used with other user defined forms of empirical expressions relating the moments to different threshold-crossing times. Several orders of approximations an presented for different threshold-crossing times, depending on the number of moments involved. For example, the worst-case error of a first- to seventh-order (single to seven moments) approximation of 50% RC delay is 1650%, 192.26%, 11.31%, 3.37%, 2.57%, 2.56%, and 1.43%, respectively. This technique is very useful to obtain information about certain signal metrics, such as delay and rise time directly without having to compute the whole time domain waveform. In addition, if the whole waveform is required it can be easily determined by interpolation between different threshold-crossing points. The presented technique works for both step and nonstep inputs, including piecewise-linear waveforms. Yehea I. Ismail, Chirayu S. Amin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2004 | Utilizing the effect of relative delay on energy dissipation in low-power on-chip busesabstractThis paper presents an analysis of how the power dissipation of on-chip buses is affected by introducing a relative delay between the switching lines. Relative delay is shown to reduce the dissipated power of oppositely switching lines while causing a power penalty for similarly switching lines. A new low-power bus scheme that uses this effect is proposed and analyzed. As the introduced delay increases, the achieved power reduction increases while decreasing the bus throughput. Thus, a tradeoff between power reduction and throughput is required when selecting the imposed relative delay. The proposed low-power scheme, dynamic delayed line bus (DDL) scheme, led to a power reduction of up to 25%, 33%, and 42% when applied to data, address, and differential buses, respectively. Simple DDL hardware is designed and implemented in a 0.18-/spl mu/m TSMC CMOS technology and applied to a 4500-/spl mu/m long Metal4 bus. Circuit simulation results for different bus widths are presented. Maged Ghoneima, Yehea I. Ismail |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2004 | Modeling skin and proximity effects with reduced realizable RL circuitsabstractOn-chip conductors such as clock- and power-distribution networks require accurately modeling skin and proximity effects. Furthermore, to incorporate skin and proximity effects in the existing generic simulation tools such as SPICE, simple-frequency independent-lumped element-circuit models are needed. A rule based RL circuit model is proposed in this paper that is realizable and predicts skin and proximity effects accurately in the frequency range of interest. With this circuit model, wires are characterized by a few parallel branches of resistors and inductors while proximity effect is captured by mutual inductance between inductors in different RL circuits. Shizhong Mei, Yehea I. Ismail |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | Realizable RLCK circuit crunchingabstractReduction of an extracted netlist is an important pre-processing step for techniques such as model order reduction in the design and analysis of VLSI circuits. This paper describes a method for realizable reduction of extracted RLCK netlists by node elimination. The method is much faster than model order reduction techniques and hence is appropriate as a pre-processing step. The proposed method eliminates nodes with time constants below a user specified time constant. By giving the freedom to the user to select a critical point in the spectrum of nodal time constants, this method provides an option to make a tradeoff between accuracy and reduction. The proposed method preserves the dc characteristics and the first two moments at all nodes. It also recognizes and eliminates all the redundant inductances generated by the extraction tools. The proposed method naturally reduces to TICER [13] in the absence of any inductances. Chirayu S. Amin, Masud H. Chowdhury, Yehea I. Ismail |
DAC | 3 |
| 2003 | Optimum positioning of interleaved repeaters In bidirectional busesabstractIt is shown in this paper that the optimum position of interleaved repeaters for minimum delay and noise is not the midpoint as commonly practiced. A closed form solution for the optimum position has been derived in this paper and verified by simulation. Bi-directional buses with the optimum interleaved repeater position are compared to commonly used bi-directional buses and shown to provide an improvement of as much as 100% in both the propagation delay and bit-rate per unit area. The area of the induced noise pulse on victim lines is shown to be zero indicating that the aggressor lines are virtually static when optimum repeater positioning is used. Maged Ghoneima, Yehea I. Ismail |
DAC | 2 |
| 2003 | Efficient model order reduction including skin effectabstractSkin effect makes interconnect resistance and inductance frequency dependent. This paper addresses the problem of efficiently estimating the signal characteristics of any RLC network when skin effect is significant, which complicates interconnect simulation. In this paper, a new type of moments is defined that simplifies the interconnect simulation, namely, the square root moments. The time to calculate the square root moments is similar to the time to calculate the traditional moments, and the new moments preserve the recursive properties of the traditional moments. Hence, the method introduced here can handle the more complex problem of interconnect simulation with skin effect at almost no overhead compared to constant element interconnect simulation. Using the square root moments, higher order approximations can be reached as compared to traditional moments. Also, the PVL method is modified to implicitly match the square root moments. The simulation results reveal the high accuracy of the proposed methods as well as the apparent variation in the signal characteristics caused by skin effect. Shizhong Mei, Chirayu S. Amin, Yehea I. Ismail |
DAC | 3 |
| 2003 | Weibull Based Analytical Waveform Model
Chirayu S. Amin, Florentin Dartu, Yehea I. Ismail |
ICCAD | 3 |
| 2003 | Improved model-order reduction by using spacial information in momentsabstractThe new concept of multinode moment matching (MMM) is introduced in this paper. The MMM technique simultaneously matches the moments at several nodes of a circuit using explicit moment matching around s=0. As compared to the well known single-point moment matching (SMM) techniques (such as asymptotic waveform evaluation), MMM has several advantages. First, the number of moments required by MMM is significantly lower than SMM for a reduced-order model of the same accuracy, which directly translates into computational efficiency. This higher computational efficiency of MMM as compared to SMM increases with the number of inputs to the circuit. Second, MMM has much better numerical stability as compared to SMM. This characteristic allows MMM to calculate an arbitrarily high-order approximation of a linear system, achieving the required accuracy for systems with complex responses. Finally, MMM is highly suitable for parallel-processing techniques especially for higher order approximations while SMM has to calculate the moments sequentially and cannot be adapted to parallel processing techniques. Yehea I. Ismail |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2002 | Efficient model order reduction via multi-node moment matchingabstractThe new concept of Multi-node Moment Matching (MMM) is introduced in this paper. The MMM technique simultaneously matches the moments at several nodes of a circuit using explicit moment matching around s=0. As compared to the well-known Single-point Moment Matching (SMM) techniques (such as AWE), MMM has several advantages. First, the number of moments required by MMM is significantly lower than SMM for a reduced order model of the same accuracy, which directly translates into computational efficiency. This higher computational efficiency of MMM as compared to SMM increases with the number of inputs to the circuit. Second, MMM has much better numerical stability as compared to SMM. This characteristic allows MMM to calculate an arbitrarily high order approximation of a linear system, achieving the required accuracy for systems with complex responses. Finally, MMM is highly suitable for parallel processing techniques especially for higher order approximations while SMM has to calculate the moments sequentially and cannot be adapted to parallel processing techniques. Yehea I. Ismail |
ICCAD | 1 |
| 2002 | DTT: direct truncation of the transfer function - an alternative tomoment matching for tree structured interconnectabstractA method is introduced to evaluate time domain signals within RLC trees with arbitrary accuracy in response to any input signal. This method depends on finding a low frequency reduced-order transfer function by direct truncation of the exact transfer function at different nodes of an RLC tree. The method is numerically accurate for any order of approximation, which permits approximations to be determined with a large number of poles appropriate for approximating RLC trees with underdamped responses. The method is computationally efficient with a complexity linearly proportional to the number of branches in an RLC tree. A common set of poles is determined that characterizes the responses at all of the nodes of an RLC tree which further enhances the computational efficiency. Stability is guaranteed by the DTT method for low-order approximations with less than five poles. Such low-order approximations are useful for evaluating monotone responses exhibited by RC circuits. Yehea I. Ismail, Eby G. Friedman |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2002 | On-chip inductance cons and prosabstractProvides a high level survey of the increasing effects of on-chip inductance. These effects are classified into desirable and nondesirable effects. Among the undesirable effects of on-chip inductance are higher interconnect coupling noise and substrate coupling, challenges for accurate extraction, the required modifications of the infrastructure of CAD tools, and the inevitably slower CAD tools as compared to RC-based tools. Among the desirable effects is lower power consumption, less need for repeaters, faster signal rise time, and less delay uncertainty. The viability of design methodologies considering on-chip inductance is briefly discussed. Yehea I. Ismail |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2002 | Guest editorial: special issue on on-chip inductance in high-speed integrated circuitsabstractHE appropriate interconnect model has changed several times over the past two decades due to aggressive technology scaling. New, more accurate interconnect models were introduced when parasitic effects that were negligible in earlier technologies, could no longer be ignored. Currently, RC models are used to analyze high resistance nets while capacitive models are used for less resistive interconnects. However, on-chip inductance is becoming increasingly important since integrated circuits now operate at frequencies where the inductive impedance of thick wide wires is comparable to wire resistance and line lengths are long enough, relative to signal rise times, for transmission line behavior to become significant. Furthermore, this trend shows every indication of spreading beyond the relatively few lines it now affects. Operating frequencies that have increased dramatically over the past decade, are expected to maintain the same rate of increase over the next decade approaching 10 GHz by the year 2012. The use of thick upper level metals, large die sizes, and low resistance copper interconnect—already used in many commercial CMOS technologies, will likewise continue. Finally, because large die sizes enable more system integration, the use of long thick wide wires, once largely devoted to global clock distribution networks, will spread to critical data-buses and control lines. This special issue deals with the design and analysis of integrated circuits including parasitic on-chip inductance. It does not deal with intentionally designed structures like the spiral inductors used in RF circuits and LC tank oscillators. As evidence that the subject material is still under investigation, the papers in this special issue do not present a unified approach to modeling parasitic inductance. The most notable difference centers around the use of 2-D (or loop inductance) models versus more general 3-D inductance models, and at the center of this debate is the question of current loops formed by return currents. The 3-D model advocates take the position that the return current paths are fundamentally unknown and present sophisticated analysis methods to cope with the increased complexity of 3-D models. The 2-D advocates insist that the return currents are equal and opposite to the interconnect currents and can therefore, employ simpler models. While this debate is apparent in the second paper which favors 2-D models and the third, fourth, and sixth papers which favor 3-D models, several other papers also imply a preference by presenting analytical methods that are suitable to 2-D models only. The paper summary below highlights when notable, the authors stated or implied preference to 2-D or 3-D models. In the first paper, Ismail presents several analytical methods for including inductive effects in both timing and noise analysis. The analytical models he presents are only applicable to Yehea I. Ismail, Byron Krauter |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2001 | Exploiting the on-chip inductance in high-speed clock distribution networksabstractOn-chip inductance effects can be used to improve the performance of high-speed integrated circuits. Specifically, inductance improves the signal slew rate (the rise time), virtually eliminates short-circuit power consumption and reduces the area of the active devices and repeaters inserted to optimize the performance of long interconnects. These positive effects suggest the development of design strategies that benefit from on-chip inductance. An example of a clock distribution network is presented to illustrate the process in which inductance can be used to improve the performance of high-speed integrated circuits. Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2000 | Sensitivity of interconnect delay to on-chip inductanceabstractInductance extraction has become an important issue in the design of high speed CMOS circuits. Two characteristics of on-chip inductance are discussed in this paper that can significantly simplify the extraction of on-chip inductance, The first characteristic is that the sensitivity of a signal waveform to errors in the inductance values is low, particularly the propagation delay and the rise time. It is quantitatively shown in this paper that the error in the propagation delay and rise time is below 9.4% and 5.9%, respectively, assuming a 30% relative error in the extracted inductance. If an RC model is used for the same example, the corresponding errors are 51% and 71%, respectively, The second characteristic is that the magnitude of the on-chip inductance is a slowly varying function of the width of a wire and the geometry of the surrounding wires. These two characteristics can be exploited by using simplified techniques that permit approximate and sufficiently accurate values of the on-chip inductance to be determined with high computational efficiency. Yehea I. Ismail, Eby G. Friedman |
ISCAS | 1 |
| 2000 | Equivalent Elmore delay for RLC treesabstractClosed-form solutions for the 50% delay, rise time, overshoots, and settling time of signals in an RLC tree are presented. These solutions have the same accuracy characteristics of the Elmore delay for RC trees and preserves the simplicity and recursive characteristics of the Elmore delay. Specifically, the complexity of calculating the time domain responses at all the nodes of an RLC tree is linearly proportional to the number of branches in the tree and the solutions are always stable. The closed-form expressions introduced here consider all damping conditions of an RLC circuit including the underdamped response, which is not considered by the Elmore delay due to the nonmonotone nature of the response. The continuous analytical nature of the solutions makes these expressions suitable for design methodologies and optimization techniques. Also, the solutions have significantly improved accuracy as compared to the Elmore delay for an overdamped response. The solutions introduced here for RLC trees can be practically used for the same purposes that the Elmore delay is used for RC trees. Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2000 | Effects of inductance on the propagation delay and repeater insertion in VLSI circuitsabstractA closed-form expression for the propagation delay of a CMOS gate driving a distributed RLC line is introduced that is within 5% of dynamic circuit simulations for a wide range of RLC loads. It is shown that the error in the propagation delay if inductance is neglected and the interconnect is treated as a distributed RC line can be over 35% for current on-chip interconnect. It is also shown that the traditional quadratic dependence of the propagation delay on the length of the interconnect for RC lines approaches a linear dependence as inductance effects increase. On-chip inductance is therefore expected to have a profound effect on traditional high-performance integrated circuit (IC) design methodologies. The closed-form delay model is applied to the problem of repeater insertion in RLC interconnect. Closed-form solutions are presented for inserting repeaters into RLC lines that are highly accurate with respect to numerical solutions. RC models can create errors of up to 30% in the total propagation delay of a repeater system as compared to the optimal delay if inductance is considered. The error between the RC and RLC models increases as the gate parasitic impedances decrease with technology scaling. Thus, the importance of inductance in high-performance very large scale integration (VLSI) design methodologies will increase as technologies scale. Yehea I. Ismail, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1999 | Effects of Inductance on the Propagation Delay and Repeater Insertion in VLSI CircuitsabstractA closed form expression for the propagation delay of a CMOS gate driving a distributed RLC line is introduced that is within 5% of dynamic circuit simulations for a wide range of RLC loads.It is shown that the traditional quadratic dependence of the propagation delay on the length of an RC line approaches a linear dependence as inductance effects increase.The closed form delay model is applied to the problem of repeater insertion in RLC interconnect.Closed form solutions are presented for inserting repeaters into RLC lines that are highly accurate with respect to numerical solutions.An RC model as compared to an RLC model creates errors of up to 30% in the total propagation delay of a repeater system.Considering inductance in repeater insertion is also shown to significantly save repeater area and power consumption.The error between the RC and RLC models increases as the gate parasitic impedances decrease which is consistent with technology scaling trends.Thus, the importance of inductance in high performance VLSI design methodologies will increase as technologies scale. Yehea I. Ismail, Eby G. Friedman |
DAC | 1 |
| 1999 | Equivalent Elmore Delay for RLC TreesabstractArticle Free Access Share on Equivalent Elmore delay for RLC trees Authors: Yehea I. Ismail Department of Electrical and Computer Engineering, University of Rochester, Rochester, New York Department of Electrical and Computer Engineering, University of Rochester, Rochester, New YorkView Profile , Eby G. Friedman Department of Electrical and Computer Engineering, University of Rochester, Rochester, New York Department of Electrical and Computer Engineering, University of Rochester, Rochester, New YorkView Profile , Jose L. Neves IBM Microelectronics, 1580 Route 52, East Fishkill, New York IBM Microelectronics, 1580 Route 52, East Fishkill, New YorkView Profile Authors Info & Claims DAC '99: Proceedings of the 36th annual ACM/IEEE Design Automation ConferenceJune 1999 Pages 715–720https://doi.org/10.1145/309847.310041Published:01 June 1999Publication History 7citation1,519DownloadsMetricsTotal Citations7Total Downloads1,519Last 12 Months76Last 6 weeks24 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
DAC | 1 |
| 1999 | Inductance Effects in RLC TreesabstractA closed form solution for characterizing voltage-based signals in an RLC tree is presented. This closed form solution is used to derive figures of merit to characterize the effects of inductance at a specific node in an RLC tree. The effective damping factor of the signal at a specific node in an RLC tree is shown to be a useful figure of merit, As the effective damping factor of a signal increases, an RC model is sufficiently accurate to characterize that waveform. The rise time of the input signal driving an RLC tree is another factor characterizing the importance of inductance. As the rise time of the input signal becomes much larger than the effective LC time constant at a specific node within an RLC tree, the signal at this node does not exhibit the effects of inductance. Evidence is provided showing that using a single line analysis to determine the importance of including inductance to characterize a tree structured interconnect line is invalid in many cases and can lead to erroneous conclusions. Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
Great Lakes Symposium on VLSI | 1 |
| 1999 | Repeater insertion in tree structured inductive interconnectabstractThe effects of inductance on repeater insertion in RLC trees is the focus of the paper. An algorithm is introduced to insert and size repeaters within an RLC tree to optimize a variety of possible cost functions such as minimizing the maximum path delay, the skew between branches, or a combination of area, power, and delay. The algorithm has a complexity proportional to the square of the number of possible repeater positions, permitting a repeater solution to be chosen that is close to the global minimum. The repeater insertion algorithm is used to insert repeaters within several copper based interconnect trees to minimize the maximum path delay based on both an RC model and an RLC model. The two buffering solutions are compared using the AS/X dynamic circuit simulator. It is shown that as inductance effects increase, the area and power consumed by the inserted repeaters to minimize the path delays of an RLC tree decreases. By including inductance in the repeater insertion methodology, the interconnect is modeled more accurately as compared to an RC model, permitting average savings in area, power, and delay of 40.8%, 15.6%, and 6.7%, respectively, for a variety of copper based interconnect trees from a 0.25 /spl mu/m CMOS technology. The average savings in area, power, and delay increases to 62.2%, 57.2% and 9.4%, respectively, when using five times faster devices with the same interconnect trees. Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
ICCAD | 1 |
| 1999 | Figures of merit to characterize the importance of on-chip inductanceabstractA closed-form solution for the output signal of a CMOS inverter driving an RLC transmission line is presented. This solution is based on the alpha power law for deep submicrometer technologies. Two figures of merit are presented that are useful for determining if a section of interconnect should be modeled as either an RLC or an RC impedance. The damping factor of a lumped RLC circuit is shown to be a useful criterion. The second useful figure of merit considered in this paper is the ratio of the rise time of the input signal at the driver of an interconnect line to the time of flight of the signals across the line. AS/X circuit simulations of an RLC transmission line and a five section RC II circuit based on a 0.25-/spl mu/m IBM CMOS technology are used to quantify and determine the relative accuracy of an RC model. One primary result of this paper is evidence demonstrating that a range for the length of the interconnect exists for which inductance effects are prominent. Furthermore, it is shown that under certain conditions, inductance effects are negligible despite the length of the section of interconnect. Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1998 | Figures of Merit to Characterize the Importance of On-Chip InductanceabstractA closed form solution for the output signal of a CMOS inverter driving an RLC transmission line is presented. This solution is based on the alpha power law for deep submicrometer technologies. Two figures of merit are presented that are useful for determining if a section of interconnect should be modeled as either an RLC or an RC impedance. The damping factor of a lumped RLC circuit is shown to be a useful figure of merit. The second useful figure of merit considered in this paper is the ratio of the rise time of the input signal at the driver of an interconnect line to the time of flight of the signals across the line. AS/X circuit simulations of an RLC transmission line and a five section RC II circuit based on a 0.25 µm IBM CMOS technology are used to quantify and determine the relative accuracy of an RC model. One primary result of this study is evidence demonstrating that a range for the length of the interconnect exists for which inductance effects are prominent. Furthermore, it is shown that under certain conditions, inductance effects are negligible despite the length of the section of interconnect. Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
DAC | 1 |
| 1998 | Dynamic and Short-Circuit Power of CMOS Gates Driving Lossless Transmission LinesabstractThe dynamic and short-circuit power consumption of a CMOS gate driving an LC transmission line as a limiting case of an RLC transmission line is investigated in this paper. Closed form solutions for the output voltage and short-circuit power of a CMOS gate driving an LC transmission line are presented These solutions agree with AS/X circuit simulations within 11% error for a wide range of transistor widths and line impedances. The ratio of the short-circuit to dynamic power is shown to be less than 7% for CMOS gates driving LC transmission lines where the line is matched or underdriven. The total power consumption is expected to decrease as inductance effects become more significant as compared to an RC dominated interconnect. Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
Great Lakes Symposium on VLSI | 1 |
| 1998 | Power dissipated by CMOS gates driving lossless transmission linesabstractThe dynamic and short-circuit power consumption of a CMOS gate driving an LC transmission line as a limiting case of an RLC transmission line is investigated in this paper. Closed form solutions for the output voltage and short-circuit power of a CMOS gate driving an LC transmission line are presented. These solutions agree with AS/X simulations within 11% error for a wide range of transistor widths and line impedances. The ratio of the short-circuit to dynamic power is less than 7% for CMOS gates driving LC transmission lines where the line is matched or underdriven. Therefore, the total power consumption is expected to decrease as inductance effects becomes more significant is compared to an RC model of the interconnect. Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
ISLPED | 1 |