EDBT 2026 Demo / reviewers in the wild / expert
Eby G. Friedman
dblp:65/1497
· DBLP profile ↗
191ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0002-5549-7160ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 189 · 3 first-author · 16 since 2021Software engineering, systems software and programming languages · 4Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Inductorless dynamic logic based on 2 ϕ -Josephson junctions
Ana Mitrovic, Eby G. Friedman |
Integr. | 2 |
| 2025 | Network-on-Interposer Co-Design for Heterogeneous Chiplet-Based Integrated SystemsabstractA co-design methodology is proposed for heterogeneous integrated systems with active interposers. Bus-based topologies are impractical for heterogeneous systems due to the latency of the chiplet interfaces which degrades the data transfer process. System-on-chip and Internet-of-Things applications present very different tradeoffs for interposers than compute-based systems. A network-on-interposer methodology is proposed which balances cost and system performance. Active interposers are shown to exhibit less than half the latency of passive interposers when interfacing between 7 nm and 45 nm chiplets. Design considerations for both passive and active interposer-based systems are explored in this paper, such as maximum bandwidth, chiplet placement, voltage conversion, and network area. Andres Ayes, Eby G. Friedman, Marilyn Wolf |
ISCAS | 2 |
| 2025 | Partitioning SFQ Circuits for Serial Biasing in Heterogeneous CircuitsabstractThe significant bias currents in large scale single flux quantum (SFQ) systems can greatly tax the power delivery system and lead to system failure. Current recycling is a technique that reduces bias currents in SFQ systems by dividing similar circuits into several blocks or partitions. Each block is placed on a different ground plane and serially biased. Since the circuit block is placed on a different ground, a driver-receiver pair (DRP) is used to transfer clock or data signals between the circuit blocks. Current recycling is extended here to different circuit blocks with varying bias current requirements, described as a heterogeneous circuit. Current recycling in heterogeneous circuits is feasible with current biasing techniques: dummy Josephson transmission lines (JTL) can be inserted into those systems with small bias imbalances, while resistive trees can be used in those systems with large bias imbalances. A partitioning algorithm for heterogeneous circuits is proposed to address the large bias imbalance that occurs after inserting DRPs. This partitioning algorithm uses a combination of the Kernighan–Lin (KL) and Fiduccia-Mattheyses (FM) algorithms while minimizing the number of DRPs. The heterogeneous partitioning algorithm is demonstrated on several ISCAS’89 benchmark circuits, incorporating both resistive tree and DRP insertion to achieve an average reduction of 29.7% in current, 31% reduction in DRPs, and a 40.9% reduction in the number of Josephson junctions. Tejumadejesu Oluwadamilare, Eby G. Friedman |
ISCAS | 2 |
| 2025 | High Efficiency Multiply-Accumulator Using Ternary Logic and Ternary Approximate AlgorithmabstractA multiply-accumulator, often abbreviated as a MAC unit, is central to a multitude of computational tasks, particularly those tasks (such as neural networks) involving array-based mathematical computations. The quest for novel methods to efficiently store and process data in a MAC has become imperative. Recently, ternary logic has attracted significant attention due to its higher information density than conventional binary systems. However, though numerous studies have showcased ternary arithmetic circuits, advancements in ternary-based vector processing have been notably scarce. To bridge this gap, this work undertakes comprehensive study into the optimization of ternary MAC units. Firstly, we propose various ternary approximate algorithms which shows 30%-less power consumption and only 2% computation error when compared with the accurate design. Secondly, we design sophisticated ternary circuits and obtain 74%~80% lower power-delay-product (PDP) than previous works. Finally, we evaluate the proposed ternary MAC unit using both carbon-nanotube field-effect transistor (CNTFET) and silicon-based 180 nm CMOS processes. The simulation results show the ternary circuit is better than binary circuit in terms of both area (~45% less) and power (~30% less), highlighting its strong potential for practical applications. Wanting Wen, Guangchao Zhao, Wanbo Hu, Ziye Li, Xingli Wang, Eby G. Friedman, Beng Kang Tay, Shaolin Ke, Mingqiang Huang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | Test Modules for Enhanced Testability of Single Flux Quantum Integrated Circuits
Abdelrahman G. Qoutb, Jamil Kawa, Eby G. Friedman |
J. Electron. Test. | 3 |
| 2024 | Power Aware Placement of On-Chip Voltage RegulatorsabstractIn traditional power delivery networks, the on-chip supply voltage is provided by board-level converters. Due to the significant distance between the converter and the load, variations in the load current are not effectively managed, producing a significant voltage drop at the point-of-load. To mitigate this issue, modern high-performance systems utilize on-chip voltage regulators. Due to the close proximity to the load, these regulators can quickly respond to fluctuations in the input voltage or load current, providing superior power quality. Integrated voltage regulators however require significant area, limiting the number of on-chip regulators. An algorithm for distributing on-chip voltage regulators is presented in this article. The algorithm is accelerated using the acrlong IMT, enabling the analysis of arbitrarily sized power grids. The power quality is maximized with a limited number of regulators. Practical scenarios are supported, such as limited current capacity and restricted placement. Several orders of magnitude speedup in the placement process is demonstrated while achieving up to 88% reduction in the maximum voltage drop. Rassul Bairamkulov, Eby G. Friedman |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | Quasi-Adiabatic Clock Networks in 3-D Voltage Stacked SystemsabstractPower delivery in three-dimensional (3-D) integrated systems poses several challenges such as high current densities, large voltage drops due to multiple levels of resistive vertical interconnect, and significant switching noise originating from transient currents within different layers. Voltage stacking is a power delivery technique that is highly compatible with 3-D integration due to the physical proximity between layers, enabling the efficient transfer of recycled current. Power noise in clock networks is, however, not inherently addressed by 3-D voltage stacking. In this brief, a quasi-adiabatic technique between multiple clock networks within 3-D voltage stacked systems is proposed. The technique exploits the proximity of the clock networks to enable mutual charging and discharging when the clock signals transition to the same voltage. During this transition, the clock distribution networks are isolated from the power grid, reducing simultaneous switching noise and current load. The maximum current is reduced by an additional 13% as compared to only voltage stacking, the maximum voltage noise is reduced by up to 72% when the clock networks are isolated from the power grids, and the clock networks pull nearly 50% less charge from the source. The proposed technique is evaluated on a 7 nm predictive technology model. Andres Ayes, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | Thermal Exploration of RSFQ Integrated CircuitsabstractThe maturing of rapid single flux quantum (RSFQ) circuits into a VLSI complexity technology has focused the need for advanced design and analysis capabilities. For example, the temperature dependence of RSFQ circuits has created a need for a thermal analysis methodology, including an accurate thermal model and related partitioning algorithm. The Josephson critical current density and the superconductive properties of the interconnects become compromised at elevated temperatures. Heating of Josephson junctions (JJs) and niobium interconnects results in reduced margins and/or functional failure. A methodology for evaluating the thermal properties of RSFQ integrated circuits, targeting large-scale systems, is presented here. This methodology comprises a thermal model and a multistage partitioning algorithm. The algorithm, based on a layout of the IC, partitions the circuit into blocks. An average thermal model is applied to the partitioned structure, producing a netlist for thermal simulation. The hot spots are also determined with a threshold temperature indicating functional failure. A peak thermal model is presented to detect hotspots based on the threshold temperature. The model is evaluated at several block complexities and validated using a numerical solver. An error of less than 1% between the model and numerical simulations is achieved. The algorithm is applied to the AMD2901 benchmark circuit composed of more than 300000 heating elements. The thermal profile for AMD2901 is generated in under 68 min. Ana Mitrovic, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | Linear Clock Tree Topology for Dynamic Source Synchronous and Fully Synchronous 3-D Interfaces
Andres Ayes, Eby G. Friedman |
Integr. | 2 |
| 2023 | Thermal Optimization of Hybrid Cryogenic Computing SystemsabstractHeterogeneous computing exploits several disparate technologies within a single system. The different components of a heterogeneous system are often placed within separate temperature zones. Selecting an appropriate operating temperature strongly affects the dissipated power, cooling power (heat load), system performance, and ambient temperature. To this date, no multitemperature design methodology exists. To overcome this limitation, a framework for thermal optimization of heterogeneous computing systems is presented in this article. The effects of operating temperature on delay and power consumption are characterized based on a graph representation of the system. In addition, thermal interactions among the components within a system are considered to accurately evaluate the total power consumption and heat load. In a practical case study, the target temperature of each component within a quantum computing system is determined to minimize the total power under target performance constraints. Nurzhan Zhuldassov, Rassul Bairamkulov, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | Double Magnetic Tunnel Junction-Based Nonvolatile LogicabstractIn exascale computing, a huge amount of data is processed in real-time. Conventional CMOS-based computing paradigms follow the read, compute, and write back mechanism, which consumes significant power and time to compute and store data. in situ computation - where data are processed within the memory system - is considered a platform for exascale computation. A spin transfer torque perpendicular magnetic tunnel junction (PMTJ) is a nonvolatile memory device with several potential advantages (fast read/write, high endurance, and CMOS compatibility) to become a next generation memory solution. PMTJ offers the possibility of constructing both standalone and embedded RAM as well as MTJ-based VLSI computing. The double magnetic tunnel junction (DMTJ) is an emerging device composed of two serially connected PMTJs. In this paper, a DMTJ-based multi-bit memory cell that also provides a nonvolatile logic compute capability is presented. The multi-level cell provides both a high speed read/write multi-bit memory cell and a nonvolatile AND, OR, and NOT logic gate that computes and stores input data in real-time with a delay of 70 ns in a 32 nm CMOS technology node. Abdelrahman G. Qoutb, Eby G. Friedman |
ISCAS | 2 |
| 2022 | Converter Topologies for On-Package Voltage StackingabstractThe rise of mobile technologies and cloud computing has increased the importance of energy consumption. On-package voltage stacking, where current is recycled between multiple cores, is a potentially effective solution to this growing issue. Two converters, a load-to-load ladder buck converter and a bus-to-load isolated resonant converter, are particularly appropriate for on-package voltage stacking. A four core system composed of either converter topology is evaluated under several current mismatch scenarios in terms of transient and DC voltage drops, voltage ripple, settling time, and power efficiency. The load-to-load buck converter exhibits higher power efficiency and smaller transient voltage drops as compared to the bus-to-load resonant converter. The resonant converter is lower cost, smaller in area, and provides isolation. Nurzhan Zhuldassov, Kan Xu, Eby G. Friedman |
ISCAS | 3 |
| 2022 | QuCTS - Single-Flux Quantum Clock Tree SynthesisabstractSuperconductive rapid single-flux quantum (RSFQ) is an emerging cryogenic technology, promising a significant boost in performance and ultralow power consumption. The operating frequency achieved by RSFQ digital integrated circuits is several orders of magnitude greater than traditional CMOS circuits. The fundamental difference of RSFQ circuits, however, renders traditional clocking techniques appropriate for CMOS unsuitable for RSFQ technology. Most RSFQ logic gates, such as AND and OR, are sequential in nature. The number of pipeline stages is therefore significantly greater in RSFQ as compared to CMOS, complicating the clock distribution network design process. This issue is further exacerbated with the need for splitters to achieve a fanout greater than one and the need for transmission lines rather than ordinary metallic wires as in CMOS. In this work, QuCTS—single-flux Quantum (SFQ) Clock Tree Synthesis—is presented. QuCTS utilizes a two-stage framework for synthesizing clock networks. In the clock skew scheduling stage, the clock signal arrival time of each gate is chosen to maximize the robustness of the circuit to timing variations. In the clock tree synthesis stage, the layout of the clock distribution network is generated based on a novel delay equilibration technique. QuCTS is the first clock tree synthesis tool for RSFQ circuits utilizing useful clock skew. The synthesized network satisfies the clock arrival time requirements while minimizing the associated overhead, such as the interconnect length and number of delay elements. The tool is validated on a set of benchmark circuits. In a prototypical case study, a clock tree is generated for the AMD2901 with 1049 clock sinks in 53 min while satisfying the clock arrival time. Rassul Bairamkulov, Tahereh Jabbari, Eby G. Friedman |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | SPROUT - Smart Power Routing Tool for Board-Level Exploration and PrototypingabstractThe board-level power network design process is governed by system-level parameters, such as the number of layers and the ball grid array (BGA) pattern. These parameters influence the characteristics of the resulting system, such as power, speed, and cost. Evaluating the impact of these parameters is however challenging. To estimate the reduction in impedance if, for example, additional BGA balls are dedicated to the power delivery system, adjustments to the board layout, and an additional impedance extraction process are required. These processes are poorly automated, requiring significant time and labor. Automating power network exploration and prototyping can greatly enhance the board-level power delivery design process by increasing the number of possible design options. With power network exploration and prototyping, the effects of the system parameters on the electrical characteristics can be better understood, providing valuable insight into the early stages of the design process. SPROUT—an automated algorithm for prototyping printed circuit board (PCB) power networks—is presented here. This tool includes the first fully automated algorithm for board-level power network layout synthesis. Two board-level industrial power networks are synthesized using SPROUT. The impedance of the resulting layouts exhibits good agreement with manual PCB layouts while significantly reducing the design time. The tool is used to explore area/impedance tradeoffs in a three-rail system, providing useful data to enhance the PCB design process. Rassul Bairamkulov, Abinash Roy, Nagarajan Mahalingam, Vaishnav Srinivas, Eby G. Friedman |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | SPROUT - Smart Power ROUting Tool for Board-Level Exploration and PrototypingabstractThe board-level power network design process is governed by system-level parameters such as the number of layers and the ball grid array (BGA) pattern. These parameters influence the characteristics of the resulting system, such as power, speed, and cost. Evaluating the impact of these parameters is, however, challenging. To estimate the reduction in impedance if, for example, additional BGA balls are dedicated to the power delivery system, adjustments to the board layout and an additional impedance extraction process are required. These processes are poorly automated, requiring significant time and labor. Automating power network exploration and prototyping can greatly enhance the power delivery design process by increasing the number of possible design options. With power network exploration and prototyping, the effects of the system parameters on the electrical characteristics can be better understood, providing valuable insight into the early design stages. SPROUT - an automated algorithm for prototyping printed circuit board (PCB) power networks - is presented here. This tool includes the first fully automated algorithm for board-level power network layout synthesis. Two board-level industrial power networks are synthesized using SPROUT where a ball grid array is connected with a power management IC and decoupling capacitors across four voltage domains. The impedance of the resulting layouts is in good agreement with manual PCB layouts while requiring 95% less design time. Rassul Bairamkulov, Abinash Roy, Mali Nagarajan, Vaishnav Srinivas, Eby G. Friedman |
DAC | 5 |
| 2021 | MTJ-Based Dithering for Stochastic Analog-to-Digital ConversionabstractThe stochastic behavior of magnetic tunnel junctions (MTJ) finds use in many applications - from analog-to- digital conversion to neuromorphic computing. In this paper, a dithering method exploiting the stochastic behavior of an MTJ, based on the voltage controlled magnetic anisotropy effect, is proposed. This method is used to measure low frequency analog signals over an interval. The circuit is composed of two MTJ devices that convert an analog signal into a series of high resistance and low resistance stochastic states, creating a dithering effect. The behavior of the input signal is extracted from the switching frequency. The binary nature of the output signal reduces the complexity, enabling a small, fast, and energy efficient circuit for extracting different forms of information from an analog signal. Ana Mitrovic, Eby G. Friedman |
ISCAS | 2 |
| 2020 | Graph-Based Power Network Routing for Board-Level High Performance SystemsabstractUnlike the routing of on-chip power delivery networks which is a highly automated process, the routing of board-level power nets is usually performed manually. The process is complicated by geometric and electrical constraints that impose restrictions on the routing process. An automated board-level power routing algorithm is presented here which provides efficient generation and refinement of power network geometries at the layout level. In a case study, a routing path connecting a power management integrated circuit to a ball grid array is routed using the automated tool, producing a low impedance network while complying with metal resource and geometric constraints. Rassul Bairamkulov, Eby G. Friedman, Abinash Roy, Mali Nagarajan, Vaishnav Srinivas |
ISCAS | 2 |
| 2020 | Bias Distribution in ERSFQ VLSI CircuitsabstractRapid single flux quantum (RSFQ) circuits have recently attracted considerable attention as a promising beyond CMOS technology for exascale computing. Unlike conventional CMOS circuits, RSFQ circuits require a specific bias current delivered to each Josephson junction, making robust bias networks an issue of great importance for large scale integration. ERSFQ is an energy efficient, inductive bias scheme for RSFQ circuits, where power dissipation is drastically lowered by eliminating the bias resistors while the cell library remains unchanged. An ERSFQ bias scheme requires the introduction of multiple circuit elements - current limiting Josephson junctions, bias inductors, and Josephson transmission lines. Multiple guidelines exist for the effective design of these structures. In this paper, additional parameter guidelines and design techniques are presented to decrease physical area and dynamic power dissipation while improving bias margins. These guidelines and techniques are applicable to automating the synthesis of bias networks to enable large scale ERSFQ circuits. Gleb Krylov, Eby G. Friedman |
ISCAS | 2 |
| 2020 | Spintronic/CMOS-Based Thermal SensorsabstractStacking more systems into a compact area or scaling devices to increase the density of integration are two approaches to provide greater functional complexity. Excessive heat generated as a result of these technology advancements leads to an increase in leakage power and degradation in system reliability. Hence, a thermal aware system composed of hundreds of distributed thermal sensor nodes is needed. Such a system requires an efficient thermal sensor placed close to the thermal hotspots, small in size, fast response, and CMOS compatibility. In this paper, two hybrid spintronic/CMOS circuits are proposed. These circuits exhibit a low power consumption of 11.9 μW during the on-state, a linearity (R2) of 0.96 over the industrial temperature range of operation (-40 to 125)oC, and a sensitivity of 3.78 mV/K. Abdelrahman G. Qoutb, Eby G. Friedman |
ISCAS | 2 |
| 2020 | Multi-Bit CNT TSV for 3-D ICsabstractThrough substrate vias (TSVs) are a seminal component of three-dimensional (3-D) integrated circuits (ICs). Each TSV typically carries a single signal between two adjacent layers of a 3-D structure. A multi-bit carbon nanotube TSV is proposed in this paper to increase the number of I/Os among layers within 3-D ICs. The proposed multi-bit TSV can propagate multiple independent signals due to the high anisotropy of the carbon nanotubes. The electrical properties of each bit within a two-bit TSV and the electrical interactions between the bits are compared to a theory-based electrical model, exhibiting high accuracy. The passive elements deviate by up to 4%, and the S-parameters of the system deviate by up to 1.5% from numerical analysis. Capacitive coupling and leakage current between the bits of the two-bit TSV model have also been evaluated. The structure exhibits negligible noise coupling (less than 1%) and a peak leakage current of 631.7 μA. Boris Vaisband, Ange Maurice, Chong Wei Tan, Beng Kang Tay, Eby G. Friedman |
ISCAS | 5 |
| 2020 | Challenges in High Current On-Chip Voltage Stacked SystemsabstractDue to the increasing throughput of high performance integrated circuits, the power consumption of recent high performance computing systems has grown significantly, leading to high on-chip current demand. The large current flowing within the power delivery network leads to challenging issues such as electromigration, low power efficiency, and thermal hotspots. As a technique to reduce on-chip current demand, voltage stacking has become a topic of growing interest within the industrial and academic communities. The challenges of on-chip voltage stacking are however significant. The limitations of relying on on-chip decoupling capacitors when load imbalances occur within a high current system are reviewed. To manage these load imbalances, a ladder topology switched capacitor converter is proposed to regulate the voltages between layers within a voltage stacked system. A 20X improvement in voltage drop is demonstrated on a case study. The current path within a voltage stacked system is quite different from a standard system. A horizontal current path is formed due to the serial connection between layers, producing large parasitic impedances within the power network. The on-chip power network within a voltage stacked system therefore requires careful consideration and specialized design techniques. Kan Xu, Eby G. Friedman |
ISCAS | 2 |
| 2020 | Distributed Port Assignment for Extraction of Power Delivery NetworksabstractThe stringent requirements of power noise on complex multi-domain power delivery networks (PDN), and the complicated relationship between signal integrity and power integrity (PI) have led to an ever challenging PI sign-off process. A lumped PDN model is widely used, where the power network is treated as a two-port network with the impedances extracted by an electromagnetic solver. A distributed model of the power network is however preferred during a PI sign-off flow, providing a more accurate circuit model for time domain simulations. Hundreds or even thousands of ports need to be properly evaluated during the PDN extraction process, which can be computationally expensive and error prone. A Python tool is described here to enable a fast and configurable process for distributed port assignment during the PDN extraction process. An enhanced automation flow, integrated with the Python tool, has also been developed to support early power network exploration. In one case study, a 360X speedup in the port assignment process is achieved while revealing a high risk power network within the package. The proposed automation flow is versatile and highly adaptive for different power network topologies. Kan Xu, Eby G. Friedman, Mikhail Popovich, Gregory Sizikov |
ISCAS | 2 |
| 2020 | Cryogenic Dynamic LogicabstractCloud computing is increasing the demand for large scale, energy efficient, and fast computing systems. One circuit technique satisfying these goals is dynamic logic. Furthermore, since portability is not required for cloud computing centers, these systems can support cryogenic operation. Cryogenic dynamic circuits eliminate the seminal issue of these circuits, loss of logic state due to leakage currents. The operation of dynamic CMOS circuits operating at cryogenic temperatures is discussed in this paper. For a 160 nm MOSFET technology, dynamic logic at room temperature can operate as low as 180 kHz. The same circuit in a cryogenic temperature can operate at DC. The state in a dynamic logic circuit operating at cryogenic temperatures is shown to remain indefinitely. This property makes low frequency testing of dynamic logic more feasible, supporting the development of complex VLSI circuits targeting high frequency applications such as cloud computing. Nurzhan Zhuldassov, Eby G. Friedman |
ISCAS | 2 |
| 2020 | Power Delivery Exploration Methodology Based on Constrained OptimizationabstractThe conventional power network design process requires iterative modifications to the existing power network to eliminate hot spots and to converge to target impedance parameters. At later stages in the IC design process, this procedure may require significant time and human resources due to the limited flexibility to accommodate necessary changes. Power delivery exploration during early stages of the design process may bring considerable savings to the system development effort. The number of iterations may be greatly reduced by choosing the initial parameters sufficiently close to the optimum. This paper presents a power delivery exploration framework based on constrained global optimization. The power network parameters are estimated at early stages of the development process, while considering both electrical and nonelectrical factors, such as area and cost. A Laplace transform-based circuit simulator is described that is well suited for optimization purposes due to the high computational efficiency when a large number of iterations is required. The proposed framework has been applied to the distribution of voltage domains in a large scale complex integrated system, while minimizing the cost of the decoupling capacitor placement. The optimal number of voltage rails are determined, demonstrating an approximately 40% lower on-chip area than alternative solutions. Rassul Bairamkulov, Kan Xu, Mikhail Popovich, Juan Ochoa, Vaishnav Srinivas, Eby G. Friedman |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2020 | Distributed Pass Gates in Power Delivery Systems With Digital Low-Dropout RegulatorsabstractOn-chip digital low-dropout (LDO) regulators enable fast dynamic voltage scaling, reducing power consumption. Integrating these regulators into a highly resistive environment has complicated the design of power delivery systems. With the increasing sensitivity of complex integrated systems to power noise, effective approaches to distribute on-chip LDOs are needed due to the limited metal resources. In this article, a methodology is proposed to distribute the pass gates of a system of on-chip digital LDOs. The distribution of the pass gates considers the location of the load currents to reduce voltage variations across the power grid. The proposed pass gate distribution topology reduces the maximum voltage variations across the grid, on average, by two to three times under nonuniform load distributions. Albert Ciprut, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | Design Methodology for Distributed Large-Scale ERSFQ Bias NetworksabstractRapid single-flux quantum (RSFQ) circuits have recently attracted considerable attention as a promising cryogenic beyond CMOS technology for exascale computing. Energy-efficient RSFQ (ERSFQ) is an energy-efficient, inductive bias scheme for RSFQ circuits, where the power dissipation is drastically lowered by eliminating the bias resistors, while the cell library remains unchanged. An ERSFQ bias scheme requires the introduction of multiple circuit elements-current limiting Josephson junctions, bias inductors, and feeding Josephson transmission lines (FJTLs). In this article, parameter guidelines and design techniques for ERSFQ circuits are presented. The proposed guidelines enable more robust circuits resistant to severe variations in supplied bias currents. Trends are considered, and advantageous tradeoffs are discussed for the different components within a bias network. The guidelines provide a means to decrease the size of an FJTL and, thereby, reduce the physical area, power dissipation, and overall bias currents, supporting further increases in circuit complexity. A distributed approach to ERSFQ FJTL is also presented to simplify placement and minimize the effects of the parasitic inductance of the bias lines. This methodology and related circuit techniques are applicable to automating the synthesis of bias networks to enable large-scale ERSFQ circuits. Gleb Krylov, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | Distributed Spintronic/CMOS Sensor Network for Thermal-Aware SystemsabstractRecent developments in IC technology rely on device scaling and 3-D integration, resulting in billions of devices compacted into a small area. These trends degrade the lifetime and reliability of the system due to an increase in temperature caused by high power densities. Dynamically managing a system based on the thermal characteristics is important to mitigate this issue. An on-chip thermal-aware system composed of hundreds of distributed thermal sensors is proposed. A hybrid spintronic/CMOS-based thermal sensor is described that exploits the thermal response and small area of an antiparallel magnetic tunnel junction. The sensor cell consumes as little as 500 pJ to read 1024 thermal sensor nodes and generates a thermal map of a system composed of 32 × 32 thermal sensors. The sensor cell exhibits a thermal linearity (R2) up to 0.983 and a thermal sensitivity of 1.91 mV/K over the commercial temperature range of 0 °C to 85 °C while consuming 32 μW. Abdelrahman G. Qoutb, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2019 | PMTJ Temperature Sensor Utilizing VCMAabstractThermal aware systems are able to control distributed CMOS blocks based on the local temperature to enhance system power, speed, and reliability. The ultimate objective is multiple in situ temperature sensors, close to the CMOS device layer, distributed over the die, physically small, and leaking almost zero power. Magnetic tunnel junctions (MTJ) are able to provide this capability. An MTJ is a CMOS compatible device, fabricated within the metallic layers, above the CMOS device layers. A method for using an MTJ as a thermal sensor is presented. The method operates MTJs in an antiparallel state where an MTJ is more sensitive to thermal variations as compared to the parallel state. The method exploits device magnetism, thermal stability, and resistance with respect to an applied sensing voltage to sense the ambient temperature. The results are based on experimentally extracted parameters of a perpendicular and voltage controlled magnetic anisotropy MgO|CoFeB MTJ. A change in the device antiparallel resistance of up to 16 Ω per degree Kelvin at a sensing voltage of 0.2 volts is exhibited. Abdelrahman G. Qoutb, Eby G. Friedman |
ISCAS | 2 |
| 2019 | Stability of On-Chip Power Delivery Systems With Multiple Low-Dropout RegulatorsabstractLow-dropout (LDO) voltage regulators have become prominent elements of on-chip power delivery systems due to the increasing importance of separate voltage domains, fast dynamic voltage scaling, and the need for high-quality power. As the number of LDOs sharing a common power network increases, the stability of the power grid can degrade if the resonance frequency due to the off-chip parasitic impedances is not sufficiently separated from the unity gain frequency of the regulator. In this paper, this source of instability in power delivery networks, consisting of multiple LDO regulators with a balanced load, is described. The effect of the number of LDOs on the resonance frequency of the power delivery network is evaluated to determine the stability of the system when the LDOs operate under similar load conditions. Albert Ciprut, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | MTJ Magnetization Switching Mechanisms for IoT ApplicationsabstractDifferent magnetization mechanisms, structures, and electrical properties of MTJ-based MRAM are described in the context of IoT applications. MTJ-based MRAM provides non-volatility (high retention time), low power, and high speed. This memory has a broad variety of applications. One important example is IoT. A comparative study of the magnetization mechanisms provides insight into which MTJ structures and magnetization mechanisms best support different IoT applications. Abdelrahman G. Qoutb, Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 2 |
| 2018 | Versatile Framework for Power Delivery ExplorationabstractOver the past decades, aggressive voltage scaling combined with increased power demands has placed stringent requirements on on-chip power quality. Unwanted voltage fluctuations and droops may cause a variety of issues, ranging from glitch power to device malfunction. If revealed at the later stages of the design process, mitigation techniques may become unbearably costly in both time and money. A framework for exploratory power delivery optimization is described to enhance the power delivery network during early stages of the design process in accordance with design specifications. The power delivery design process is converted into a constrained minimization problem, consisting of design metrics combined into objective and constraint functions. The framework supports the optimization of the power network characteristics while considering external, non-electrical design specifications, such as cost and area, providing a comprehensive network analysis capability. In one case study, a 15% reduction in decoupling capacitor placement along with a 38.6% reduction in power consumption is achieved while satisfying performance and power quality constraints. Rassul Bairamkulov, Kan Xu, Eby G. Friedman, Mikhail Popovich, Juan Ochoa, Vaishnav Srinivas |
ISCAS | 3 |
| 2018 | Hybrid Write Bias Scheme for Non-Volatile Resistive Crossbar ArraysabstractCrossbar arrays based on non-volatile resistive devices are planned for future memory systems due to the scalability and performance as compared to conventional charge based memory systems. To enhance the feasibility of these resistive memory systems, the energy consumption needs to be reduced. The write operation of a resistive memory based on a one-selector-one-resistor crossbar array consumes significant energy. The energy consumed by a crossbar array is dependent on the device and interconnect characteristics as well as the bias scheme. While the device and circuit parameters are the same for a specific application, the bias scheme of an array can be tuned to improve the energy efficiency. In this paper, an intelligent write scheme is proposed to provide a hybrid bias scheme. The proposed system adaptively sets the bias schemes to enhance energy efficiency. The most energy efficient bias scheme depends upon several parameters such as the size of the array, nonlinearity factor, and number of selected cells. For a specific array size and device characteristics, a power delivery system is described that sets the bias voltages based on the number of selected cells. Energy improvements of more than 2× are demonstrated with this hybrid bias scheme. Albert Ciprut, Eby G. Friedman |
ISCAS | 2 |
| 2018 | Behavioral Verilog-A Model of Superconductor-Ferromagnetic TransistorabstractThe superconductor-ferromagnetic transistor (SFT) is a novel cryogenic device with the potential to greatly enhance traditional single flux quantum (SFQ) circuits. Since SFT devices are under active development, compact models are necessary to evaluate this device in novel circuits. In this paper, a simplified compact model of a three terminal SFT device is proposed. The model fits the general I-V characteristics of existing devices with 7.4% mean absolute error, while also capturing the transient behavior of the device. The model has been implemented in Verilog-A and simulated in Cadence Spectre. The proposed model enables the simulation of SFQ circuits containing SFT devices, and is reconfigurable to support developments in SFT technology. Gleb Krylov, Eby G. Friedman |
ISCAS | 2 |
| 2018 | Heterogeneous 3-D ICs as a platform for hybrid energy harvesting in IoT systems
Boris Vaisband, Eby G. Friedman |
Future Gener. Comput. Syst. | 2 |
| 2018 | Exploratory design of on-chip power delivery for 14, 10, and 7 nm and beyond FinFET ICs
Kan Xu, Ravi Patel 0001, Praveen Raghavan, Eby G. Friedman |
Integr. | 4 |
| 2018 | Energy-Efficient Write Scheme for Nonvolatile Resistive Crossbar Arrays With SelectorsabstractThe write operation of a resistive memory based on a one-selector-one-resistor (1S1R) crossbar array consumes significant energy and is dependent on the device and circuit characteristics as well as the bias scheme. In this paper, the energy efficiency of a crossbar array of a 1S1R configuration during a write operation is explored for the V/2 and V/3 bias schemes. The characteristics that affect the most energy-efficient bias scheme are demystified. The write energy of a crossbar array is modeled in terms of the array size, number of selected cells, and nonlinearity factor. For a specific array size and selector technology, the number of selected cells during a write operation can affect the choice of bias scheme. The effect of leakage current due to partially biased unselected cells is explored. Furthermore, an energy-efficient write operation based on a hybrid bias scheme is proposed to reduce the write energy. This write operation adaptively sets the bias scheme based on the number of selected cells to enhance overall energy efficiency. Energy improvements of more than two times are demonstrated with this hybrid bias scheme. Albert Ciprut, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Test point insertion for RSFQ circuitsabstractA test point insertion technique for enhanced testability in superconductive RSFQ logic is proposed here. Test point insertion reduces the overhead of a set/scan chain while maintaining most of the functionality. The SFQ multiplexers are replaced with mergers and blocking gates. The multiplexer control signals are replaced with a gated clock signal. Clocked blocking gates are used to disable undesired data input. The proposed technique requires 35% fewer Josephson junctions as compared to multiplexers when applied to a 64-bit register. This technique can be utilized to evaluate error prone SFQ circuits and for built-in self-test of SFQ compatible memory systems. Gleb Krylov, Eby G. Friedman |
ISCAS | 2 |
| 2017 | Hybrid energy harvesting in 3-D IC IoT devicesabstractThree-dimensional integrated circuits are a natural platform for IoT devices. IoT devices exhibit a small footprint, integrate disparate technologies, and require long term sustainability (extremely low power or self powered). A hybrid energy harvesting system within a three-dimensional integrated circuit is proposed in this paper. The harvesting system exploits different types of energy available from the ambient (electromagnetic, solar, thermal, and kinetic). Integration of the hybrid harvesting system onto a three-dimensional platform ensures that each type of harvested energy can be individually optimized. In addition, lower parasitic impedances are exhibited within the three-dimensional structure, leading to improved efficiencies in the energy harvesting process. For an example IoT system, the power requirements are less than 57% of the power delivered to the load. Boris Vaisband, Eby G. Friedman |
ISCAS | 2 |
| 2017 | Modeling Size Limitations of Resistive Crossbar Array With Cell SelectorsabstractDue to recent developments in emerging memory technologies, resistive crossbar arrays have gained increasing importance. The size of the crossbar arrays is, however, limited due to challenges brought by the interconnect resistance, sneak path currents, and the physical area of the peripheral circuitry. In this paper, three figures of merit that characterize the limitations of resistive crossbar arrays with selectors are described, such as the driver resistance, voltage degradation across the cell, and read margin. The models, exhibiting good agreement with SPICE, are compared with different biasing schemes during both write and read operations. These models are also used to predict the device requirements of resistive crossbar arrays with selectors and to project parameter values, such as the nonlinearity factor, on-state resistance, and tolerable interconnect resistance per cell for large-scale crossbar arrays. Albert Ciprut, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Exploratory Power Noise Models of Standard Cell 14, 10, and 7 nm FinFET ICsabstractThe physical dimensions of standard cells constrain the dimensions of power networks, affecting the on-chip power noise. An exploratory modeling methodology is presented for estimating power noise in advanced technology nodes. The models are evaluated for 14, 10, and 7 nm technologies to assess the impact on performance. Scaled technologies are shown to be more sensitive to power noise, resulting in potential loss of performance enhancements achieved by scaling. Stripes between local track rails is evaluated as a means to reduce power noise, exhibiting up to 56.5% improvement in power noise for the 7 nm technology node. A strong dependence on the width of a stripe is observed, indicating that fewer wide stripes are more favorable then many thin stripes. As a promising alternative material for power network interconnects, graphene is shown to exhibit good potential in reducing power noise. The effects of different scaling scenarios of local power rails on power noise are also discussed. Ravi Patel 0001, Kan Xu, Eby G. Friedman, Praveen Raghavan |
ACM Great Lakes Symposium on VLSI | 3 |
| 2016 | Design models of resistive crossbar arrays with selector devicesabstractDue to recent developments in emerging memory technologies such as MRAM/RRAM, resistive crossbar arrays have gained increasing importance. The size of the crossbar arrays is, however, limited due to challenges brought by the interconnect resistance, sneak path currents, and the physical area of the peripheral circuitry. In this paper, three figures of merit that characterize the limitations of resistive crossbar arrays with selectors, the driver resistance, voltage degradation across the cell, and read margin, are discussed. Models are described that exhibit good agreement with SPICE, exhibiting a maximum error of 6.5% for the worst case voltage degradation during a write operation and 6% during a read operation for voltage ratios above, respectively, 0.5 and 0.25. Furthermore, these models are used to predict device requirements of resistive crossbar arrays with selectors and to project parameter values such as the nonlinearity factor, resistance in the on state, and tolerable interconnect resistance per cell for large scale crossbar arrays. Albert Ciprut, Eby G. Friedman |
ISCAS | 2 |
| 2016 | Power noise in 14, 10, and 7 nm FinFET CMOS technologiesabstractA methodology is described to determine the distribution of current within a power network for use in CMOS standard cell integrated circuits based on exploratory information about the power network. Models are presented to extrapolate the noise within power networks in 14, 10, and 7 nm CMOS technologies. Stripes, interconnect between local power rails, are evaluated as a means to reduce power noise, resulting in a 56.5% reduction in noise for the 7 nm CMOS technology node. Ravi Patel 0001, Eby G. Friedman, Praveen Raghavan |
ISCAS | 2 |
| 2016 | Layer ordering to minimize TSVs in heterogeneous 3-D ICsabstractA layer ordering algorithm to minimize the total number of TSVs within heterogeneous 3-D integrated circuits is described in this paper. Different constraints may complicate the process of ordering the layers within a 3-D system. These constraints are (1) any two layers must be adjacent, (2) a layer must be placed at a specific location, and (3) a layer must be separated from another layer(s). The algorithm generates an optimal layer order given the number of I/Os among all layers. Certain layers can be pre-assigned to specific locations within the 3-D structure. The application of the algorithm to multiple layer 3-D structures significantly reduces the number of TSVs and occupied area as compared to a random layer assignment. The area overhead of a random solution as compared to the optimal solution for unconstrained 3-D systems (without pre-assigned layers) with three to ten layers is, respectively, ~24,090 μm2to ~854,469 μm2. In constrained 3-D systems (with pre-assigned layers), the area overhead for an eight layer 3-D system with one to six assigned layers ranges up to ~249, 240 μm2. Boris Vaisband, Eby G. Friedman |
ISCAS | 2 |
| 2016 | Adaptive power gating of 32-bit Kogge Stone adder
Alexander E. Shapiro, Francois Atallah, Kyugseok Kim, Jihoon Jeong, Jeff Fischer, Eby G. Friedman |
Integr. | 6 |
| 2016 | Back to the Future: Current-Mode Processor in the Era of Deeply Scaled CMOSabstractThis paper explores the use of MOS current-mode logic (MCML) as a fast and low noise alternative to static CMOS circuits in microprocessors, thereby improving the performance, energy efficiency, and signal integrity of future computer systems. The power and ground noise generated by an MCML circuit is typically 10-100× smaller than the noise generated by a static CMOS circuit. Unlike static CMOS, whose dominant dynamic power is proportional to the frequency, MCML circuits dissipate a constant power independent of clock frequency. Although these traits make MCML highly energy efficient when operating at high speeds, the constant static power of MCML poses a challenge for a microarchitecture that operates at the modest clock rate and with a low activity factor. To address this challenge, a single-core microarchitecture for MCML is explored that exploits the C-slow retiming technique, and operates at a high frequency with low complexity to save energy. This design principle contrasts with the contemporary multicore design paradigm for static CMOS that relies on a large number of gates operating in parallel at the modest speeds. The proposed architecture generates 10-40× lower power and ground noise, and operates within 13% of the performance (i.e., 1/ExecutionTime) of a conventional, eight-core static CMOS processor while exhibiting 1.6× lower energy and 9% less area. Moreover, the operation of an MCML processor is robust under both systematic and random variations in transistor threshold voltage and effective channel length. Yanwei Song, Mahdi Nazm Bojnordi, Alexander E. Shapiro, Eby G. Friedman, Engin Ipek |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2016 | Reducing Switching Latency and Energy in STT-MRAM Caches With Field-Assisted WritingabstractA field-assisted spin-torque transfer magnetoresistive RAM (STT-MRAM) cache is presented for the use in high-performance energy-efficient microprocessors. Adding field assistance reduces the switching latency by a factor of 4. An array model is developed to evaluate the switching energy for different field currents and array sizes. Several STT-MRAM-based cells demonstrate a 55% energy reduction as compared with an SRAM cache subsystem. As compared with STT-MRAM caches with subbank buffering and differential writes, a field-assisted STT-MRAM cache improves the system performance by 28%, with a 6.7% increase in energy. Ravi Patel 0001, Qing Guo 0004, Engin Ipek, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2016 | Power Efficient Level Shifter for 16 nm FinFET Near Threshold CircuitsabstractSince the minimum feature size has shrunk beyond the sub-30-nm node, power density has become the major factor in modern microprocessors. Techniques such as dynamic voltage scaling operating down to near threshold voltage levels and supporting multiple voltage domains have become necessary to reduce dynamic as well as static power. A key component of these techniques is a level shifter that serves different voltage domains. This level shifter must be high speed and power efficient. The proposed level shifter translates voltages ranging from 250 to 790 mV, and exhibits 42% shorter delay, 45% lower energy consumption, and 48% lower static power dissipation. In addition, the proposed level shifter exhibits symmetric rise and fall transition times with up to 12% skew at the extreme conditions over the maximum range of voltages. Alexander E. Shapiro, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Noise Coupling Models in Heterogeneous 3-D ICsabstractModels of coupling noise from an aggressor module to a victim module by way of through silicon vias (TSVs) within heterogeneous 3-D integrated circuits (ICs) are presented in this paper. Existing TSV models are enhanced for different substrate materials within heterogeneous 3-D ICs. Each model is adapted to each substrate material according to the local noise coupling characteristics. The 3-D noise coupling system is evaluated for isolation efficiency over frequencies of up to 100 GHz. Isolation improvement techniques, such as reducing the ground network inductance and increasing the distance between the aggressor and victim modules, are quantified in terms of noise improvements. A maximum improvement of 73.5 dB for different ground network impedances and a difference of 38.5 dB in isolation efficiency for greater separation between the aggressor and victim modules are demonstrated. Compact, accurate, and computationally efficient models are extracted from the transfer function for each of the heterogeneous substrate materials. The reduced transfer functions are used to explore different manufacturing and design parameters to evaluate coupling noise across multiple 3-D planes. Boris Vaisband, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | PNS-FCR: Flexible Charge Recycling Dynamic Circuit Technique for Low-Power MicroprocessorsabstractDue to the superior speed and area characteristics, dynamic circuits are widely applied in data paths and other time critical components in modern microprocessors. The high switching activity of dynamic circuits, however, consumes significant power. In this paper, a p-type/n-type dynamic circuit selection (PNS) algorithm and a flexible charge recycling (FCR) design methodology are proposed to achieve high power efficiency in data paths. The effects of technology scaling, data path width, design complexity, clock skew, and environmental conditions are discussed. Simulation results show that the power consumption of an arithmetic and logic unit (ALU) with the proposed PNS-FCR can be reduced by up to 60% as compared with a conventional ALU. An 8-bit ALU test circuit has also been manufactured based on a 0.35-μm Global Foundries technology, demonstrating the power and area efficiency of the proposed methodology. Na Gong, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Architecting a MOS current mode logic (MCML) processor for fast, low noise and energy-efficient computing in the near-threshold regimeabstractNear-threshold computing (NTC) is an effective technique for improving the energy efficiency of a CMOS microprocessor, but suffers from a significant performance loss and an increased sensitivity to voltage noise. MOS current-mode logic (MCML), a differential logic family, maintains a low voltage swing and a constant current, making it inherently fast and low-noise. These traits make MCML a natural selection to implement an NTC processor; however, MCML suffers from a high static power regardless of the clock frequency or the level of switching activity, which would result in an inordinate energy consumption in a large scale IC. To address this challenge, this paper explores a single-core microarchitecture for MCML that takes advantage of C-slow retiming technique, and runs at a high frequency with low complexity to save energy. This design principle is opposite to the contemporary multicore design paradigm for static CMOS that relies on a large number of gates running in parallel at modest speeds. When compared to an eight-core static CMOS processor operating in the near-threshold regime, the proposed processor exhibits 3x higher performance, 2x lower energy, and 10 x lower voltage noise, while maintaining a similar level of power dissipation. Yanwei Song, Mahdi Nazm Bojnordi, Alexander E. Shapiro, Engin Ipek, Eby G. Friedman |
ICCD | 6 |
| 2015 | 3-D floorplanning algorithm to minimize thermal interactionsabstractAn algorithm for including relative thermal interactions among different circuit modules within a 3-D system is introduced in this paper. Application of the proposed algorithm on MCNC and GSRC benchmark circuits is presented. The thermal behavior of a heterogeneous 3-D structure, consisting of a different number of modules and substrate materials, is evaluated to emulate the heat transfer characteristics of practical heterogeneous 3-D systems. The algorithm lowers thermal interactions between different modules while maintaining the peak temperature within a practical range. The thermal characteristics of the floorplan are evaluated using HotSpot and HotSpot Detailed 3-D and compared to a random floorplan. The recorded peak temperatures are within the practical range of on-chip temperatures. Boris Vaisband, Eby G. Friedman |
ISCAS | 2 |
| 2015 | Inductive coupling effects in large TSV arraysabstractThe effects of inductive coupling among TSVs within large TSV arrays are investigated in this paper. A comparison of the equivalent inductance of a paired TSV model and arrayed TSV macromodel is presented for three TSV distribution topologies, grouped, lined, and uniform, within the power network. Modified closed-form expressions are proposed to determine the equivalent inductance of a TSV in these large TSV arrays. Simulation results show that this method achieves hundred times speed improvement as compared to an electromagnetic field solver while maintaining accuracy within 7%. Kan Xu, Eby G. Friedman |
ISCAS | 2 |
| 2015 | Energy efficient adaptive clustering of on-chip power delivery systems
Inna Partin-Vaisband, Eby G. Friedman |
Integr. | 2 |
| 2015 | Scaling trends of power noise in 3-D ICs
Kan Xu, Eby G. Friedman |
Integr. | 2 |
| 2015 | Multistate Register Based on Resistive RAMabstractIn recent years, memristive technologies, such as resistive random access memory (RRAM), have emerged. These technologies are usually considered as alternates for static RAM, dynamic RAM, and Flash. In this paper, a novel digital circuit, the multistate register, is proposed. The multistate register is different from conventional types of memory, and is used to store multiple data bits, where only a single bit is active and the remaining data bits are idle. The active bit is stored within a CMOS flip flop, while the idle bits are stored in an RRAM crossbar co-located with the flip flop. It is demonstrated that additional states require an area overhead of 1.4% per state for a 64-state register. The use of multistate registers as pipeline registers is demonstrated for a novel multithreading architecture-continuous flow multithreading (CFMT), where the total area overhead in the CPU pipeline is only 2.5% for 16 threads compared with a single thread CMOS pipeline. The use of multistate registers in the CFMT microarchitecture enables higher performance processors (40% average performance improvement) with relatively low energy (6.5% average energy reduction) and area overhead. Ravi Patel 0001, Shahar Kvatinsky, Eby G. Friedman, Avinoam Kolodny |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Experimental Analysis of Thermal Coupling in 3-D Integrated CircuitsabstractA 3-D test circuit examining thermal propagation within a through-silicon via-based 3-D integrated stack has been designed, fabricated, and tested. Design insight into thermal coupling in 3-D integrated circuits (ICs) through both experiment and simulation is provided, and suggestions to mitigate thermal effects in 3-D ICs are offered. Two wafers are vertically bonded to form a 3-D stack. Intraplane and interplane thermal coupling is investigated through single-point heat generation using resistive thermal heaters and temperature monitoring through four-point resistive measurements. Thermal paths are identified and analyzed based on the metric of thermal resistance per unit length. The peak steady-state temperature due to die location within a 3-D stack is described. The reduction in peak temperature through fan-based active cooling is also reported. Thermal propagation from a heat source located on the backside of the silicon is examined with both back metal and on-chip thermal sensors. A comparison of thermal coupling between two different heat sources on the same device plane is also provided. Ioannis Savidis, Boris Vaisband, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Field driven STT-MRAM cell for reduced switching latency and energyabstractA field driven approach to STT-MRAM switching is proposed as a method for reducing the switching latency of an MTJ in high performance caches. An MRAM array model is presented to characterize the switching energy and maximum achievable reduction in energy using the field driven approach. The switching latency per bit is reduced by more than a factor of ten. The resultant switching energy per bit is reduced by 82% as compared to a standard STT-MRAM. Ravi Patel 0001, Engin Ipek, Eby G. Friedman |
ISCAS | 3 |
| 2014 | Computationally efficient clustering of power supplies in heterogeneous real time systemsabstractHigh quality power delivery for on-chip high performance integrated circuits is a significant design challenge in modern functionally diverse systems with multiple power domains. To provide a high quality power delivery system with dynamically changing voltages and transient currents, the onchip power needs to be regulated in real time, within each power domain. To exploit the advantages of existing switching and linear power supplies, a heterogeneous power delivery system has recently been proposed based on the principle of separation of power conversion and regulation, lowering the overall energy loss while requiring small on-chip area. The power efficiency of the system is shown to be a strong function of the clustering of the power supplies - the specific configuration in which power converters and regulators are co-designed. A recursive clustering algorithm with polynomial computational complexity is proposed for an optimal real time power distribution system with minimum power losses. The proposed algorithm is evaluated on IBM power grid benchmark circuits and two multi-power domain circuits, yielding up to a 21% increase in power efficiency, and orders of magnitude speedup in runtime with the proposed recursive clustering algorithm. Inna Partin-Vaisband, Eby G. Friedman |
ISCAS | 2 |
| 2014 | Thermal conduction path analysis in 3-D ICsabstractThe on-going effort of integrating heterogeneous circuits as well as the increasing length of global interconnect are driving the semiconductor community towards 3-D integrated circuits. In this work, thermal paths within a 3-D stack are investigated using the HotSpot simulator, and the results are compared to experimental data of a fabricated two layer stack with a single back metal layer. Resistive heaters and sensors measure the heat flow in both the horizontal and vertical dimensions. The dependence of the thermal conductivity on temperature is integrated into the thermal simulation process. At high temperatures (~ 80°C), this effect is responsible for inaccuracies in the temperature and thermal resistance of up to, respectively, 20% and 28%. As confirmed by simulation, those horizontal paths that lie mostly within the silicon layer conduct more heat as compared to the vertical paths, since the thermal conductivity of silicon dioxide is ~ 200 times smaller than the thermal conductivity of silicon. Boris Vaisband, Ioannis Savidis, Eby G. Friedman |
ISCAS | 3 |
| 2014 | Memristor-Based Material Implication (IMPLY) Logic: Design Principles and MethodologiesabstractMemristors are novel devices, useful as memory at all hierarchies. These devices can also behave as logic circuits. In this paper, the IMPLY logic gate, a memristor-based logic circuit, is described. In this memristive logic family, each memristor is used as an input, output, computational logic element, and latch in different stages of the computing process. The logical state is determined by the resistance of the memristor. This logic family can be integrated within a memristor-based crossbar, commonly used for memory. In this paper, a methodology for designing this logic family is proposed. The design methodology is based on a general design flow, suitable for all deterministic memristive logic families, and includes some additional design constraints to support the IMPLY logic family. An IMPLY 8-bit full adder based on this design methodology is presented as a case study. Shahar Kvatinsky, Guy Satat, Nimrod Wald, Eby G. Friedman, Avinoam Kolodny, Uri C. Weiser |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2014 | Digitally Controlled Pulse Width Modulator for On-Chip Power ManagementabstractA digitally controlled current starved pulse width modulator (PWM) is described in this paper. The current from the power grid to the ring oscillator is controlled by a header circuit. By changing the header current, the pulse width of the switching signal generated at the output of the ring oscillator is dynamically controlled, permitting the duty cycle to vary between 25% and 90%. A duty cycle to voltage converter is used to ensure the accuracy of the system under process, voltage, and temperature (PVT) variations. A ring oscillator with two header circuits is proposed to control both duty cycle and frequency of the operation. Analytic closed-form expressions for the operation of a PWM are provided. The accuracy and performance of the proposed PWM is evaluated with 22-nm CMOS predictive technology models under PVT variations. An error of less than 3.1% and 4.4% in the duty cycle, respectively, with and without constant frequency control is reported for the PWM. A constant operation frequency with less than 1.25% period variation is demonstrated. The proposed PWM is appropriate for dynamic voltage scaling systems due to the small on-chip area and high accuracy under PVT variations. Inna Partin-Vaisband, Mahmood J. Azhar, Eby G. Friedman, Selçuk Köse |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | AC-DIMM: associative computing with STT-MRAMabstractWith technology scaling, on-chip power dissipation and off-chip memory bandwidth have become significant performance bottlenecks in virtually all computer systems, from mobile devices to supercomputers. An effective way of improving performance in the face of bandwidth and power limitations is to rely on associative memory systems. Recent work on a PCM-based, associative TCAM accelerator shows that associative search capability can reduce both off-chip bandwidth demand and overall system energy. Unfortunately, previously proposed resistive TCAM accelerators have limited flexibility: only a restricted (albeit important) class of applications can benefit from a TCAM accelerator, and the implementation is confined to resistive memory technologies with a high dynamic range (RHigh/RLow), such as PCM. Qing Guo 0004, Ravi Patel 0001, Engin Ipek, Eby G. Friedman |
ISCA | 5 |
| 2013 | Current profile of a microcontroller to determine electromagnetic emissionsabstractA methodology is proposed to determine the current profile early in the design process to accurately estimate electromagnetic emissions. Design information describing the clock and power distribution network topologies and the placement and sizing of the decoupling capacitors is used to determine the current signatures of individual circuit blocks and the entire system. The proposed methodology is incorporated within the integrated circuit emission model (ICEM). Current profiles for various circuits with different characteristics are determined using the proposed model. Selçuk Köse, Eby G. Friedman, Radu M. Secareanu, Olin L. Hartin |
ISCAS | 2 |
| 2013 | Digitally controlled wide range pulse width modulator for on-chip power suppliesabstractA digitally controlled current starved pulse width modulator is described in this paper. The current from the power grid to the ring oscillator is controlled by a header circuit. By changing the header current, the pulse width of the switching signal generated at the output of the ring oscillator is dynamically controlled, permitting the duty cycle to vary between 50% and 90%. A duty cycle to voltage converter is used to ensure the accuracy of the system under process, voltage, and temperature (PVT) variations. The accuracy and performance of the proposed digitally controlled pulse width modulator is evaluated with 22 nm CMOS predictive technology models under PVT variations. The proposed pulse width modulator is appropriate for dynamic voltage scaling systems due to the small on-chip area and high accuracy under process, voltage, and temperature variations. Although the frequency of the switching signal is affected by changes in the duty cycle, the frequency variations are typically negligible. Selçuk Köse, Inna Partin-Vaisband, Eby G. Friedman |
ISCAS | 3 |
| 2013 | Timing-driven variation-aware synthesis of hybrid mesh/tree clock distribution networks
Ameer Abdelhadi, Ran Ginosar, Avinoam Kolodny, Eby G. Friedman |
Integr. | 4 |
| 2013 | Power Network Optimization Based on Link Breaking MethodologyabstractA link breaking methodology is introduced to reduce voltage degradation within mesh structured power distribution networks. The resulting power distribution network combines a single power distribution network to lower the network impedance, and multiple networks to reduce noise coupling among the circuits. Since the sensitivity to supply voltage variations within a power distribution network can vary among various circuits, the proposed methodology reduces the voltage drop at the more sensitive circuits, while penalizes the less sensitive circuits. Each circuit can behave as an aggressor as well as a victim. The methodology utilizes two matrices describing the aggressiveness and sensitivity of a circuit. The proposed methodology is evaluated for multiple case studies, demonstrating a reduction in the voltage drop in the sensitive circuits. Based on these case studies, the voltage is improved by 5% at those nodes with the highest sensitivity. The voltage prior to application of the link breaking methodology is 96% of the ideal power supply voltage. Lowering the noise on the power network enhances the maximum operating frequency by 16% by utilizing the proposed link breaking methodology. The link breaking methodology has also been compared with a multiple voltage domain methodology, achieving 7% improvement in operating frequency. Renatas Jakushokas, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Active Filter-Based Hybrid On-Chip DC-DC Converter for Point-of-Load Voltage RegulationabstractAn active filter-based on-chip DC-DC voltage converter for application to distributed on-chip power supplies in multivoltage systems is described in this paper. No inductor or output capacitor is required in the proposed converter. The area of the voltage converter is therefore significantly less than that of a conventional low-dropout (LDO) regulator. Hence, the proposed circuit is appropriate for point-of-load voltage regulation for noise sensitive portions of an integrated circuit. The performance of the circuit has been verified with Cadence Spectre simulations and fabricated with a commercial 110 nm complimentary metal oxide semiconductor (CMOS) technology. The area of the voltage regulator is 0.015 mm2and delivers up to 80 mA of output current. The transient response with no output capacitor ranges from 72 to 192 ns. The parameter sensitivity of the active filter is also described. The advantages and disadvantages of the active filter-based, conventional switching, linear, and switched capacitor voltage converters are compared. The proposed circuit is an alternative to classical LDO voltage regulators, providing a means for distributing multiple local power supplies across an integrated circuit while maintaining high current efficiency and fast response time within a small area. Selçuk Köse, Simon M. Tam, Sally Pinzon, Bruce McDermott, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2012 | Link breaking methodology: mitigating noise within power networksabstractA link breaking methodology is introduced to reduce voltage degradation within mesh structured power distribution networks. The resulting power distribution network combines a single power distribution network to lower the network impedance, and multiple networks to reduce noise coupling among the circuits. Since the sensitivity to supply voltage variations within a power distribution network can vary among different circuits, the proposed methodology reduces the voltage drop at the more sensitive circuits, while penalizes the less sensitive circuits. The proposed methodology is evaluated for two case studies, demonstrating a reduction in the voltage drop in sensitive circuits. Based on these case studies, the voltage is improved by, on average, 4% at those nodes with the highest sensitivity. The voltage after application of the link breaking methodology is, on average, 96% of the ideal power supply voltage. Lowering the noise on the power network enhances, on average, the maximum operating frequency by 11% by utilizing the proposed link breaking methodology. Renatas Jakushokas, Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 2 |
| 2012 | Energy metrics for power efficient crosslink and mesh topologiesabstractClock distribution networks are an essential element of a synchronous digital circuit, a significant power consumer and highly sensitive to process, voltage, and temperature variations. Mesh- and crosslink-based topologies reliably compensate for skew variations in these networks, albeit with a significant increase in dissipated power as compared to variation-sensitive low power clock trees. Existing crosslink-based methods, however, only address skew from an algorithmic perspective at the network topology level. Guidelines for inserting crosslinks within a buffered low power clock tree are provided in this paper. Physical constraints, such as the size of the crosslink and exact location between the driving and load buffers, are analytically described. Metrics to determine the most energy efficient non-tree topology are provided based on closed-form expressions, and verified with simulation. Inna Partin-Vaisband, Eby G. Friedman, Ran Ginosar, Avinoam Kolodny |
ISCAS | 2 |
| 2012 | Arithmetic encoding for memristive multi-bit storage
Ravi Patel 0001, Eby G. Friedman |
VLSI-SoC | 2 |
| 2012 | Efficient algorithms for fast IR drop analysis exploiting locality
Selçuk Köse, Eby G. Friedman |
Integr. | 2 |
| 2011 | Fast algorithms for IR voltage drop analysis exploiting localityabstractClosed form expressions and related algorithms for fast power grid analysis are proposed in this paper. The IR voltage drop at an arbitrary point in a power distribution network is determined. Two algorithms are described for non-uniform voltage supplies and non-uniform current loads distributed throughout a power grid. The principle of spatial locality is exploited to accelerate the proposed power grid analysis method. Analysis of the non-uniform power grids utilizes the principle of spatial locality. Since no iterations are required for the proposed IR drop analysis, the proposed algorithms are over 70 times faster for smaller power grids composed of less than five million nodes and over 180 times faster for larger power grids composed of more than 25 million nodes as compared to existing methods. The proposed method exhibits less than 0.5% error. Selçuk Köse, Eby G. Friedman |
DAC | 2 |
| 2011 | Memristor-based IMPLY logic design procedureabstractMemristors can be used as logic gates. No design methodology exists, however, for memristor-based combinatorial logic. In this paper, the design and behavior of a memristive-based logic gate - an IMPLY gate - are presented and design issues such as the tradeoff between speed (fast write times) and correct logic behavior are described, as part of an overall design methodology. A memristor model is described for determining the write time and state drift. It is shown that the widely used memristor model - a linear ion drift memristor - is impractical for characterizing an IMPLY logic gate, and a different memristor model is necessary such as a memristor with a current threshold. Shahar Kvatinsky, Avinoam Kolodny, Uri C. Weiser, Eby G. Friedman |
ICCD | 4 |
| 2011 | Clock distribution models of 3-D integrated systemsabstractClock distribution topologies in a three-tier 3-D integrated circuit are explored. Models of three different clock topologies are applied to determine the root to leaf delay. The models incorporate the impedance of the 3-D via between planes based on closed-form expressions of the resistance, inductance, and capacitance of a through silicon via (TSV). The resulting modeled delays are compared to experimental data. Good agreement between simulation and experimental data is achieved. Ioannis Savidis, Vasilis F. Pavlidis, Eby G. Friedman |
ISCAS | 3 |
| 2011 | Multi-Layer Interdigitated Power Distribution NetworksabstractHigher operating frequencies and greater power demands have increased the requirements on the power and ground network. Simultaneously, due to the larger current loads, current densities are increasing, making electromigration an important design issue. In this paper, methods for optimizing a multi-layer interdigitated power and ground network are presented. Based on the resistive and inductive (both self- and mutual) impedance, a closed-form solution for determining the optimal power and ground wire width is described, producing the minimum impedance for a single metal layer. Electromigration is considered, permitting the appropriate number of metal layers to be determined. The tradeoff between the network impedance and current density is investigated. Based on 65-, 45-, and 32-nm CMOS technologies, the optimal width as a function of metal layer is determined for different frequencies, suggesting important trends for interdigitated power and ground networks. Renatas Jakushokas, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | Shielding Methodologies in the Presence of Power/Ground NoiseabstractDesign guidelines for shielding in the presence of power/ground (P/G) noise are presented in this paper. The effect of P/G noise on crosstalk is analyzed for different line lengths, line widths, and interconnect driver resistances. Considering the P/G noise, a shield line can degrade rather than enhance signal integrity due to increased P/G noise coupling on the victim line. A$2\pi$RLC interconnect model is used to investigate the effects of both coupling capacitance and mutual inductance on the crosstalk noise. Physical spacing and shield insertion are compared in terms of the coupling noise on the victim line for several technology nodes. Boundary conditions are also provided to determine the effective range of spacing and shield insertion in the presence of P/G noise. Additionally, the effects of technology scaling on P/G noise and shielding efficiency are discussed, and related design tradeoffs are addressed. Selçuk Köse, Emre Salman, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Clock Distribution Networks in 3-D Integrated Systemsabstract3-D integration is an important technology that addresses fundamental limitations in on-chip interconnects. Several design issues related to 3-D circuits, such as multiplane synchronization, however, need to be addressed. A comparison of three 3-D clock distribution network topologies is presented in this paper. Good agreement is shown between the modeled and experimental results of a 3-D test circuit composed of three device planes. Successful operation of the 3-D test circuit at 1.4 GHz is demonstrated. Clock skew, clock delay, signal slew, and power dissipation measurements for the different clock topologies are also provided. The measurements suggest that each topology provides certain advantages and disadvantages in terms of different performance criteria. The proper choice, consequently, of a clock distribution network is not dictated by a single design objective but rather by the overall 3-D system design requirements including availability of resources and number of bonded planes. Vasilis F. Pavlidis, Ioannis Savidis, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | A Distributed Filter Within a Switching Converter for Application to 3-D Integrated CircuitsabstractA design methodology for distributing a buck converter filter for application to 3-D circuits is described. The 3-D filter exploits transmission line properties, permitting the generation and distribution of power supplies to different planes. As compared to a conventional LC filter, the proposed filter only requires on-chip capacitors without the use of on-chip inductors. Additionally, the physical structure of the filter simultaneously enables the distribution of the current to the load while filtering the switching signal at the input. A case study in a 0.18- μm CMOS 3-D technology demonstrates the generation of a 1.2 V power supply delivering 700 mA peak current. Jonathan Rosenfeld, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | Linear and Switch-Mode Conversion in 3-D CircuitsabstractA methodology for DC-DC conversion in three-dimensional (3-D) circuits is described in this paper. The proposed approach exploits both linear and switching buck converters with different conversion ranges, thereby increasing power efficiency. By replacing the traditional LC filter within a switching converter with a distributed filter, a significant increase in efficiency is demonstrated. Additionally, the physical structure of the filter simultaneously enables the distribution of high current to the load while filtering the switching signal at the input. Design guidelines and expressions are developed, achieving good agreement with simulations based on the MIT Lincoln Laboratories CMOS/SOI 180-nm 3-D technology. The proposed converter distributes 2.5 A maximum current, achieving conversion from 3.3 V to 2.5 V and 1 V with, respectively, 74% and 44% power efficiency. Jonathan Rosenfeld, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Timing-driven variation-aware nonuniform clock mesh synthesisabstractClock skew variations adversely affect timing margins, limiting performance, reducing yield, and may also lead to functional faults. Non-tree clock distribution networks, such as meshes and crosslinks, are employed to reduce skew and also to mitigate skew variations. However, these networks incur an increase in dissipated power while consuming significant metal resources. Several methods have been proposed to trade off power and wires to reduce skew. In this paper, an efficient algorithm is presented to reduce skew variations rather than skew, and prioritize the algorithm for critical timing paths, since these paths are more sensitive to skew variations. The algorithm has been implemented for a standard 65 nm cell library using standard EDA tools, and has been tested on several benchmark circuits. As compared to other methods, experimental results show a 37% average reduction in metal consumption and 39% average reduction in power dissipation, while insignificantly increasing the maximum skew. Ameer Abdelhadi, Ran Ginosar, Avinoam Kolodny, Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 4 |
| 2010 | Line width optimization for interdigitated power/ground networksabstractHigher operating frequencies have increased the importance of inductance in power and ground networks. The effective inductance of the power and ground network can be reduced with an interdigitated structure. A closed-form solution for determining the optimal power and ground wire width is described, producing the minimum impedance within an interdigitated structure. The optimal wire width is determined under different physical network dimensions and signal frequencies, suggesting useful trends for interdigitated power and ground networks. In addition, a closed-form expression for the optimal wire width is determined that minimizes the voltage drop over a single metal layer of an interdigitated power/ground network. Renatas Jakushokas, Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | On-chip point-of-load voltage regulator for distributed power suppliesabstractAn ultra-low area, current efficient voltage regulator appropriate for distributed point-of-load voltage regulation in high performance integrated circuits (ICs) is described in this paper. The proposed voltage regulator is a hybrid combination of a switching voltage regulator and a linear voltage regulator. The voltage regulator can supply over 100 mA current while generating 0.9 volts from a 1.2 input voltage. The current efficiency exceeds 99% while the load regulation to a step current ranges between 40 ns and 60 ns. No output capacitor is required to ensure stability. Hence, the required on-chip area is as small as 0.026 mm^2 for the proposed voltage regulator which is approximately four to six times smaller than area efficient low dropout regulators. The proposed circuit therefore provides a means for distributing multiple local power supplies across an integrated circuit, maintaining high current efficiency and small area. Selçuk Köse, Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | Methodology to achieve higher tolerance to delay variations in synchronous circuitsabstractA methodology is proposed for designing robust circuits exhibiting higher tolerance to process and environmental variations. This higher tolerance is achieved by exploiting the interdependence between the setup and hold times, reducing the delay uncertainty caused by variations. An algorithm is proposed to determine the interdependent setup-hold pair of a register. A data path designed with the proposed setup-hold pair improves the overall tolerance to variations. The methodology is evaluated for several technologies to determine the overall reduction in delay uncertainty. Emre Salman, Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | An intra-chip free-space optical interconnectabstractContinued device scaling enables microprocessors and other systems-on-chip (SoCs) to increase their performance, functionality, and hence, complexity. Simultaneously, relentless scaling, if uncompensated, degrades the performance and signal integrity of on-chip metal interconnects. These systems have therefore become increasingly communications-limited. The communications-centric nature of future high performance computing devices demands a fundamental change in intra- and inter-chip interconnect technologies. Alok Garg, Berkehan Ciftcioglu, Jianyun Hu, Ioannis Savidis, Rebecca Berman, Peng Liu 0016, Michael C. Huang 0001, Hui Wu 0007, Eby G. Friedman, Gary Wicks, Duncan Moore |
ISCA | 12 |
| 2010 | Globally integrated power and clock distribution networkabstractThe global networks within a conventional integrated circuits (IC) consists of three major types: power, ground, and clock networks. These three networks consumes most of the metal resources in the highest metal layers. The signals traversing the power and clock distribution networks are fundamentally different in terms of signal frequency and current flow. Combining the power and clock network into a globally integrated network is therefore possible. In this paper, the general concept of a globally integrated power and clock (GIPAC) system is proposed. The circuitry supporting this GIPAC system is also presented. Simulation results based on a 90 nm CMOS technology demonstrate the potential of GIPAC. Renatas Jakushokas, Eby G. Friedman |
ISCAS | 2 |
| 2010 | Methodology for multi-layer interdigitated power and ground network designabstractHigher operating frequencies and greater power demands have increased the requirements on the power and ground network. Simultaneously, due to the larger current loads, current densities are increasing, making electromigration an important design issue. The optimal wire width for an interdigitated power and ground network is based on the resistive and inductive (both self- and mutual) impedance. In this paper, a methodology for optimizing a multi-layer interdigitated power and ground network is presented, reducing the current density and impedance of a network. Based on 65 nm, 45 nm, and 32 nm CMOS technologies, the optimal width as a function of metal layer is determined for different frequencies, suggesting important trends for interdigitated power and ground networks. Renatas Jakushokas, Eby G. Friedman |
ISCAS | 2 |
| 2010 | Compact substrate models for efficient noise coupling and signal isolation analysisabstractCurrent propagation within a lightly doped substrate is approximated with a half-ellipse to efficiently estimate substrate resistances. As opposed to existing work, the proposed model contains only one fitting parameter. Compact models are also developed to determine the isolation efficiency of several commonly used structures such as a guard ring and triple well. The accuracy of these models is verified by comparing the models with a commercial substrate extraction tool based on a boundary element method. These models are used to compare several isolation structures within an industrial mixed-signal circuit with a lightly doped substrate. Renatas Jakushokas, Emre Salman, Eby G. Friedman, Radu M. Secareanu, Olin L. Hartin, Cynthia L. Recker |
ISCAS | 3 |
| 2010 | An area efficient fully monolithic hybrid voltage regulatorabstractA hybrid voltage regulator module for an on-chip DC-DC voltage converter is proposed in this paper. The circuit is appropriate for point-of-load voltage regulation due to an ultra area efficient architecture. The proposed voltage regulator is a hybrid combination of a switching DC-DC voltage converter and a low-dropout regulator exploiting active circuitry rather than bulky passive devices within the filter structure. The proposed circuit can supply over 100 mA current while generating 0.9 volts from a 1.2 input voltage, exhibiting a high current efficiency of greater than 99%. The on-chip area is 0.026 mm2which is 500 times smaller than a monolithic buck converter and four times smaller than an LDO. The proposed regulator provides a means for distributing multiple local power supplies across an integrated circuit while providing high current efficiency. Selçuk Köse, Eby G. Friedman |
ISCAS | 2 |
| 2010 | Fast algorithms for power grid analysis based on effective resistanceabstractThe size of on-chip power distribution networks is increasing with each technology generation. Accurate and computationally efficient analysis of these power distribution networks has therefore become increasingly challenging. High performance power distribution networks are generally implemented as a uniform mesh structure. The uniformity of these power distribution networks can be exploited for fast, accurate nodal analysis. A closed form expression is presented here for determining the voltage at any arbitrary node in a power distribution network. The error of the proposed method as compared with SPICE is less than 0.2%. Since no iterations are required, the proposed method significantly outperforms previously proposed power grid analysis techniques in terms of computational speed while exhibiting low error. Selçuk Köse, Eby G. Friedman |
ISCAS | 2 |
| 2010 | Resource Based Optimization for Simultaneous Shield and Repeater InsertionabstractA new approach for resource based optimization for high performance integrated circuits is presented. The methodology is applied to simultaneous shield and repeater insertion, resulting in minimum coupling noise under power, delay, and area constraints. Design expressions exhibiting parabolic noise behavior are compared with SPICE simulations. Due to the parabolic coupled noise behavior, the minimum noise is established. A design case is compared with only shielding and only repeater insertion techniques, exhibiting enhanced performance for different resources. Renatas Jakushokas, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Unified Logical Effort - A Method for Delay Evaluation and Minimization in Logic Paths With RC InterconnectabstractThe unified logical effort (ULE) model for delay evaluation and minimization in paths composed of CMOS logic gates and resistive wires is presented. The method provides conditions for timing optimization while overcoming the limitations of standard logical effort (LE) in the presence of interconnects. The condition for optimal gate sizing in a logic path with long wires is also presented. This condition is achieved when the delay component due to the gate input capacitance is equal to the delay component due to the effective output resistance of the gate. The ULE delay model unifies the problems of gate sizing and repeater insertion: In the case of negligible interconnect, the ULE method converges to standard LE optimization, yielding tapered gate sizes. In the case of long wires, the solution converges toward uniform sizing of gates and repeaters. The technique is applied to various types of logic paths to demonstrate the influence of wire length, gate type, and technology. Arkadiy Morgenshtein, Eby G. Friedman, Ran Ginosar, Avinoam Kolodny |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Corrections to "Unified Logical Effort - A Method for Delay Evaluation and Minimization in Logic Paths With RC Interconnect" [May 10 689-696]abstractIn the above titled paper (ibid., vol. 18, no. 5, pp. 689-696, May 10), the formula and the caption in Fig. 3 appeared incorrectly. The correct figure is presented here along with an explanation. Arkadiy Morgenshtein, Eby G. Friedman, Ran Ginosar, Avinoam Kolodny |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Design challenges in high performance three-dimensional circuitsabstractThe initial focus of the presentation will be on reviewing the fundamental trends specific to 3-D circuits and systems, including the many opportunities and challenges of this exciting new technology. A short review of the MIT Lincoln Laboratories 3-D manufacturing technology will follow. A summary of some primary issues in the physical design of 3-D systems will be reviewed. This discussion will be followed by a review of current research in the area of on-chip 3-D computer network topologies; specifically, 3-D networks-on-chip. A discussion of the so-called Rochester Cube will then be presented in the context of its relative impact and importance. Circuit design issues will be discussed and experimental results will be reviewed. The presentation will conclude with a review of some near-term and long term research problems in 3-D systems. Categories & Subject Descriptors: B.7 INTEGRATED CIRCUITS, B.7.1 Types and Design Styles, Advanced technologies, Algorithms implemented in hardware, Microprocessors and microcomputers, VLSI (very large scale integration) General Terms: Performance, Design, Reliability Bio Eby G. Friedman received the B.S. degree from Lafayette College in 1979, and the M.S. and Ph.D. degrees from the University of California, Irvine, in 1981 and 1989, respectively, all in electrical engineering. From 1979 to 1991, he was with Hughes Aircraft Company. He has been with the Department of Electrical and Computer Engineering at the University of Rochester since 1991, where he is a Distinguished Professor, and the Director of the High Performance VLSI/IC Design and Analysis Laboratory. He is also a Visiting Professor at the Technion Israel Institute of Technology. His current research and teaching interests are in high performance synchronous digital and mixed-signal microelectronic circuit design. He is the author of more than 325 papers and book chapters, several patents, and the author or editor of ten books in the fields of high speed and low power CMOS design techniques, high speed interconnect, and the theory and application of synchronous clock and power distribution networks. He previously was the Editor-in-Chief of the IEEE Transactions on Very Large Scale Integration (VLSI) Systems, and a recipient of the University of Rochester Graduate Teaching Award, and a College of Engineering Teaching Excellence Award. Dr. Friedman is a Senior Fulbright Fellow and an IEEE Fellow. Copyright is held by the author/owner(s). GLSVLSI’09, May 10–12, 2009, Boston, Massachusetts, USA. ACM 978-1-60558-522-2/09/05. Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 1 |
| 2009 | Simultaneous shield and repeater insertionabstractResource based optimization for high performance integrated circuits is presented. The methodology is applied to simultaneous shield and repeater insertion, resulting in minimum coupling noise under power, delay, and area constraints. Design expressions exhibiting parabolic noise behavior are compared with SPICE simulations. Due to the parabolic coupled noise behavior, the minimum noise is established. Good agreement between the analytic results and SPICE simulations is shown. Renatas Jakushokas, Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 2 |
| 2009 | Contact merging algorithm for efficient substrate noise analysis in large scale circuitsabstractA methodology is proposed to efficiently estimate the substrate noise generated by large scale aggressor circuits. Small spatial voltage differences within the ground distribution network of an aggressor circuit are exploited to reduce the overall number of input ports before the substrate extraction process. Specifically, the substrate of an aggressor circuit is partitioned into voltage domains where each domain is represented by a single substrate contact. The remaining ports of the substrate within that domain are ignored to reduce the computational complexity. A linear time algorithm is developed to identify these voltage domains and generate an equivalent contact. A reduction of more than four orders of magnitude in the number of extracted substrate resistances is demonstrated while introducing 20% error in the peak-to-peak value of the substrate noise voltage. Emre Salman, Renatas Jakushokas, Eby G. Friedman, Radu M. Secareanu, Olin L. Hartin |
ACM Great Lakes Symposium on VLSI | 3 |
| 2009 | Power efficient tree-based crosslinks for skew reductionabstractClock distribution networks are an important design issue that is highly dependent on delay variations and load imbalances, while requiring power efficiency. Existing mesh solutions significantly increase the dissipated power, whereas existing link based methods only address skew caused by variations and do not consider power consumption. The power dissipated by the inserted crosslinks within a buffered clock tree is investigated in this paper, and is shown to be a strong function of the resistance and capacitance of the crosslink. A crosslink may be power efficient despite the presence of short-circuit currents caused by multiple drivers in a non-tree clock network. The power characteristics of crosslink size and placement are also discussed, showing that the crosslink is best placed as close as possible to the target leaves of the tree. Crosslink insertion as both an alternative and complement to buffer sizing for low power skew reduction is also considered. Inna Partin-Vaisband, Ran Ginosar, Avinoam Kolodny, Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 4 |
| 2009 | Minimizing Noise Via Shield and Repeater InsertionabstractTwo techniques, shield and repeater insertion, are simultaneously investigated. Based on resource optimization, the relationship among noise, power, and delay is investigated. Coupling noise as a function of power dissipation is shown to behave parabolically. Due to this parabolic behavior, the minimum noise can be established. The resulting design expressions are compared with SPICE simulations, exhibiting good agreement. A design case is compared with only shielding and only repeater insertion techniques, exhibiting enhanced performance for different resources. Renatas Jakushokas, Eby G. Friedman |
ISCAS | 2 |
| 2009 | Shielding Methodologies in the Presence of Power/Ground NoiseabstractDesign guidelines for shielding in the presence of power/ground (P/G) noise are presented in this paper. The effect of noise in the P/G network is analyzed for various line lengths, line widths, and interconnect driver resistances. A 2pi RLC model is used to investigate the effect of both coupling capacitance and mutual inductance on the crosstalk noise. For a range of shield lengths and widths, a shield line can degrade signal integrity by increasing the crosstalk noise on the victim line. Different physical spacing and shield insertion methods are compared for various parameters in terms of the coupling noise on the victim line for a 65 nm technology node. Selçuk Köse, Emre Salman, Eby G. Friedman |
ISCAS | 3 |
| 2009 | Interconnect-Based Design Methodologies for Three-Dimensional Integrated CircuitsabstractDesign techniques for three-dimensional (3-D) ICs considerably lag the significant strides achieved in 3-D manufacturing technologies. Advanced design methodologies for two-dimensional circuits are not sufficient to manage the added complexity caused by the third dimension. Consequently, design methodologies that efficiently handle the added complexity and inherent heterogeneity of 3-D circuits are necessary. These 3-D design methodologies should support robust and reliable 3-D circuits while considering different forms of vertical integration, such as system-in-package and 3-D ICs with fine grain vertical interconnections. Global signaling issues, such as clock and power distribution networks, are further exacerbated in vertical integration due to the limited number of package pins, the distance of these pins from other planes within the 3-D system, and the impedance characteristics of the through silicon vias (TSVs). In addition to these dedicated networks, global signaling techniques that incorporate the diverse traits of complex 3-D systems are required. One possible approach, potentially significantly reducing the complexity of interconnect issues in 3-D circuits, is 3-D networks-on-chip (NoC). Design methodologies that exploit the diversity of 3-D structures to further enhance the performance of multiplane integrated systems are necessary. The longest interconnects within a 3-D circuit are those interconnects comprising several TSVs and traversing multiple physical planes. Consequently, minimizing the delay of the interplane nets is of great importance. By considering the nonuniform impedance characteristics of the interplane interconnects while placing the TSVs, the delay of these nets is decreased. In addition, the difference in electrical behavior between the horizontal and vertical interconnects suggests that asymmetric structures can be useful candidates for distributing the clock signal within a 3-D circuit. A 3-D test circuit fabricated with a 180 nm silicon-on-insulator (SOI) technology, manufactured by MIT Lincoln Laboratories, exploring several clock distribution topologies is described. Correct operation at 1 GHz has been demonstrated. Several 3-D NoC topologies incorporating dissimilar 3-D interconnect structures are reviewed as a promising solution for communication limited systems-on-chip (SoC). Appropriate performance models are described to evaluate these topologies. Several forms of vertical integration, such as system-in-package and different candidate technologies for 3-D circuits, such as SOI, are considered. The techniques described in this paper address fundamental interconnect structures in the 3-D design process. Several interesting research problems in the design of 3-D circuits are also discussed. Vasilis F. Pavlidis, Eby G. Friedman |
Proc. IEEE | 2 |
| 2009 | Quasi-Resonant Interconnects: A Low Power, Low Latency Design MethodologyabstractDesign and analysis guidelines for quasi-resonant interconnect networks (QRN) are presented in this paper. The methodology focuses on developing an accurate analytic distributed model of the on-chip interconnect and inductor to obtain both low power and low latency. Excellent agreement is shown between the proposed model and SpectraS simulations. The analysis and design of the inductor, insertion point, and driver resistance for minimum power-delay product is described. A case study demonstrates the design of a quasi-resonant interconnect, transmitting a 5 Gb/s data signal along a 5 mm line in a TSMC 0.18-mum CMOS technology. As compared to classical repeater insertion, an average reduction of 91.1% and 37.8% is obtained in power consumption and delay, respectively. As compared to optical links, a reduction of 97.1% and 35.6% is observed in power consumption and delay, respectively. Jonathan Rosenfeld, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Identification of Dominant Noise Source and Parameter Sensitivity for Substrate CouplingabstractA simple, yet physically intuitive macrolevel model is presented to identify the dominant substrate coupling mechanism at the early stages of the design process, while simultaneously considering multiple parameters. Furthermore, the sensitivity of substrate noise to these parameters is evaluated, demonstrating the nonmonotonic dependence of noise on rise time. The design implications of the proposed analysis are discussed, identifying the preferred noise reduction technique for a specific set of operating points. Emre Salman, Eby G. Friedman, Radu M. Secareanu, Olin L. Hartin |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Methodology for Efficient Substrate Noise Analysis in Large-Scale Mixed-Signal CircuitsabstractA methodology is proposed to efficiently analyze substrate noise coupled to a sensitive block due to an aggressor digital block in large-scale mixed-signal circuits. The methodology is based on identifyingvoltage domainson the substrate by exploiting the small spatial voltage differences on the ground distribution network of the aggressor circuit. Specifically, similarly biased regions on the substrate short-circuited by the ground network are determined, and each of these regions is represented by a single equivalent input port to the substrate. The remaining ports within that domain are ignored to reduce the computational complexity of the extraction process. An algorithm with linear time complexity is proposed to merge those substrate contacts exhibiting a voltage difference smaller than a specified value, identifying a voltage domain. An equivalent contact is placed at the geometric mean of the merged contacts, ignoring all of the remaining ports such as the source/drain junctions of the devices. The ground network impedance is updated for each merged contact based on the proposed algorithm to maintain sufficient accuracy of the noise voltage. The substrate with reduced input ports is extracted using an existing extraction tool to analyze the noise at the sense node. As compared to the full extraction of an aggressor circuit, the methodology achieves a reduction of more than four orders of magnitude in the number of extracted substrate resistors with a peak-to-peak error of 24%. Emre Salman, Renatas Jakushokas, Eby G. Friedman, Radu M. Secareanu, Olin L. Hartin |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2008 | Transient simulation of on-chip transmission lines via exact pole extractionabstractAn accurate and efficient solution for the transient response at the far end of a transmission line is proposed in this paper. Unlike approximating the poles by truncating the transfer function or matching moments, the exact poles of an interconnect system are analytically extracted. Excellent match is observed between the proposed method and Spectre simulations. With two pairs of poles, the average error for the 50% delay is 1%. Higher accuracy can be obtained with additional pairs of poles. The computational complexity of the model is proportional to the number of pole pairs. Eby G. Friedman |
ISCAS | 2 |
| 2008 | Equivalent rise time for resonance in power/ground noise estimationabstractThe non-monotonic behavior of power/ground noise with respect to the rise time tr is investigated for an inductive power distribution network with a decoupling capacitor. A time domain solution is provided for the rise time that produces resonant behavior, thereby maximizing the power/ground noise. The sensitivity of the ground noise to the decoupling capacitance Q and parasitic inductance Lgis evaluated as a function of the rise time. Increasing the decoupling capacitance is shown to efficiently reduce the noise for trles 2radic(LgCd). Alternatively, reducing the parasitic inductance Lgis shown to be effective for trges 2radic(LgCd). Emre Salman, Eby G. Friedman, Radu M. Secareanu, Olin L. Hartin |
ISCAS | 2 |
| 2008 | Input port reduction for efficient substrate extraction in large scale IC'sabstractA methodology is proposed to improve the efficiency of the substrate impedance extraction process for a large scale circuit by exploiting the circuit activity. Similarly biased regions of the substrate short-circuited by the ground network are identified to reduce the computational complexity of the extraction process. Each of these voltage domains is represented by a single equivalent input port to the substrate, merging the remaining ports within that domain. An algorithm is presented to determine these domains and generate an equivalent port for each domain. The parasitic impedance of the ground network is updated to maintain accuracy. A reduction of more than two orders of magnitude in the number of extracted substrate resistances is demonstrated while introducing 15% error in the rms value of the substrate noise voltage at the sense node. Emre Salman, Renatas Jakushokas, Eby G. Friedman, Radu M. Secareanu, Olin L. Hartin |
ISCAS | 3 |
| 2008 | Electrical modeling and characterization of 3-D viasabstractElectrical characterization of the resistance, capacitance, and inductance of inter-plane 3-D vias is presented in this paper. Both capacitive and inductive coupling between multiple 3-D vias is described as a function of the separation distance and plane location. The effects of placing a third shield via between two signal vias is investigated as a means to limit the capacitive coupling. The location of the return path is examined to determine the best placement of a 3-D via to reduce the overall loop inductance. Based on the extracted resistance, capacitance, and inductance, the L/R time constant is shown to be much larger than the RC time constant, demonstrating that the 3-D via structure investigated in this paper is inductively limited rather than capacitively limited. Ioannis Savidis, Eby G. Friedman |
ISCAS | 2 |
| 2008 | Timing-driven via placement heuristics for three-dimensional ICs
Vasilis F. Pavlidis, Eby G. Friedman |
Integr. | 2 |
| 2008 | Efficient Distributed On-Chip Decoupling Capacitors for Nanoscale ICsabstractA distributed on-chip decoupling capacitor network is proposed in this paper. A system of distributed on-chip decoupling capacitors is shown to provide an efficient solution for providing the required on-chip decoupling capacitance under existing technology constraints. In a system of distributed on-chip decoupling capacitors, each capacitor is sized based on the parasitic impedance of the power distribution grid. Various tradeoffs in a system of distributed on-chip decoupling capacitors are also discussed. Related simulation results for typical values of on-chip parasitic resistance are also presented. The worst case error is 0.003% as compared to SPICE. Mikhail Popovich, Eby G. Friedman, Radu M. Secareanu, Olin L. Hartin |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | On-Chip Power Distribution Grids With Multiple Supply Voltages for High-Performance Integrated CircuitsabstractOn-chip power distribution grids with multiple supply voltages are discussed in this paper. Two types of interdigitated and paired power distribution grids with multiple supply voltages and multiple grounds are presented. Analytic models are also developed to estimate the loop inductance in four types of proposed power delivery schemes. Two proposed schemes, fully and pseudo-interdigitated power delivery, reduce power supply voltage drops as compared to conventional interdigitated power distribution systems with dual supplies and a single ground by, on average, 15.3% and 0.3%, respectively. The performance of the proposed on-chip power distribution grids is compared to a reference power distribution grid with a single supply and a single ground. The voltage drop in fully interdigitated and fully paired power distribution grids with multiple supplies and multiple grounds is reduced, on average, by 2.7% and 2.3%, respectively, as compared to the voltage drop of an interdigitated power distribution grid with a single supply and a single ground. The proposed power distribution grids are a better alternative to a single supply voltage and a single ground power distribution system. On-chip resonances in power distribution grids with decoupling capacitors are intuitively explained in this paper, and circuit design implications are provided. It is also noted that fully interdigitated and fully paired power distribution grids with multiple supply voltages and multiple grounds are recommended to decouple power supply voltages. Mikhail Popovich, Eby G. Friedman, Michael Sotman, Avinoam Kolodny |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | Effective Radii of On-Chip Decoupling CapacitorsabstractDecoupling capacitors are widely used to reduce power supply noise. On-chip decoupling capacitors have traditionally been allocated into the white space available on a die or placed inside the rows in standard cell circuit blocks. The efficacy of on-chip decoupling capacitors depends upon the impedance of the power/ground lines connecting the capacitors to the current loads and power supplies. A design methodology for placing on-chip decoupling capacitors is presented in this paper. A maximum effective radius is shown to exist for each on-chip decoupling capacitor. Beyond this effective distance, a decoupling capacitor is ineffective. Depending upon the parasitic impedance of the power distribution system, the maximum voltage drop seen at the current load is caused either by the first droop (determined by the rise time) or by the second droop (determined by the transition time). Two criteria to estimate the minimum required on-chip decoupling capacitance are developed based on the critical parasitic impedance. In order to provide the required charge drawn by the load, the decoupling capacitor has to be charged before the next switching cycle. For an on-chip decoupling capacitor to be effective, both effective radii criteria should be simultaneously satisfied. Mikhail Popovich, Michael Sotman, Avinoam Kolodny, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2007 | Efficient placement of distributed on-chip decoupling capacitors in nanoscale ICsabstractDecoupling capacitors are widely used to reduce power supply noise. On-chip decoupling capacitors have traditionally been allocated into the white space available on the die based on an unsystematic or ad hoc approach. In this way, large decoupling capacitors are often placed at a significant distance from the current load, compromising the signal integrity of the system. This issue of power delivery cannot be alleviated by simply increasing the size of the on-chip decoupling capacitors. To be effective, the on-chip decoupling capacitors should be placed physically close to the current loads. The area occupied by the on-chip decoupling capacitor, however, is directly proportional to the magnitude of the capacitor. The minimum impedance between the on-chip decoupling capacitor and the current load is therefore fundamentally affected by the magnitude of the capacitor. A distributed on-chip decoupling capacitor network is proposed in this paper. A system of distributed on-chip decoupling capacitors is shown to provide an efficient solution for providing the required on-chip decoupling capacitance under existing technology constraints. In a system of distributed on-chip decoupling capacitors, each capacitor is sized based on the parasitic impedance of the power distribution grid. Various tradeoffs in a system of distributed on-chip decoupling capacitors are also discussed. Related simulation results for typical values of on-chip parasitic resistance are also presented. An analytic solution is shown to provide accurate distributed system. The worst case error is 0.003% as compared to SPICE. Techniques presented in this paper are applicable not only for current technologies, but also provide an efficient placement of the on-chip decoupling capacitors in future technology generations. Mikhail Popovich, Eby G. Friedman, Radu M. Secareanu, Olin L. Hartin |
ICCAD | 2 |
| 2007 | Quasi-Resonant Interconnects: A Low Power Design MethodologyabstractDesign and analysis guidelines for resonant interconnect networks are presented in this paper. The methodology focuses on developing an accurate analytic distributed model of the on-chip interconnect and inductor to obtain low power and low latency. Excellent agreement is shown between the proposed model and SpectraS simulations. The analysis and design of the inductance, the insertion point, and the driver resistance for minimum power consumption is described. A case study demonstrates the design of a resonant interconnect, transmitting a 5 Gbps data signal along a 5 mm line in a TSMC 0.18 μm CMOS technology. As compared to classical repeater insertion, an average reduction of 94.8% and 72.8% is obtained in power consumption and delay, respectively. As compared to optical links, a reduction of 98.5% and 60% is observed in power consumption and delay, respectively. Jonathan Rosenfeld, Eby G. Friedman |
ISCAS | 2 |
| 2007 | Substrate Noise Reduction Based On Noise Aware Cell DesignabstractA substrate biasing methodology is introduced based on modifying standard cells by inserting dedicated substrate contacts in those cells behaving as aggressive digital noise generators. These contacts are connected to a dedicated ground network. The proposed approach reduces two primary noise injection mechanisms: ground coupling and source/drain junction coupling. Limitations of the Kelvin biasing scheme are removed while achieving more than a 60% (9 dB) reduction in substrate noise at the cost of a 12% increase in area. Emre Salman, Eby G. Friedman, Radu M. Secareanu, Olin L. Hartin |
ISCAS | 2 |
| 2007 | Predictions of CMOS compatible on-chip optical interconnect
Mikhail Haurylau, Nicholas Nelson 0001, David H. Albonesi, Philippe M. Fauchet, Eby G. Friedman |
Integr. | 7 |
| 2007 | Wire shaping of RLC interconnects
Magdy A. El-Moursy, Eby G. Friedman |
Integr. | 2 |
| 2007 | Exploiting Setup-Hold-Time Interdependence in Static Timing AnalysisabstractA methodology is proposed to exploit the interdependence between setup- and hold-time constraints in static timing analysis (STA). The methodology consists of two phases. The first phase includes the interdependent characterization of sequential cells, resulting in multiple constraint pairs. The second phase includes an efficient algorithm that exploits these multiple pairs in STA. The methodology improves accuracy by removing optimism and reducing unnecessary pessimism. Furthermore, the tradeoff between setup and hold times is exploited to significantly reduce timing violations in STA. These benefits are validated using industrial circuits and tools, exhibiting up to 53% reduction in the number of constraint violations as well as up to 48% reduction in the worst negative slack, which corresponds to a 15% decrease in the clock period Emre Salman, Ali Dasdan, Feroze Taraporevala, Kayhan Küçükçakar, Eby G. Friedman |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2007 | 3-D Topologies for Networks-on-ChipabstractSeveral interesting topologies emerge by incorporating the third dimension in networks-on-chip (NoC). The speed and power consumption of 3D NoC are compared to that of 2D NoC. Physical constraints, such as the maximum number of planes that can be vertically stacked and the asymmetry between the horizontal and vertical communication channels of the network, are included in speed and power consumption models of these novel 3D structures. An analytic model for the zero-load latency of each network that considers the effects of the topology on the performance of a 3D NoC is developed. Tradeoffs between the number of nodes utilized in the third dimension, which reduces the average number of hops traversed by a packet, and the number of physical planes used to integrate the functional blocks of the network, which decreases the length of the communication channel, is evaluated for both the latency and power consumption of a network. A performance improvement of 40% and 36% and a decrease of 62% and 58% in power consumption is demonstrated for 3D NoC as compared to a traditional 2D NoC topology for a network size of N = 128 and N = 256 nodes, respectively. Vasilis F. Pavlidis, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2007 | Design Methodology for Global Resonant H-Tree Clock Distribution NetworksabstractDesign guidelines for resonant H-tree clock distribution networks are presented in this paper. A distributed model of a two-level resonant H-tree structure is described, supporting the design of low power, skew, and jitter resonant H-tree clock distribution networks. Excellent agreement is shown between the proposed model and SpectraS simulations. A case study is presented that demonstrates the design of a two-level resonant H-tree network, distributing a 5-GHz clock signal in a 0.18-mum CMOS technology. This example exhibits an 84% decrease in power dissipation as compared to a standard H-tree clock distribution network. The design methodology enables tradeoffs among design variables to be examined, such as the operating frequency, the size of the on-chip inductors and capacitors, the output resistance of the driving buffer, and the interconnect width. A sensitivity analysis of resonant H-tree clock distribution networks is also provided. The effect of the driving buffer output resistance, on-chip inductor and capacitor size, and signal and shielding transmission line width and spacing on the output voltage swing and power consumption is described Jonathan Rosenfeld, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | Maximum effective distance of on-chip decoupling capacitors in power distribution gridsabstractDecoupling capacitors are widely used to reduce power supply noise. On-chip decoupling capacitors have traditionally been allocated into the available white space on a die. The efficacy of on-chip decoupling capacitors depends upon the impedance of the power/ground lines connecting the capacitors to the current loads and power supplies. A maximum effective radius exists for each on-chip decoupling capacitor. Beyond this effective distance, a decoupling capacitor is completely ineffective. Two effective radii determined by the target impedance (during discharge) and charge time are presented in this paper. Depending upon the parasitic impedance of the power distribution system, the maximum voltage drop as seen at the current load is achieved either at the first droop or at the end of the switching activity (the second droop). Two criteria to estimate the minimum required on-chip decoupling capacitance are developed based on the critical parasitic impedance. To be effective, the decoupling capacitor has to be fully charged before the next switching event. A design space is described that characterizes the tolerable parasitic resistances and inductances, while restoring the charge on the decoupling capacitor within a target charge time. An overall design methodology for placing on-chip decoupling capacitors is presented in this paper. It is shown that for an on-chip decoupling capacitor to be effective, both effective radii criteria should be simultaneously satisfied. Mikhail Popovich, Eby G. Friedman, Michael Sotman, Avinoam Kolodny, Radu M. Secareanu |
ACM Great Lakes Symposium on VLSI | 2 |
| 2006 | Sensitivity evaluation of global resonant H-tree clock distribution networksabstractA sensitivity analysis of resonant H-tree clock distribution networks is presented in this paper for a TSMC 0.18 ?m CMOS technology. The analysis focuses on the effect of the driving buffer output resistance, on-chip inductor and capacitor size, and signal and shielding transmission line width and spacing on the output voltage swing and power consumption. A two level resonant H-tree network exhibits low sensitivity to these variations. Jonathan Rosenfeld, Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 2 |
| 2006 | Effective capacitance of RLC loads for estimating short-circuit powerabstractAn effective capacitance of a distributed RLC load for estimating short-circuit power is presented in this paper. Both resistive and inductive shielding effects of interconnects are considered and no iterations are required to determine the effective capacitance. The proposed method has been verified with Cadence Spectre. For a single switching input, the average error of the short-circuit power obtained with the effective capacitance is less than 2% for the example circuits as compared with an RLC /spl pi/ model. The proposed method can be used in look-up table or k-factor based models to estimate short-circuit power dissipation in CMOS gates with complex interconnects. Eby G. Friedman |
ISCAS | 2 |
| 2006 | Optimum wire tapering for minimum power dissipation in RLC interconnectsabstractThe optimum tapered structure for RLC interconnect to minimize transient power dissipation is determined. Wire tapering can reduce the power dissipated by a circuit by up to 72% as compared to uniform wire sizing. An analytic solution to determine the optimum tapered structure exhibits an error of less than 2% as compared to SPICE Magdy A. El-Moursy, Eby G. Friedman |
ISCAS | 2 |
| 2006 | Via placement for minimum interconnect delay in three-dimensional (3D) circuitsabstractThe propagation delay of interlayer 3D interconnects is investigated in this paper. For RC interconnects connecting two circuits located on different physical planes, the interconnect delay is minimized by optimally placing the non-stacked interlayer vias. The problem of determining this optimum via locations under the Elmore delay model is described as a geometric program. Simulations indicate delay improvements of up to 26% for relatively short interconnect. The proposed approach is also compared with a wire sizing algorithm. Timing-driven via placement exhibits better results both in terms of delay and power consumption Vasilis F. Pavlidis, Eby G. Friedman |
ISCAS | 2 |
| 2006 | Design methodology for global resonant H-tree clock distribution networksabstractDesign guidelines for resonant H-tree clock distribution networks are presented in this paper. A distributed model of a two level resonant H-tree is presented, supporting the design of low power, low skew, and low jitter resonant H-tree clock distribution networks. Excellent agreement is shown between the proposed model and SpectraS simulations. A case study is presented that demonstrates the design of a two level resonant H-tree network, distributing a 5 GHz clock signal in a TSMC 0.18 mum CMOS technology. The design methodology enables tradeoffs among design variables to be examined, such as the operating frequency, size of the on-chip inductors and capacitors, the output resistance of the driving buffer, and the interconnect width Jonathan Rosenfeld, Eby G. Friedman |
ISCAS | 2 |
| 2006 | On-die decoupling capacitance: frequency domain analysis of activity radiusabstractOn-die capacitances interact with the inductance and resistance of the power distribution network to supply electrical charge. A distributed model is generally required to analyze and design a power distribution network to maintain acceptable levels of supply voltage noise. An approximation method is proposed in this paper for modeling the power distribution system in the frequency domain, defining an effective decoupling capacitance and an effective decoupling radius around each switching element, both of which are frequency dependent. At high frequencies, the supply of decoupling charge is highly localized, and the effective decoupling capacitance is determined primarily by the power grid inductance. Design considerations and guidelines are presented for determining the appropriate density of on-die decoupling capacitances and required power grid parameters, depending upon the density and the frequency characteristics of the switching activity of the circuit Michael Sotman, Avinoam Kolodny, Mikhail Popovich, Eby G. Friedman |
ISCAS | 4 |
| 2006 | Low-power repeaters driving RC and RLC interconnects with delay and bandwidth constraintsabstractInterconnect plays an increasingly important role in deep-submicrometer very large scale integrated technologies. Multiple design criteria are considered in interconnect design, such as delay, power, and bandwidth. In this paper, a repeater insertion methodology is presented for achieving the minimum power in an RC interconnect while satisfying delay and bandwidth constraints. These constraints determine a design space for the number and size of the repeaters. The minimum power is shown to occur at the edge of the design space. With delay constraints, closed form solutions for the minimum power are developed, where the average error is 7% as compared with SPICE. With bandwidth constraints, the minimum power can be achieved with minimum-sized repeaters. The effects of inductance on the delay, bandwidth, and power of an RLC interconnect with repeaters are also analyzed. By including inductance, the minimum interconnect power under a delay or bandwidth constraint decreases as compared with an RC interconnect. Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | Decoupling capacitors for multi-voltage power distribution systemsabstractMultiple power supply voltages are often used in modern high-performance ICs, such as microprocessors, to decrease power consumption without affecting circuit speed. To maintain the impedance of a power distribution system below a specified level, multiple decoupling capacitors are placed at different levels of the power grid hierarchy. The system of decoupling capacitors used in power distribution systems with multiple power supplies is described in this paper. The noise at one power supply can propagate to the other power supply, causing power and signal integrity problems in the overall system. With the introduction of a second power supply, therefore, the interaction between the two power distribution networks should be considered. The dependence of the impedance and magnitude of the voltage transfer function on the parameters of the power distribution system is investigated. An antiresonance phenomenon is intuitively explained in this paper. It is shown that the magnitude of the voltage transfer function is strongly dependent on the parasitic inductance of the decoupling capacitors, decreasing with smaller inductance. Design techniques to cancel and shift antiresonant spikes out of range of the operating frequencies are presented. It is also shown that it is highly desirable to maintain the effective series inductance of the decoupling capacitors as low as possible to decrease the overshoots of the response of the dual-voltage power distribution system over a wide range of operating frequencies. A criterion for an overshoot-free voltage response is presented in this paper. It is noted that the frequency range of the overshoot-free voltage response can be traded off with the magnitude of the response. Mikhail Popovich, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | Crosstalk modeling for coupled RLC interconnects with application to shield insertionabstractOn-chip interconnect delay and crosstalk noise have become significant bottlenecks in the performance and signal integrity of deep submicrometer VLSI circuits. A crosstalk noise model for both identical and nonidentical coupled resistance-inductance-capacitance (RLC) interconnects is developed based on a decoupling technique exhibiting an average error of 6.8% as compared to SPICE. The crosstalk noise model, together with a proposed concept of effective mutual inductance, is applied to evaluate the effectiveness of the shielding technique. Junmou Zhang, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Interconnect delay minimization through interlayer via placement in 3-D ICsabstractThe dependence of the propagation delay of the interlayer 3-D interconnects on the vertical through via location and length is investigated. For a variable vertical through via location, with fixed vertical length, the optimum vertical through via location that minimizes the propagation delay of an interconnect line connecting two circuits on different planes is determined. The optimum vertical through via location and length or, equivalently, the number of physical planes traversed by the vertical through via, are determined for varying the placement of the connected circuits. Design expressions for the optimal via locations and lengths have been developed to support placement and routing algorithms for 3-D ICs. Vasilis F. Pavlidis, Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 2 |
| 2005 | On-chip power distribution grids with multiple supply voltages for high performance integrated circuitsabstractMultiple supply voltages are often utilized to decrease power dissipation in high performance integrated circuits. On-chip power distribution grids with multiple supply voltages are discussed in this paper. A power distribution grid with multiple supply voltages and multiple grounds is presented. The proposed power delivery scheme reduces power supply voltage drops as compared to conventional power distribution systems with dual supplies and a single ground by 17% on average (20% maximum). For an example power grid with decoupling capacitors placed between the power supply and ground, the proposed grid with multiple supply and multiple ground exhibits, respectively, 13% and 18% average performance improvement. The proposed power distribution grid can be an alternative to a single supply voltage and single ground power distribution system. Mikhail Popovich, Eby G. Friedman, Michael Sotman, Avinoam Kolodny |
ACM Great Lakes Symposium on VLSI | 2 |
| 2005 | An RLC interconnect model based on fourier analysisabstractBased on a Fourier series analysis, an analytic interconnect model is presented which is suitable for periodic signals, such as a clock signal. In this model, the far-end time-domain waveform is approximated by the summation of several sinusoids. Closed-form solutions of the 50% delay and overshoots/undershoots are provided when the fifth and higher order harmonics are ignored. Good accuracy is observed between the model and SPICE simulations. The model is applied to resistance-capacitance-inductance interconnect trees and the computational complexity of the model is linear with the size of the tree and the model order. The tree model is shown to be an effective method to analyze clock distribution networks. The single interconnect model is also extended to coupled multi-interconnect systems to analyze crosstalk noise and a general waveform solution is obtained. It is noted that in addition to the transition time, the period of the aggressor signal also has a significant effect on the crosstalk noise. Eby G. Friedman |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2005 | Shielding effect of on-chip interconnect inductanceabstractInterconnect inductance introduces a shielding effect which decreases the effective capacitance seen by the driver of a circuit, reducing the gate delay. A model of the effective capacitance of an RLC load driven by a CMOS inverter is presented. The interconnect inductance decreases the gate delay and increases the time required for the signal to propagate across an interconnect, reducing the overall delay to drive an RLC load. Ignoring the line inductance overestimates the circuit delay, inefficiently oversizing the circuit driver. Considering line inductance in the design process saves gate area, reducing dynamic power dissipation. Average reductions in power of 17% and area of 29% are achieved for example circuits. An accurate model for a CMOS inverter and an RLC load is used to characterize the propagation delay. The accuracy of the delay model is within an average error of less than 9% as compared to SPICE. Magdy A. El-Moursy, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Exponentially tapered H-tree clock distribution networksabstractExponentially tapered interconnect can reduce the dynamic power dissipation of clock distribution networks. A criterion for sizing H-tree clock networks is proposed. The technique reduces the power dissipated for an example clock network by up to 15% while preserving the signal transition times and propagation delays. Furthermore, the inductive behavior of the interconnects is reduced, decreasing the inductive noise. Exponentially tapered interconnects decrease by approximately 35% the difference between the overshoots in the signal at the input of a tree. As compared to a uniform tree with the same area overhead, overshoots in the signal waveform at the source of the tree are reduced by 40%. Magdy A. El-Moursy, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2004 | Optimum wire sizing of RLC interconnect with repeaters
Magdy A. El-Moursy, Eby G. Friedman |
Integr. | 2 |
| 2004 | Power characteristics of inductive interconnectabstractThe width of an interconnect line affects the total power consumed by a circuit. The effect of wire sizing on the power characteristics of an inductive interconnect line is presented in this paper. The matching condition between the driver and the load affects the power consumption since the short-circuit power dissipation may decrease and the dynamic power will increase with wider lines. A tradeoff, therefore, exists between short-circuit and dynamic power in inductive interconnects. The short-circuit power increases with wider linewidths only if the line is underdriven. The power characteristics of inductive interconnects therefore may have a great influence on wire sizing optimization techniques. An analytic solution of the transition time of a signal propagating along an inductive interconnect with an error of less than 15% is presented. The solution is useful in wire sizing synthesis techniques to decrease the overall power dissipation. The optimum linewidth that minimizes the total transient power dissipation is determined. An analytic solution for the optimum width with an error of less than 6% is presented. For a specific set of line parameters and resistivities, a reduction in power approaching 80% is achieved as compared to the minimum wire width. Considering the driver size in the design process, the optimum wire and driver size that minimizes the total transient power is also determined. Magdy A. El-Moursy, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2004 | Sleep switch dual threshold Voltage domino logic with reduced standby leakage currentabstractA circuit technique is presented for reducing the subthreshold leakage energy consumption of domino logic circuits. Sleep switch transistors are proposed to place an idle dual threshold voltage domino logic circuit into a low leakage state. The circuit technique enhances the effectiveness of a dual threshold voltage CMOS technology to reduce the subthreshold leakage current by strongly turning off all of the high threshold voltage transistors. The sleep switch circuit technique significantly reduces the subthreshold leakage energy as compared to both standard low-threshold voltage and dual threshold voltage domino logic circuits. A domino adder enters and leaves a low leakage sleep mode within a single clock cycle. The energy overhead of the circuit technique is low, justifying the activation of the proposed sleep scheme by providing a net savings in total power consumption during short idle periods. Volkan Kursun, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2004 | Scaling trends of on-chip power distribution noiseabstractThe design of power distribution networks in high-performance integrated circuits has become significantly more challenging with recent advances in process technologies. As on-chip currents exceed tens of amperes and circuit clock periods are reduced well below a nanosecond, the signal integrity of on-chip power supply has become a primary concern in the integrated circuit design. The scaling behavior of the inductive and resistance voltage drops across the on-chip power distribution networks is the subject of this paper. The existing work on power distribution noise scaling is reviewed and extended to include the scaling behavior of the inductance of the on-chip global power distribution networks in high-performance flip-chip packaged integrated circuits. As the dimensions of the on-chip devices are scaled by S, where S>1, the resistive voltage drop across the power grids remains constant and the inductive voltage drop increases by S, if the metal thickness is maintained constant. Consequently, the signal-to-noise ratio decreases by S in the case of resistive noise and by S/sup 2/ in the case of inductive noise. As compared to the constant metal thickness scenario, ideal interconnect scaling of the global power grid mitigates the unfavorable scaling of the inductive noise but exacerbates the scaling of resistive noise by a factor of S. On-chip inductive noise will, therefore, become of greater significance with technology scaling. Careful tradeoffs between the resistance and inductance of the power distribution networks will be necessary in nanometer technologies to achieve minimum power supply noise. Andrey V. Mezhiba, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2004 | Impedance characteristics of power distribution grids in nanoscale integrated circuitsabstractThe essential design characteristic of nanoscale integrated circuits is increased interconnect complexity. Conductors at different levels of the interconnect hierarchy have highly different physical and, consequently, electrical characteristics. These interconnect lines also exhibit inductive behavior due to enhanced switching speed of nanoscale devices, making interconnect design and analysis difficult. The design of robust and area efficient power distribution networks for high-speed integrated circuits has therefore become a challenging task. The impedance characteristics of multilayer power distribution grids and the relevant design implications are the subject of this paper. The power distribution network spans many layers of interconnect with disparate electrical properties. Unlike single-layer grids, the electrical characteristics of multilayer grids vary significantly with frequency. As the frequency increases, a large share of the current flow is transfered from the low-resistance upper layers to the low-inductance lower layers. The inductance of a multilayer grid therefore decreases with frequency, while the resistance increases with frequency. The lower layers of multilayer power grids provide a low-inductance current path, significantly reducing the grid impedance at high frequencies. Multilayer power distribution grids extend to the lower interconnect layers, exhibiting superior high-frequency impedance characteristics as compared to power distribution grids built exclusively within the upper, low-resistance metal layers. A significant share of metal resources to distribute the global power should therefore be allocated to the lower metal layers. An analytic model is also presented to determine the impedance characteristics of a multilayer grid from the inductive and resistive properties of the comprising individual grid layers. Andrey V. Mezhiba, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2004 | Substrate coupling in digital circuits in mixed-signal smart-power systemsabstractThis paper describes theoretical and experimental data characterizing the sensitivity of nMOS and CMOS digital circuits to substrate coupling in mixed-signal, smart-power systems. The work presented here focuses on the noise effects created by high-power analog circuits and affecting sensitive digital circuits on the same integrated circuit. The sources and mechanism of the noise behavior of such digital circuits are identified and analyzed. The results are obtained primarily from a set of dedicated test circuits specifically designed, fabricated, and evaluated for this work. The conclusions drawn from the theoretical and experimental analyses are used to develop physical and circuit design techniques to mitigate the substrate noise problems. These results provide insight into the noise immunity of digital circuits with respect to substrate coupling. Radu M. Secareanu, Scott Warner, Scott Seabridge, Cathie Burke, Juan Becerra, Thomas E. Watrobski, Christopher Morton, William Staub, Thomas Tellier, Ivan S. Kourtev, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 11 |
| 2003 | Reduced Delay Uncertainty in High Performance Clock Distribution Networks
Dimitrios Velenis, Marios C. Papaefthymiou, Eby G. Friedman |
DATE | 3 |
| 2003 | Orthogonal code generator for 3G wireless transceiversabstractOrthogonal variable spreading factor (OVSF) codes are standard in third generation UMTS cellular systems. The efficient generation of these codes is essential for reducing the area and power of wireless transceivers. In this paper, the basic properties of this family of codes are analyzed from an RTL perspective and two efficient hardware code generators are proposed. Tradeoffs and design solutions as well as low power considerations are discussed. These results represent the first reported implementation of an OVSF code generator. Boris D. Andreev, Edward L. Titlebaum, Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 3 |
| 2003 | Optimum wire sizing of RLC interconnect with repeatersabstractRepeaters are often used to drive high impedance interconnects. These lines have become highly inductive which can affect signal behavior in long interconnects. The line inductance should, therefore, be considered in determining the optimum number and size of the repeaters driving a line. A tradeoff exists, however, between the transient power dissipation and the minimum propagation delay in sizing long interconnects driven by repeaters. Optimizing the line width to achieve the minimum power delay product, however, can satisfy current high speed, low power design objectives. A reduction in power of 65% and delay of 97% is achieved for an example repeater system. Magdy A. El-Moursy, Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 2 |
| 2003 | Shielding effect of on-chip interconnect inductanceabstractInterconnect inductance introduces a shielding effect which decreases the effective capacitance seen by the driver of a circuit, reducing the gate delay. The effective capacitance of an RLC load driven by a CMOS inverter is analytically modeled. The interconnect inductance decreases the gate delay and increases the time required for the signal to propagate across an interconnect, reducing the overall signal propagation delay to drive an RLC load. Ignoring the line inductance overestimates the circuit delay, inefficiently oversizing the circuit driver. Considering line inductance in the design process saves gate area, thereby reducing the dynamic power dissipation. A reduction in power of 17% and area of 29% is achieved for an example circuit. Magdy A. El-Moursy, Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 2 |
| 2003 | Domino logic with variable threshold voltage keeperabstractA variable threshold voltage keeper circuit technique is proposed for simultaneous power reduction and speed enhancement of domino logic circuits. The threshold voltage of a keeper transistor is dynamically modified during circuit operation to reduce contention current without sacrificing noise immunity. The variable threshold voltage keeper circuit technique enhances circuit evaluation speed by up to 60% while reducing power dissipation by 35% as compared to a standard domino (SD) logic circuit. The keeper size can be increased with the proposed technique while preserving the same delay or power characteristics as compared to a SD circuit. The proposed domino logic circuit technique offers 14% higher noise immunity as compared to a SD circuit with the same evaluation delay characteristics. Forward body biasing the keeper transistor is also proposed for improved noise immunity as compared to a SD circuit with the same keeper size. It is shown that by applying forward and reverse body biased keeper circuit techniques, the noise immunity and evaluation speed of domino logic circuits are simultaneously enhanced. Volkan Kursun, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | Analysis of buck converters for on-chip integration with a dual supply voltage microprocessorabstractAn analysis of an on-chip buck converter is presented in this paper. A high switching frequency is the key design parameter that simultaneously permits monolithic integration and high efficiency. A model of the parasitic impedances of a buck converter is developed. With this model, a design space is determined that allows integration of active and passive devices on the same die for a target technology. An efficiency of 88.4% at a switching frequency of 477 MHz is demonstrated for a voltage conversion from 1.2-0.9 volts while supplying 9.5 A average current. The area occupied by the buck converter is 12.6 mm/sup 2/ assuming an 80-nm CMOS technology. An estimate of the efficiency is shown to be within 2.4% of simulation at the target design point. Full integration of a high-efficiency buck converter on the same die with a dual-V/sub DD/ microprocessor is demonstrated to be feasible. Volkan Kursun, Siva G. Narendra, Vivek De, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2002 | Efficient implementation of a complex ±1 multiplierabstractA complex ±1 multiplier is an integral element in modern CDMA communication systems, specifically as a pseudonoise code scrambler/descrambler. Therefore, an efficient implementation is essential to reduce the critical path delay, power, and area of wireless receivers. A new architecture is proposed to achieve this complex multiplier function. Tradeoffs and design solutions as well as the interface with subsequent arithmetic circuits are discussed. Simulations exhibit a significant speed improvement as compared to alternative architectures. These results are also applicable to other arithmetic circuits. Boris D. Andreev, Eby G. Friedman, Edward L. Titlebaum |
ACM Great Lakes Symposium on VLSI | 2 |
| 2002 | Low swing dual threshold voltage domino logicabstractA low swing domino logic technique is proposed to decrease power consumption without sacrificing noise immunity. With the proposed low swing domino logic circuit technique, active power consumption is reduced by up to 9.4% while improving the noise immunity by 2.6% as compared to standard domino logic circuits. It is also shown that by applying a low swing contention reduction technique, the power savings can be further increased by 6.7% while the delay can be improved by 8.6%. A simple and efficient dual threshold voltage (dual-Vt) circuit technique that incorporates low swing signals is also proposed. It is shown that the proposed dual-Vt technique reduces the standby leakage current by approximately 235 times while offering enhanced delay characteristics as compared to a standard low threshold voltage implementation. Volkan Kursun, Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 2 |
| 2002 | Properties of on-chip inductive current loopsabstractThe variation of inductance with circuit length is investigated in this paper. The nonlinear variation of inductance with length is shown to be a result of inductive coupling among circuit segments. If the distance between the forward and return current paths of a current loop is much smaller than the loop length, the inductive coupling to the forward current is similar to the coupling to the return current, resulting in negligible coupling. The inductance of these circuits therefore varies approximately linearly with length. Similarly, the effective inductive coupling between two parallel current loops is reduced through cancellation and has a negligible effect on the net inductance of a circuit. As a general rule, the inductance of circuits where the distance between the forward and return current is much smaller than the characteristic dimensions of the circuit scales linearly with circuit dimensions. This linear behavior can be used to simplify the inductance extraction and circuit analysis process. Andrey V. Mezhiba, Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 2 |
| 2002 | Managing static leakage energy in microprocessor functional unitsabstractStatic energy due to subthreshold leakage current is projected to become a major component of the total energy in high performance microprocessors. Many studies so far have examined and proposed techniques to reduce leakage in on-chip storage structures. In this study, static energy is reduced in the integer functional units by leveraging the unique qualities of dual threshold voltage domino logic. Domino logic has desirable properties that greatly reduce leakage current while providing fast propagation times. However due to the energy cost of entering the low leakage current state (sleep mode), domino logic has thus far been used only for leakage reduction in the longterm standby mode. We examine the utility of the sleep mode (while considering the aforementioned costs) when idle times are relatively short, one to a few hundred cycles, as is often the case for functional units. Using an analytical energy model suitable for architecture-level analysis, we explore the interaction of the application and technology, and the effect on energy and performance as the underlying parameters are varied, on a set of benchmarks. Our results show that if the leakage approaches the magnitude as projected in the literature, even for short idle intervals as few as ten cycles, an aggressive policy of activating the sleep mode at every idle period performs well and a more complex control strategy may not be warranted. We also propose a simple design, called Gradual Sleep, to reduce the energy impact of using the sleep mode for smaller idle periods. Steven G. Dropsho, Volkan Kursun, David H. Albonesi, Sandhya Dwarkadas, Eby G. Friedman |
MICRO | 5 |
| 2002 | DTT: direct truncation of the transfer function - an alternative tomoment matching for tree structured interconnectabstractA method is introduced to evaluate time domain signals within RLC trees with arbitrary accuracy in response to any input signal. This method depends on finding a low frequency reduced-order transfer function by direct truncation of the exact transfer function at different nodes of an RLC tree. The method is numerically accurate for any order of approximation, which permits approximations to be determined with a large number of poles appropriate for approximating RLC trees with underdamped responses. The method is computationally efficient with a complexity linearly proportional to the number of branches in an RLC tree. A common set of poles is determined that characterizes the responses at all of the nodes of an RLC tree which further enhances the computational efficiency. Stability is guaranteed by the DTT method for low-order approximations with less than five poles. Such low-order approximations are useful for evaluating monotone responses exhibited by RC circuits. Yehea I. Ismail, Eby G. Friedman |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2002 | Retiming and clock scheduling for digital circuit optimizationabstractThis paper investigates the application of simultaneous retiming and clock scheduling for optimizing synchronous circuits under setup and hold constraints. Two optimization problems are explored: (1) clock period minimization and (2) tolerance maximization to clock-signal delay variations. Exact mixed-integer linear programming formulations and efficient heuristics are given for both problems. When both long and short paths are considered, circuits optimized by the combined application of retiming and clock scheduling can achieve shorter clock periods or demonstrate greater tolerance to clock-signal delay variations than circuits optimized by retiming or clock scheduling. Experiments with benchmark circuits demonstrate the effectiveness of the combined optimization. In comparison with the best result obtained by either of the two optimizations, the joint application of retiming and clock scheduling increased operating speeds by more than 8% on the average. It also increased tolerance to clock delay variations by an average of 12% over a broad range of target clock frequencies. Larger relative improvements were achieved for shorter clock periods, thus suggesting that simultaneous retiming and clock scheduling can play an important role in high-speed design. Marios C. Papaefthymiou, Eby G. Friedman |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2002 | Inductive properties of high-performance power distribution gridsabstractThe design of high integrity, area efficient power distribution grids has become of practical importance as the portion of on-chip interconnect resources dedicated to power distribution networks in high performance integrated circuits has greatly increased. The inductive characteristics of several types of gridded power distribution networks are described in this paper. The inductance extraction program FastHenry is used to evaluate the inductive properties of grid structured interconnect. In power distribution grids with alternating power and ground lines, the inductance is shown to vary linearly with grid length and inversely linearly with the number of lines in the grid. The inductance is also relatively constant with frequency in these grid structures. These properties permit the efficient estimation of the inductive characteristics of power distribution grids. To optimize the process of allocating on-chip metal resources, inductance/area/resistance tradeoffs in high speed performance distribution grids are explored. Two tradeoff scenarios in power grids with alternating power and ground lines are considered. Andrey V. Mezhiba, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2002 | Simultaneous switching noise in on-chip CMOS power distribution networksabstractSimultaneous switching noise (SSN) has become an important issue in the design of the internal on-chip power distribution networks in current very large scale integration/ultra large scale integration (VLSI/ULSI) circuits. An inductive model is used to characterize the power supply rails when a transient current is generated by simultaneously switching the on-chip registers and logic gates in a synchronous CMOS VLSI/ULSI circuit. An analytical expression characterizing the SSN voltage is presented here based on a lumped inductive-resistive-capacitive RLC model. The peak value of the SSN voltage based on this analytical expression is within 10% as compared to SPICE simulations. Design constraints at both the circuit and layout levels are also discussed based on minimizing the effects of the peak value of the SSN voltage. Kevin T. Tang, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2001 | Clock distribution networks in synchronous digital integrated circuitsabstractClock distribution networks synchronize the flow of data signals among synchronous data paths. The design of these networks can dramatically affect system-wide performance and reliability. A theoretical background of clock skew is provided in order to better understand how clock distribution networks interact with data paths. Minimum and maximum timing constraints are developed from the relative timing between the localized clock skew and the data paths. These constraint relationships are reviewed, and compensating design techniques are discussed. The field of clock distribution network design and analysis can be grouped into a number of subtopics: 1) circuit and layout techniques for structured custom digital integrated circuits; 2) the automated layout and synthesis of clock distribution networks with application to automated placement and routing of gate arrays, standard cells and larger block-oriented circuits; 3) the analysis and modeling of the timing characteristics of clock distribution networks; and 4) the scheduling of the optimal timing characteristics of clock distribution networks based on architectural and functional performance requirements. Each of these areas is described the clock distribution networks of specific industrial circuits are surveyed and future trends are discussed. Eby G. Friedman |
Proc. IEEE | 1 |
| 2001 | Exploiting the on-chip inductance in high-speed clock distribution networksabstractOn-chip inductance effects can be used to improve the performance of high-speed integrated circuits. Specifically, inductance improves the signal slew rate (the rise time), virtually eliminates short-circuit power consumption and reduces the area of the active devices and repeaters inserted to optimize the performance of long interconnects. These positive effects suggest the development of design strategies that benefit from on-chip inductance. An example of a clock distribution network is presented to illustrate the process in which inductance can be used to improve the performance of high-speed integrated circuits. Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2000 | Transparent repeatersabstractThe concept of a “transparent repeater1,” which is an amplifier circuit designed to minimize the delay introduced by highly resistive interconnect lines in high speed digital circuits, is introduced and described in this paper. An insertion methodology for this circuit is also discussed. Defining characteristics of this circuit are: the input is connected to the output, the output generates the same sense transition as the corresponding input transition, the buffer output becomes high impedance after every transition, and the buffer may detect input transitions with low threshold voltages. Radu M. Secareanu, Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 2 |
| 2000 | Noise estimation due to signal activity for capacitively coupled CMOS logic gatesabstractThe effect of interconnect coupling capacitance on neighboring CMOS logic gates driving coupled interconnections strongly depends upon the signal activity. A transient analysis of two capacitively coupled CMOS logic gates is presented in this paper for different combinations of signal activity. The uncertainty of the effective load capacitance and propagation delay due to the signal activity is addressed. Analytical expressions characterizing the output voltage and propagation delay are also presented for different signal activity conditions. The propagation delay based on these analytical expressions is within 3% as compared to SPICE, while the estimated delay neglecting the difference between the load capacitances can exceed 45%. The logic gates should be properly sized to balance the load capacitances in order to minimize any uncertainty in the signal delay. The peak noise voltage on a quiet interconnection determined from the analytical expressions is within 4% of SPICE. Kevin T. Tang, Eby G. Friedman |
ACM Great Lakes Symposium on VLSI | 2 |
| 2000 | Sensitivity of interconnect delay to on-chip inductanceabstractInductance extraction has become an important issue in the design of high speed CMOS circuits. Two characteristics of on-chip inductance are discussed in this paper that can significantly simplify the extraction of on-chip inductance, The first characteristic is that the sensitivity of a signal waveform to errors in the inductance values is low, particularly the propagation delay and the rise time. It is quantitatively shown in this paper that the error in the propagation delay and rise time is below 9.4% and 5.9%, respectively, assuming a 30% relative error in the extracted inductance. If an RC model is used for the same example, the corresponding errors are 51% and 71%, respectively, The second characteristic is that the magnitude of the on-chip inductance is a slowly varying function of the width of a wire and the geometry of the surrounding wires. These two characteristics can be exploited by using simplified techniques that permit approximate and sufficiently accurate values of the on-chip inductance to be determined with high computational efficiency. Yehea I. Ismail, Eby G. Friedman |
ISCAS | 2 |
| 2000 | Physical design to improve the noise immunity of digital circuits in a mixed-signal smart-power systemabstractTheoretical, simulation and experimental analysis and data are presented, discussing physical design techniques which influence the noise behavior of digital circuits in a mixed-signal smart-power system. Several physical design strategies are presented to improve the noise immunity of digital circuits in smart-power systems. Radu M. Secareanu, Scott Warner, Scott Seabridge, Cathie Burke, Thomas E. Watrobski, Christopher Morton, William Staub, Thomas Tellier, Eby G. Friedman |
ISCAS | 9 |
| 2000 | Delay and power expressions characterizing a CMOS inverter driving an RLC loadabstractOn-chip parasitic inductance has become an important design issue in high speed integrated circuits. On-chip inductance may degrade on-chip signal quality, affect transmission delay, and cause additional short-circuit power dissipation. The effects of on-chip inductance on the output voltage, propagation delay, and short-circuit power of a CMOS inverter are presented in this paper. Analytic equations characterizing the output voltage are derived based on an assumption of a fast ramp input signal. Closed form expressions describing the short-circuit power are also presented, The accuracy of these analytic equations is within 10% as compared to SPICE simulations. It is demonstrated that large inductive loads and fast input transition times can increase short-circuit current. Kevin T. Tang, Eby G. Friedman |
ISCAS | 2 |
| 2000 | Transient analysis of a CMOS inverter driving resistive interconnectabstractExpressions characterizing the output voltage and propagation delay of a CMOS inverter driving a resistive-capacitive interconnect are presented in this paper. The MOS transistors are characterized by the nth power law model. In order to emphasize the nonlinear behavior of a CMOS inverter, the interconnect is modeled as a lumped RC load. The propagation delay of a CMOS inverter is characterized for both a fast ramp and a slow ramp input signal. The waveform of the output voltage based on these analytic equations is quite close to SPICE assuming a fast ramp input signal. The accuracy of the propagation delay model for both fast ramp and slow ramp input signals is within 7% as compared to SPICE simulations. Kevin T. Tang, Eby G. Friedman |
ISCAS | 2 |
| 2000 | Delay and noise estimation of CMOS logic gates driving coupled resistive-capacitive interconnectionsabstractThe effect of interconnect coupling capacitance on the transient characteristics of a CMOS logic gate strongly depends upon the signal activity. A transient analysis of CMOS logic gates driving two and three coupled resistive–capacitive interconnect lines is presented in this paper for different signal combinations. Analytical expressions characterizing the output voltage and the propagation delay of a CMOS logic gate are presented for a variety of signal activity conditions. The uncertainty of the effective load capacitance on the propagation delay due to the signal activity is also addressed. It is demonstrated that the effective load capacitance of a CMOS logic gate depends upon the intrinsic load capacitance, the coupling capacitance, the signal activity, and the size of the CMOS logic gates within a capacitively coupled system. Some design strategies are also suggested to reduce the peak noise voltage and the propagation delay caused by the interconnect coupling capacitance. Kevin T. Tang, Eby G. Friedman |
Integr. | 2 |
| 2000 | Equivalent Elmore delay for RLC treesabstractClosed-form solutions for the 50% delay, rise time, overshoots, and settling time of signals in an RLC tree are presented. These solutions have the same accuracy characteristics of the Elmore delay for RC trees and preserves the simplicity and recursive characteristics of the Elmore delay. Specifically, the complexity of calculating the time domain responses at all the nodes of an RLC tree is linearly proportional to the number of branches in the tree and the solutions are always stable. The closed-form expressions introduced here consider all damping conditions of an RLC circuit including the underdamped response, which is not considered by the Elmore delay due to the nonmonotone nature of the response. The continuous analytical nature of the solutions makes these expressions suitable for design methodologies and optimization techniques. Also, the solutions have significantly improved accuracy as compared to the Elmore delay for an overdamped response. The solutions introduced here for RLC trees can be practically used for the same purposes that the Elmore delay is used for RC trees. Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2000 | Effects of inductance on the propagation delay and repeater insertion in VLSI circuitsabstractA closed-form expression for the propagation delay of a CMOS gate driving a distributed RLC line is introduced that is within 5% of dynamic circuit simulations for a wide range of RLC loads. It is shown that the error in the propagation delay if inductance is neglected and the interconnect is treated as a distributed RC line can be over 35% for current on-chip interconnect. It is also shown that the traditional quadratic dependence of the propagation delay on the length of the interconnect for RC lines approaches a linear dependence as inductance effects increase. On-chip inductance is therefore expected to have a profound effect on traditional high-performance integrated circuit (IC) design methodologies. The closed-form delay model is applied to the problem of repeater insertion in RLC interconnect. Closed-form solutions are presented for inserting repeaters into RLC lines that are highly accurate with respect to numerical solutions. RC models can create errors of up to 30% in the total propagation delay of a repeater system as compared to the optimal delay if inductance is considered. The error between the RC and RLC models increases as the gate parasitic impedances decrease with technology scaling. Thus, the importance of inductance in high-performance very large scale integration (VLSI) design methodologies will increase as technologies scale. Yehea I. Ismail, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1999 | Effects of Inductance on the Propagation Delay and Repeater Insertion in VLSI CircuitsabstractA closed form expression for the propagation delay of a CMOS gate driving a distributed RLC line is introduced that is within 5% of dynamic circuit simulations for a wide range of RLC loads.It is shown that the traditional quadratic dependence of the propagation delay on the length of an RC line approaches a linear dependence as inductance effects increase.The closed form delay model is applied to the problem of repeater insertion in RLC interconnect.Closed form solutions are presented for inserting repeaters into RLC lines that are highly accurate with respect to numerical solutions.An RC model as compared to an RLC model creates errors of up to 30% in the total propagation delay of a repeater system.Considering inductance in repeater insertion is also shown to significantly save repeater area and power consumption.The error between the RC and RLC models increases as the gate parasitic impedances decrease which is consistent with technology scaling trends.Thus, the importance of inductance in high performance VLSI design methodologies will increase as technologies scale. Yehea I. Ismail, Eby G. Friedman |
DAC | 2 |
| 1999 | Equivalent Elmore Delay for RLC TreesabstractArticle Free Access Share on Equivalent Elmore delay for RLC trees Authors: Yehea I. Ismail Department of Electrical and Computer Engineering, University of Rochester, Rochester, New York Department of Electrical and Computer Engineering, University of Rochester, Rochester, New YorkView Profile , Eby G. Friedman Department of Electrical and Computer Engineering, University of Rochester, Rochester, New York Department of Electrical and Computer Engineering, University of Rochester, Rochester, New YorkView Profile , Jose L. Neves IBM Microelectronics, 1580 Route 52, East Fishkill, New York IBM Microelectronics, 1580 Route 52, East Fishkill, New YorkView Profile Authors Info & Claims DAC '99: Proceedings of the 36th annual ACM/IEEE Design Automation ConferenceJune 1999 Pages 715–720https://doi.org/10.1145/309847.310041Published:01 June 1999Publication History 7citation1,519DownloadsMetricsTotal Citations7Total Downloads1,519Last 12 Months76Last 6 weeks24 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
DAC | 2 |
| 1999 | Maximizing Performance by Retiming and Clock Skew SchedulingabstractThe application of retiming and clock skew scheduling for improving the operating speed of synchronous circuits under setup and hold constraints is investigated in this paper. It is shown that when both long and short paths are considered, circuits optimized by the simultaneous application of retiming and clock scheduling can achieve shorter clock periods than optimized circuits generated by applying either of the two techniques separately. A mixed-integer linear programming formulation and an efficient heuristic are given for the problem of simultaneous retiming and clock skew scheduling under setup and hold constraints. Experiments with benchmark circuits demonstrate the efficiency of this heuristic and the effectiveness of the combined optimization. All of the test circuits show improvement. For more than half of them, the maximum operating speed increases by more than 21% over the optimized circuits obtained by applying retiming or clock skew scheduling separately. Marios C. Papaefthymiou, Eby G. Friedman |
DAC | 3 |
| 1999 | Minimizing Sensitivity to Delay Variations in High-Performance Synchronous CircuitsabstractThis paper investigates retiming and clock skew scheduling for improving the tolerance of synchronous circuits to delay variations. It is shown that when both long and short paths are considered, circuits optimized by the combined application of the two techniques are more tolerant to delay variations than when optimized by either of the two techniques separately. A novel mixed-integer linear programming formulation is given for simultaneous retiming and clock scheduling with a target clock period and tolerance under setup and hold constraints. Experiments with LGSynth93 and ISCAS89 benchmark circuits demonstrate the effectiveness of the combined optimization. For half of the test circuits, tolerance to delay variations increased by at least 23% over the separate application of retiming and clock scheduling. Moreover, for two thirds of the test circuits, maximum tolerance improved by at least 11%. Marios C. Papaefthymiou, Eby G. Friedman |
DATE | 3 |
| 1999 | Inductance Effects in RLC TreesabstractA closed form solution for characterizing voltage-based signals in an RLC tree is presented. This closed form solution is used to derive figures of merit to characterize the effects of inductance at a specific node in an RLC tree. The effective damping factor of the signal at a specific node in an RLC tree is shown to be a useful figure of merit, As the effective damping factor of a signal increases, an RC model is sufficiently accurate to characterize that waveform. The rise time of the input signal driving an RLC tree is another factor characterizing the importance of inductance. As the rise time of the input signal becomes much larger than the effective LC time constant at a specific node within an RLC tree, the signal at this node does not exhibit the effects of inductance. Evidence is provided showing that using a single line analysis to determine the importance of including inductance to characterize a tree structured interconnect line is invalid in many cases and can lead to erroneous conclusions. Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
Great Lakes Symposium on VLSI | 2 |
| 1999 | Noise Immunity of Digital Circuits in Mixed-Signal Smart Power SystemsabstractExperimental data describing circuit and physical design issues that influence the noise immunity of digital latches in mixed-signal smart power circuits are described and discussed. The principal result of this paper is the characterization of the conditions under which substrate noise generated by high power analog circuitry affects digital latches. The experimental data characterize a variety of different noise mitigation techniques for the particular process technology circuit structures, signal/clocking interdependencies, and related conditions. Radu M. Secareanu, Ivan S. Kourtev, Juan Becerra, Thomas E. Watrobski, Christopher Morton, William Staub, Thomas Tellier, Eby G. Friedman |
Great Lakes Symposium on VLSI | 8 |
| 1999 | Repeater insertion in tree structured inductive interconnectabstractThe effects of inductance on repeater insertion in RLC trees is the focus of the paper. An algorithm is introduced to insert and size repeaters within an RLC tree to optimize a variety of possible cost functions such as minimizing the maximum path delay, the skew between branches, or a combination of area, power, and delay. The algorithm has a complexity proportional to the square of the number of possible repeater positions, permitting a repeater solution to be chosen that is close to the global minimum. The repeater insertion algorithm is used to insert repeaters within several copper based interconnect trees to minimize the maximum path delay based on both an RC model and an RLC model. The two buffering solutions are compared using the AS/X dynamic circuit simulator. It is shown that as inductance effects increase, the area and power consumed by the inserted repeaters to minimize the path delays of an RLC tree decreases. By including inductance in the repeater insertion methodology, the interconnect is modeled more accurately as compared to an RC model, permitting average savings in area, power, and delay of 40.8%, 15.6%, and 6.7%, respectively, for a variety of copper based interconnect trees from a 0.25 /spl mu/m CMOS technology. The average savings in area, power, and delay increases to 62.2%, 57.2% and 9.4%, respectively, when using five times faster devices with the same interconnect trees. Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
ICCAD | 2 |
| 1999 | Clock skew scheduling for improved reliability via quadratic programmingabstractThis paper considers the problem of determining an optimal clock skew schedule for a synchronous VLSI circuit. A novel formulation of clock skew scheduling as a constrained quadratic programming (QP) problem is introduced. The concept of a permissible range, or a valid interval, for the clock skew of each local data path is key to this QP approach. From a reliability perspective, the ideal clock schedule corresponds to each clock skew within the circuit being at the center of the respective permissible range. However, this ideal clock schedule is nor practically implementable because of limitations imposed by the connectivity among the registers within the circuit. To evaluate the reliability, a quadratic cost function is introduced as the Euclidean distance between the ideal schedule and a given practically feasible clock schedule. This cost function is the minimization objective of the described algorithms for the solution of the previously mentioned quadratic program. Furthermore, the work described here substantially differs from previous research in that it permits complete control over specific clock signal delays or skews within the circuit. Specifically, the algorithms described here can be employed to obtain results with explicitly specified target values of important clock delays/skews with a circuit, such as for example, the clock delays/skews for I/O registers. An additional benefit is a potential reduction in clock period of up to 10%. An efficient mathematical algorithm is derived for the solution of the QP problem with O(r/sup 3/) run time complexity and O(r/sup 2/) storage complexity, where r is the number of registers in the circuit. The algorithm is implemented as a C++ program and demonstrated on the ISCAS'89 suite of benchmark circuits as well as on a number of industrial circuits. The work described here yields additional insights into the correlation between circuit structure and circuit timing by characterizing the degree to which specific signal paths limit the overall performance and reliability of a circuit. This information is directly applicable to logic and architectural synthesis. Ivan S. Kourtev, Eby G. Friedman |
ICCAD | 2 |
| 1999 | Interconnect coupling noise in CMOS VLSI circuitsabstractInterconnect between a CMOS driver and re- ceiver can be modeled as a 1ossy transmission line in high speed CMOS VLSI circuits as transition times become comparable to or less than the time of flight delay of the signal through the low resistivity interconnect. In this paper, closed form expressions for the coupling noise between adjacent interconnect are presented to estimate the coupling noise voltage on a quiet line. These expressions are based on an assumption that the interconnections are loosely coupled, where the effect of the coupling noise on the waveform of the active line is small and can be ne- glected. It is demonstrated that the output impedance of the CMOS driver should preferably be comparable to the interconnect impedance in order to reduce the propagation delay of the CMOS driver stage. Kevin T. Tang, Eby G. Friedman |
ISPD | 2 |
| 1999 | Figures of merit to characterize the importance of on-chip inductanceabstractA closed-form solution for the output signal of a CMOS inverter driving an RLC transmission line is presented. This solution is based on the alpha power law for deep submicrometer technologies. Two figures of merit are presented that are useful for determining if a section of interconnect should be modeled as either an RLC or an RC impedance. The damping factor of a lumped RLC circuit is shown to be a useful criterion. The second useful figure of merit considered in this paper is the ratio of the rise time of the input signal at the driver of an interconnect line to the time of flight of the signals across the line. AS/X circuit simulations of an RLC transmission line and a five section RC II circuit based on a 0.25-/spl mu/m IBM CMOS technology are used to quantify and determine the relative accuracy of an RC model. One primary result of this paper is evidence demonstrating that a range for the length of the interconnect exists for which inductance effects are prominent. Furthermore, it is shown that under certain conditions, inductance effects are negligible despite the length of the section of interconnect. Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1998 | Figures of Merit to Characterize the Importance of On-Chip InductanceabstractA closed form solution for the output signal of a CMOS inverter driving an RLC transmission line is presented. This solution is based on the alpha power law for deep submicrometer technologies. Two figures of merit are presented that are useful for determining if a section of interconnect should be modeled as either an RLC or an RC impedance. The damping factor of a lumped RLC circuit is shown to be a useful figure of merit. The second useful figure of merit considered in this paper is the ratio of the rise time of the input signal at the driver of an interconnect line to the time of flight of the signals across the line. AS/X circuit simulations of an RLC transmission line and a five section RC II circuit based on a 0.25 µm IBM CMOS technology are used to quantify and determine the relative accuracy of an RC model. One primary result of this study is evidence demonstrating that a range for the length of the interconnect exists for which inductance effects are prominent. Furthermore, it is shown that under certain conditions, inductance effects are negligible despite the length of the section of interconnect. Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
DAC | 2 |
| 1998 | Dynamic and Short-Circuit Power of CMOS Gates Driving Lossless Transmission LinesabstractThe dynamic and short-circuit power consumption of a CMOS gate driving an LC transmission line as a limiting case of an RLC transmission line is investigated in this paper. Closed form solutions for the output voltage and short-circuit power of a CMOS gate driving an LC transmission line are presented These solutions agree with AS/X circuit simulations within 11% error for a wide range of transistor widths and line impedances. The ratio of the short-circuit to dynamic power is shown to be less than 7% for CMOS gates driving LC transmission lines where the line is matched or underdriven. The total power consumption is expected to decrease as inductance effects become more significant as compared to an RC dominated interconnect. Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
Great Lakes Symposium on VLSI | 2 |
| 1998 | Power dissipated by CMOS gates driving lossless transmission linesabstractThe dynamic and short-circuit power consumption of a CMOS gate driving an LC transmission line as a limiting case of an RLC transmission line is investigated in this paper. Closed form solutions for the output voltage and short-circuit power of a CMOS gate driving an LC transmission line are presented. These solutions agree with AS/X simulations within 11% error for a wide range of transistor widths and line impedances. The ratio of the short-circuit to dynamic power is less than 7% for CMOS gates driving LC transmission lines where the line is matched or underdriven. Therefore, the total power consumption is expected to decrease as inductance effects becomes more significant is compared to an RC model of the interconnect. Yehea I. Ismail, Eby G. Friedman, José Neves 0002 |
ISLPED | 2 |
| 1997 | Incorporating interconnect, register, and clock distribution delays into the retiming processabstractA retiming algorithm is presented which includes the effects of variable register, clock distribution, and interconnect delays. These delay components are incorporated into the retiming process by assigning register electrical characteristics (RECs) to each edge in the graph representation of a synchronous circuit. A matrix, called the sequential adjacency matrix (SAM), is presented that contains all path delays. Timing constraints for each data path are derived from this matrix. Vertex lags are assigned ranges rather than single values as in existing retiming algorithms. The approach used in the proposed algorithm is to initialize these ranges with unbounded values and to continuously tighten these ranges using localized timing constraints until an optimal solution is obtained. A branch and bound method is offered for the general retiming problem where the REC values are arbitrary. Certain monotonicity constraints can be placed on the REC values to permit the use of standard linear programming methods, thereby requiring significantly less computational time. These conditions and the feasibility of their application to practical circuits are presented. The algorithm is demonstrated on modified benchmark circuits and both increased clock frequencies and the elimination of all race conditions are observed. Tolga Soyata, Eby G. Friedman, James H. Mulligan Jr. |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1996 | Optimal Clock Skew Scheduling Tolerant to Process VariationsabstractA methodology is presented in this paper for determining an optimal set of clock path delays for designing high performance VLSI/ULSI-based clock distribution networks.This methodology emphasizes the use of non-zero clock skew to reduce the system-wide minimum clock period.Although choosing (or scheduling) clock skew values has been previously recognized as an optimization technique for reducing the minimum clock period, difficulty in controlling the delays of the clock paths due to process parameter variations has limited its effectiveness.In this paper the minimum clock period is reduced using intentional clock skew by calculating a permissible clock skew range for each local data path while incorporating process dependent delay values of the clock signal paths.Graph-based algorithms are presented for determining the minimum clock period and for selecting a range of processtolerant clock skews for each local data path in the circuit, respectively.These algorithms have been demonstrated on the ISCAS-89 suite of circuits.Furthermore, examples of clock distribution networks with intentional clock skew are shown to tolerate worst case clock skew variations of up to 30% without causing circuit failure while increasing the system-wide maximum clock frequency by up to 20% over zero skew-based systems. José Neves 0002, Eby G. Friedman |
DAC | 2 |
| 1996 | Design methodology for synthesizing clock distribution networks exploiting nonzero localized clock skewabstractAn integrated top-down design methodology is presented in this brief for synthesizing high performance clock distribution networks based on application dependent localized clock skew. The methodology is divided into four phases: (1) determining an optimal clock skew schedule composed of a set of nonzero clock skew values and the related minimum clock path delays; (2) designing the topology of the clock distribution network with delays assigned to each branch based on the circuit hierarchy, the aforementioned clock skew schedule, and minimizing process and environmental delay variations; (3) designing circuit structures to emulate the delay values assigned to the individual branches of the clock tree; and (4) designing the physical layout of the clock distribution network. The clock distribution network synthesis methodology is based on CMOS technology. The clock lines are transformed from distributed resistive capacitive interconnect lines into purely capacitive interconnect lines by partitioning the RC interconnect lines with inverting repeaters. Variations in process parameters are considered during the circuit design of the clock distribution network to guarantee a race-free circuit. Nominal errors of less than 2.5% for the delay of the clock paths and 7% for the clock skew between any two registers belonging to the same global data path as compared with SPICE Level-3 are demonstrated. José Neves 0002, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1995 | Minimizing Power Dissipation in Non-Zero Skew-Based Clock Distribution Networks
José Neves 0002, Eby G. Friedman |
ISCAS | 2 |
| 1995 | Monotonicity Constraints on Path Delays for Efficient Retiming with Localized Clock Skew and Variable Register DelayabstractClock skew and delay characteristics associated with practical registers are significant factors affecting the retiming of synchronous circuits. Although work recently reported using branch and bound techniques offers a means for effective retiming taking these factors into account, the computational complexity involved is substantially greater than that associated with less general retiming algorithms that use standard linear programming methods. This paper presents sufficient conditions among values of localized clock skew and register characteristics which permit the retiming process to be achieved with a considerable reduction in computational complexity. The application of these conditions to some practical synchronous circuits is illustrated. Tolga Soyata, Eby G. Friedman, James H. Mulligan Jr. |
ISCAS | 2 |
| 1995 | A unified design methodology for CMOS tapered buffersabstractIn this paper, the various disparate approaches to CMOS tapered buffer design are unified into an integrated design methodology. Circuit speed, power dissipation, physical area, and system reliability are the four performance criteria of concern in tapered buffers, and each places a separate, often conflicting, constraint on the design of a tapered buffer. Enhanced short-channel tapered buffer design equations are presented for propagation delay and power dissipation, as well as a new split-capacitor model of hot-carrier reliability of tapered buffers and a two-component physical area model. Each performance criterion is individually investigated and analyzed with respect to the number of stages and tapering factor, and the interaction of the four criteria is examined to develop both a qualitative and a quantitative understanding of the various design tradeoffs. The creation of process dependent look-up tables for optimal buffer design is described, and a methodology to apply these look-up tables to application-specific tapered buffers for both unconstrained and constrained systems is developed.> Brian S. Cherkauer, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1994 | Retiming with non-zero clock skew, variable register, and interconnect delay
Tolga Soyata, Eby G. Friedman |
ICCAD | 2 |
| 1994 | Unification of Speed, Power, Area & Reliability in CMOS Tapered Buffer DesignabstractCircuit speed, power dissipation, physical area, and system reliability are the four performance criteria of concern in tapered buffer design. Each places a separate, often conflicting constraint on the design of a tapered buffer. Enhanced short-channel tapered buffer design equations are developed for propagation delay and power dissipation, as well as a new split-capacitor model of hot-carrier reliability and a two-component physical area model. Each performance criterion is independently investigated and analyzed, and the interaction of the four criteria is examined to develop both a qualitative and a quantitative understanding of the various design tradeoffs. These disparate approaches to tapered buffer design are unified into a convenient, integrated design methodology.> Brian S. Cherkauer, Eby G. Friedman |
ISCAS | 2 |
| 1994 | Forum: From 100 Milliwatts/MIPS to 10 Microwatts/MIPSabstractThe design and application of low power VLSI-based circuits and systems is the focus of this forum and paper. Attention is placed on issues related to low power systems, such as choice of implementing semiconductor technology, power supply voltage (e.g., 3.3 V, 2.5 V, 1.0 V), CAD for low power design and synthesis, micropower digital and analog circuit design techniques, system architectural tradeoffs for low power, selective clocking for power management, low power synchronization strategies, self-calibration circuit techniques for low power, and low power figures of merit.> Eby G. Friedman, Eric A. Vittoz, David J. Allstot, Erik P. Harris |
ISCAS | 1 |
| 1994 | Circuit Synthesis of Clock Distribution Networks Based on Non-Zero Clock SkewabstractA methodology is presented in this paper for synthesizing clock distribution networks by inserting circuit structures to emulate the delay values assigned to specified branches of the clock tree. These clock distribution networks are designed with localized non-zero clock skew so as to improve circuit performance and reliability. The design methodology is targeted for CMOS technology. The clock lines are transformed from distributed resistive-capacitive interconnect lines into purely capacitive interconnect lines by partitioning the RC interconnect lines with inverting repeaters. The inverters are specified by the geometric size of the transistors, the slope of the ramp shaped input/output waveform, and the output load capacitance. The branch delay model integrates both an inverter delay model and an interconnect delay model. Maximum errors of less than 3% for the delay of the clock paths and 6% for the clock skew between any two registers belonging to the same global data path are obtained as compared to SPICE Level-3.> José Neves 0002, Eby G. Friedman |
ISCAS | 2 |
| 1994 | Channel width tapering of serially connected MOSFET's with emphasis on power dissipationabstractTransistor channel width tapering in serial MOSFET chains is shown in this paper to simultaneously decrease propagation delay, power dissipation, and physical area of VLSI circuits. Tapering is the process of decreasing the size of each MOSFET transistor width along a serial chain such that the largest transistor is connected to the power supply and the smallest is connected to the output node. A detailed explanation of the effects of tapering on the output waveform is presented with specific emphasis on the power dissipation of tapered chains. It is demonstrated that in many cases tapering decreases delay and changes the shape of the output waveform such that the time during which a load inverter is conducting short-circuit current is reduced. This decrease in short-circuit current also occurs in many cases where tapering may not offer a speed advantage. In addition, dynamic CV/sup 2/f power dissipation of the serial chain is reduced. In those circuits where tapering does not decrease propagation delay, tapering permits a designer to tradeoff speed for a reduction in both short-circuit and dynamic power dissipation, a tradeoff not normally available with untapered chains. Thus the total power consumed by a serial chain of MOSFET's, as well as its propagation delay and area, can be reduced by channel width tapering. A design system for determining when tapering is appropriate, selecting the amount of tapering, and synthesizing the physical layout is presented. Physical layout issues unique to tapering are discussed, and fabricated test structures are described.> Brian S. Cherkauer, Eby G. Friedman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1993 | The Effects of Channel Width Tapering on the Power Dissipation of Serially Connected MOSFETs
Brian S. Cherkauer, Eby G. Friedman |
ISCAS | 2 |
| 1993 | Clock Distribution Design in VLSI Circuits. An Overview
Eby G. Friedman |
ISCAS | 1 |
| 1993 | Integration of Clock Skew and Register Delays into a Retiming Algorithm
Tolga Soyata, Eby G. Friedman, James H. Mulligan Jr. |
ISCAS | 2 |