Matthew R. Guthaus

dblp:70/1802 · also Matthew Guthaus · DBLP profile ↗
← Back
48ranked-venue papers
12as first author
6since 2021 · last 2026
0000-0002-8113-4531ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 48 · 12 first-author · 6 since 2021
YearPublicationVenuePosition
2026 A Fast and Accurate Surrogate Model for Clock-Mesh Timing Analysis
abstract
Clock meshes are an essential technique in high-performance VLSI systems to minimize skew and handle On-Chip Variation (OCV), especially in nanometer technologies. However, analyzing meshes is difficult due to reconvergent paths and multi-source drivers. The industrial standard is to use SPICE simulation since static timing analysis (STA) tools can not handle mesh loops. SPICE simulations are accurate, but slow, and approximate models miss critical effects like input slew and input skew. In this work, we propose a Graph Neural Network based surrogate model of the clock mesh represented as a graph with augmented structural and physical features. Trained on SPICE data, our model achieves high accuracy with average delay error of 0.64ps on unseen real designs versus 17.51ps from prior approximate models, while achieving speed-ups up to 3900x over multi-threaded SPICE simulation enabling faster and accurate analysis for clock meshes. Furthermore, we propose a curriculum learning strategy that further improves our model for OCV analysis.
Muhammad Hadir Khan, Matthew R. Guthaus
ACM Great Lakes Symposium on VLSI2
2026 An Open-Source Flow for Single-Phase, Edge-Triggered to Two-Phase, Non-Overlapping Clocking Conversion
abstract
Two-phase clocking offers significant advantages in timing margin and clock flexibility, yet its adoption remains limited due to the absence of automation in modern design flows. Managing strict non-overlap and 180° phase separation introduces complexity in RTL implementation and timing closure, leaving two-phase clocking rare in practice. This paper presents the first fully automated two-phase clocking flow integrated into OpenROAD Flow Scripts (ORFS). Our methodology automatically transforms flip-flop-based RTL into two-phase latch-based designs using Yosys technology mapping, ABC retiming, dual clock tree synthesis, two-phase correctness validation, and full physical design from RTL-to-GDS. We implement clock-gated and recirculation mux variants, where clock-gated achieves an average 29.2% power reduction and 50% latch count reduction over recirculation mux. Both variants are compared against flip-flop baselines, demonstrating timing closure through time borrowing on a design that failed timing with flip-flops.
Paolo Pedroso, Lee-Way Wang, Matthew R. Guthaus
ACM Great Lakes Symposium on VLSI3
2025 Invited: Mapping Two Decades of Innovation: Lessons from 25 Years of ISPD Research
abstract
The design automation research community has driven the evolution of integrated circuits from a handful of transistors in the 1960s to billions today. The International Symposium on Physical Design (ISPD) has been instrumental in tackling challenges like scaling complexities, hardware security, and the exponential growth in transistor counts. This study conducts a comprehensive bibliometric analysis of ISPD publications using Natural Language Processing, machine learning, and network analysis. It explores research themes, collaboration dynamics, and global contributions through citation networks, co-authorship graphs, geographical and spatial mapping, and topic modeling. Key areas of focus include Physical Design Optimization, Power Efficiency, and Emerging Technologies, with prominent topics such as placement, routing, clock skew, lithography, machine learning, and hardware security. The analysis highlights the evolution of foundational techniques like placement and routing while identifying emerging trends such as AI-driven design automation. These insights provide a roadmap for sustaining innovation in physical design over the next 25 years.
Gona Rahmaniani, Matthew R. Guthaus, Laleh Behjat
ISPD2
2024 GAT-Steiner: Rectilinear Steiner Minimal Tree Prediction Using GNNs
abstract
The Rectilinear Steiner Minimum Tree (RSMT) problem is a fundamental problem in VLSI placement and routing and is known to be NP-hard. Traditional RSMT algorithms spend a significant amount of time on finding Steiner points to reduce the total wire length or use heuristics to approximate producing sub-optimal results. We show that Graph Neural Networks (GNNs) can be used to predict optimal Steiner points in RSMTs with high accuracy and can be parallelized on GPUs. In this paper, we propose GAT-Steiner, a graph attention network model that correctly predicts 99.846% of the nets in the ISPD19 benchmark with an average increase in wire length of only 0.480% on suboptimal wire length nets. On randomly generated benchmarks, GAT-Steiner correctly predicts 99.942% with an average increase in wire length of only 0.420% on suboptimal wire length nets.
Bugra Onal, Eren Dogan, Muhammad Hadir Khan, Matthew R. Guthaus
ICCAD4
2023 SRAM Design with OpenRAM in SkyWater 130nm
abstract
OpenRAM is an open-source framework for the development of memories with an initial focus on SRAMs. OpenRAM provides an application interface for netlist, layout, and characterization to create designs using either open-source or commercial verification and simulation tools. The first silicon in SkyWater$\mathbf{130}nm$has been successfully verified which includes a 32-bit 1-kilobyte dual-port SRAM macro. This paper presents the first macro design, test setup, test results, and enhancements on a subsequent tape-out.
Jesse Cirimelli-Low, Muhammad Hadir Khan, Samuel Crow, Amogh Lonkar, Bugra Onal, Andrew D. Zonenberg, Matthew R. Guthaus
ISCAS7
2023 OpenSpike: An OpenRAM SNN Accelerator
abstract
This paper presents a spiking neural network (SNN) accelerator made using fully open-source EDA tools, process design kit (PDK), and memory macros synthesized using Open-RAM. The chip is taped out in the 130 nm SkyWater process and integrates over 1 million synaptic weights, and offers a reprogrammable architecture. It operates at a clock speed of 40 MHz, a supply of 1.8 V, uses a PicoRV32 core for control, and occupies an area of 33.3 mm2. The throughput of the accelerator is 48,262 images per second with a wallclock time of 20.72$\mu \mathbf{s}$, at 56.8 GOPS/W. The spiking neurons use hysteresis to provide an adaptive threshold (i.e., a Schmitt trigger) which can reduce state instability. This results in high performing SNNs across a range of benchmarks that remain competitive with state-of-the-art, full precision SNNs. The design is open sourced and available online: https://githuh.com/sJmth/OpenSpike
Farhad Modaresi, Matthew R. Guthaus, Jason Kamran Eshraghian
ISCAS2
2019 Fast and Area-Efficient SRAM Word-Line Optimization
abstract
A word line driver controls the access of cells in a row in Static Random Access Memories (SRAMs) and has a significant impact on SRAM speed and power consumption. When gate delay is the dominant factor, simple models are a good guideline for fast word lines. However, routing wire delay is significant when the row size is large, which causes these designs to be suboptimal. This paper presents an analytical optimization technique using a delay model that includes gate delay, wire resistance, and wire capacitance to optimize high-performance word line driver topologies for SRAMs. The proposed methodology has a maximum 45% delay improvement and 42% buffer cost reduction.
James E. Stine, Matthew R. Guthaus
ISCAS3
2019 Automated Synthesis of Multi-Port Memories and Control
abstract
High performance systems often employ multi-ported memories to enhance the throughput and flexibility of the memory. Existing SRAM compilers offer limited control over the SRAM design and port configurations while SRAMs are commonly dual-ported. Experimental designs could benefit from design exploration of multi-port configurations. We propose an open-source, multi-port solution that extends the OpenRAM memory compiler. A parameterized bitcell is presented which can support any combination of read, write, and read-write ports. The bitcell layout is generated for these port combinations, and the SRAM layout can support any combination of two ports. In addition, support for multi-port characterization and functional testing ensures correctness and incorporation into design methodologies.
Hunter Nichols, Michael Grimes, Jennifer Sowash, Jesse Cirimelli-Low, Matthew R. Guthaus
VLSI-SoC5
2019 Bottom-Up Approach for High Speed SRAM Word-line Buffer Insertion Optimization
abstract
The delay of a square SRAM array is dominated by the bit line delay due to the high capacitance per unit length attached to the bit line. Hence, SRAM arrays are usually longer in the word line direction. However, the word line delay also increases dramatically in a simple naive topology and can be a dominating factor when the word line dimension is much longer than that of the bit line. Therefore, word line optimization is an important part of SRAM delay optimization. Buffer insertion, which is commonly used for long interconnects, can also be used to improve word line delay. This paper proposes an approach to place and size the buffers to reduce word line and overall SRAM delay. The proposed methodology improves the read critical path delay by 15.7%, at the cost of only 5.26% extra area in a 128 Kbit SRAM.
Matthew R. Guthaus
VLSI-SoC2
2018 DCMCS: Highly Robust Low-Power Differential Current-Mode Clocking and Synthesis
Riadul Islam, Hany Ahmed Fahmy, Ping-Yao Lin, Matthew R. Guthaus
IEEE Trans. Very Large Scale Integr. Syst.4
2017 Energy Savings and Performance Improvement in Subthreshold Using Adaptive Body Bias
abstract
In subthreshold operation, circuits are more sensitive to the impact of parametric variation due to reduced supply voltages. To meet timing specification and ensure reliable operation, circuits require compensation techniques that mitigate variation. We developed a design methodology to use adaptive forward body bias and reduce worst case 3-sigma active energy, delay and standby energy caused by threshold voltage variation. We validated this methodology on the ISCAS85 benchmarks and improved the worst-case metrics in each case, with no loss of performance. Our approach reduces worst-case standby energy and worst-case active energy by up to 21.06% and 18.80%, respectively, on average.
Rajsaktish Sankaranarayanan, Matthew R. Guthaus
ACM Great Lakes Symposium on VLSI2
2017 Timing speculative SRAM
abstract
Static Random Access Memories (SRAMs) are considered a major bottleneck in high performance System-on-Chip (SoC) design and there is a large demand for high performance SRAMs with minimal energy consumption. Time speculation techniques such as Razor ease timing guardbands to improve performance or reduce energy consumption. The state-of-the-art approach has high area and energy overheads due to the error detection logic. This study proposes a timing speculative SRAM that extends the existing Replica Bitline Column to detect read timing failures. We also extend the SRAM decode logic to protect from incorrect write operations. We demonstrate our Replica-based Timing Speculative SRAM (RTS) is an energy and area efficient design alternative to prior techniques such as Razor. Our proposed design is 22% to 58% more energy efficient in reading operations and it has an error detection mechanism which is 35% to 73% more area efficient that Razor-enabled SRAM.
Elnaz Ebrahimi 0001, Matthew R. Guthaus, Jose Renau
ISCAS2
2017 Architectural opportunities for novel dynamic EMI shifting (DEMIS)
abstract
Processors emit non-trivial amounts of electromagnetic radiation, creating interference in frequency bands used by wireless communication technologies such as cellular, WiFi and Bluetooth. We introduce the problem of in-band radio frequency noise as a form of electromagnetic interference (EMI) to the computer architecture community as a technical challenge to be addressed.
Daphne I. Gorman, Matthew R. Guthaus, Jose Renau
MICRO2
2017 CMCS: Current-Mode Clock Synthesis
abstract
In a high-performance VLSI design, the clock network consumes a significant amount of power. While most existing methodologies use voltage-mode (VM) signaling, these clock distributions lose a tremendous amount of dynamic power to charge/discharge the large global clock capacitance. New circuit approaches for current-mode (CM) clocking save significant clock power, but have been limited to only symmetric networks, while most application specific integrated circuits have asymmetric clock distributions. In this paper, we propose the first CM clock synthesis (CMCS) methodology to reduce the overall clock network power with low skew. The method can integrate with traditional clock routing followed by transmitter and receiver sizing. We validate the proposed methodology using ISPD 2009 and 2010 industrial benchmarks using an extracted SPICE model distributed in 1.4-275.6-mm2area and consists of 81-2249 sinks. This methodology saves 39%-84% average power with similar skew on the benchmarks using 45-nm CMOS technology simulation of clock frequencies range from 1-3 GHz. In addition, the CMCS methodology takes 2.4-9.1× less running time and consumes 20%-26% less transistor area compared with synthesized, buffered VM clock distributions.
Riadul Islam, Matthew R. Guthaus
IEEE Trans. Very Large Scale Integr. Syst.2
2016 OpenRAM: an open-source memory compiler
abstract
Computer systems research is often inhibited by the availability of memory designs. Existing Process Design Kits (PDKs) frequently lack memory compilers, while expensive commercial solutions only provide memory models with immutable cells, limited configurations, and restrictive licenses. Manually creating memories can be time consuming and tedious and the designs are usually inflexible. This paper introduces OpenRAM, an open-source memory compiler, that provides a platform for the generation, characterization, and verification of fabricable memory designs across various technologies, sizes, and configurations. It enables research in computer architecture, system-on-chip design, memory circuit and device research, and computer-aided design.
Matthew R. Guthaus, James E. Stine, Samira Ataei, Mehedi Sarwar
ICCAD1
2016 A 64 kb differential single-port 12T SRAM design with a bit-interleaving scheme for low-voltage operation in 32 nm SOI CMOS
abstract
In this paper, a novel differential single-port 12T SRAM bitcell is presented. This bitcell uses a read buffer to eliminate read disturbance, improves the read stability and achieves read static noise margin equal to its hold static noise margin. Using a column-based select signal this bitcell provides a half-select free feature, facilitating a bit-interleaving structure to reduce multi-bit soft errors by conventional error correcting code techniques. By boosting the wordline and select signal voltage, this bitcell can read and write with no error at 300 mV while data can be held down to 250 mV in standby mode. Bitline leakage suppression in 12T bitcell allows more bitcells per bitline for high density SRAMs and provides faster read operation. This paper also introduces OpenRAM, an open-source memory compiler, that provides a platform for the generation, characterization, and verification of fabricable memory designs across various technologies, sizes, and configurations. Using OpenRAM, a 64 kb 12T SRAM macro is designed in IBM 32 nm SOI CMOS technology that operates down to 0.3 V with 50 MHz operating frequency while it functions at 0.9 V with 2.2 GHz operating frequency, as well.
Samira Ataei, James E. Stine, Matthew R. Guthaus
ICCD3
2015 Switched capacitor quasi-adiabatic clocks
abstract
Clock Distribution Networks (CDNs) in high speed designs can consume 30-50% of the total chip dynamic power. Adiabatic clock circuits can save some of this power, but these depend on a time varying power supply which is difficult to implement in practice. In this paper, we present the first quasi-adiabatic clock circuit with a constant supply voltage at high speeds. Our proposed adiabatic clocks attain an average 23% clock power savings with better slew rate and the same skew compared to traditional buffered clocks.
Hany Ahmed Fahmy, Ping-Yao Lin, Riadul Islam, Matthew R. Guthaus
ISCAS4
2015 Multi-frequency resonant clocks
abstract
Clock distribution networks consume a significant portion of total chip power in high-performance designs. Resonant clocks are one proposed method to lower this power in modern designs as well as a fewer required clock buffers. Recent resonant solutions are limited to optimal performance at one particular frequency which is problematic since dynamic frequency scaling is often used to lower overall system power. This paper introduces the first scheme to produce a clock distribution network with a tunable resonant frequency. Experimental results show the resonant frequency ranges from 1.2GHz to 2.6GHz while saving up to 41% of the power on the clock distribution network when compared to the non-resonant distribution.
Benjamin M. LaCara, Ping-Yao Lin, Matthew R. Guthaus
ISCAS3
2015 LC resonant clock resource minimization using compensation capacitance
abstract
Distributed-LC resonant clock distribution is a viable technique to reduce clock distribution network (CDN) dynamic power. However, resonant clocks can require significant on-chip resources to form the inductors and decoupling capacitors which discourages adoption. This paper uses a compensation capacitor (Cc) to reduce the overhead of the on-chip inductor and capacitor resources without changing the performance of a distributed-LC resonant clock. Analysis on the ISPD clock benchmarks show nearly 12% reduction in passive device area compared to previous resonant clocks while still saving 49.9% power over traditional buffered clocks.
Ping-Yao Lin, Hany Ahmed Fahmy, Riadul Islam, Matthew R. Guthaus
ISCAS4
2014 Current-mode clock distribution
abstract
We propose a new paradigm for clock distribution that uses current, rather than voltage, to distribute a global clock signal with reduced power consumption. While current-mode (CM) signaling has been used in one-to-one signals, this is the first usage in a one-to-many clock distribution network. To accomplish this, we create a new high-performance current-mode pulsed flip-flop (CMPFF) using a representative 45 nm CMOS technology. When the CMPFF is combined with a CM transmitter, the first CM clock distribution network exhibits 45.2% lower average power compared to traditional voltage mode clocks.
Riadul Islam, Matthew R. Guthaus
ISCAS2
2013 Embedded tutorials: Embedded tutorial 1: Cell-aware test-from gates to transistors
abstract
Devices manufactured in 20 nm and smaller geometry technologies will potentially be very large by today's standards, they will also have new characteristics implied by things like process variability and adoption of FinFET transistors. The industry has cumulatively adopted more and more sophisticated fault models that use timing as well as layout information. There is a growing body of experimental data showing it is still insufficient. The next area of focus will be the quality of test. Cell-aware test is one of the most promising approaches developed over the last five years aimed at improving the quality of test while maintaining the efficiency of gate-level approach. This approach combines two levels of abstraction to provide trade-offs between accuracy and efficiency. The first step creates the cell-aware test library models. It starts with standard cell libraries and performs layout extraction. Realistic defects (bridges and opens) are injected into the SPICE netlist, and analog fault simulation is performed to determine the conditions under which the defects are detected. Those conditions are aggregated to create a compact and efficient representation of the libraries for ATPG done at the gate-level. Generation of library views for cell-aware test is performed only once for a given standard cell library. The final cell-aware ATPG generates the high quality test patterns based on the cell-aware library views. This guarantees that the investment in gate-level ATPG infrastructure could be efficiently utilized. The technology has been used on a number of high-volume industrial designs. The experimental data show a significant increase of defect coverage and the corresponding improvement of defect rate.
Janusz Rajski, Miodrag Potkonjak, Adit D. Singh, Abhijit Chatterjee, Zainalabedin Navabi, Matthew R. Guthaus, Sezer Gören 0001
VLSI-SoC6
2013 Revisiting automated physical synthesis of high-performance clock networks
abstract
High-performance clock distribution has been a challenge for nearly three decades. During this time, clock synthesis tools and algorithms have strove to address a myriad of important issues helping designers to create faster, more reliable, and more power efficient chips. This work provides a complete discussion of the high-performance ASIC clock distribution using information gathered from both leading industrial clock designers and previous research publications. While many techniques are only briefly explained, the references summarize the most influential papers on a variety of topics for more in-depth investigation. This article also provides a thorough discussion of current issues in clock synthesis and concludes with insight into future research and design challenges for the community at large.
Matthew R. Guthaus, Gustavo Wilke, Ricardo Augusto da Luz Reis
ACM Trans. Design Autom. Electr. Syst.1
2012 Library-aware resonant clock synthesis (LARCS)
abstract
Clock grids are often used in high-performance ASIC designs because of their low skew and robustness to variations. Resonant clock grids have the potential to reduce the power consumption of these high-performance clocks without sacrificing the skew and robustness of a clock grid. We present the first methodology to synthesize high-performance distributed resonant LC tank clock grids that utilize a pre-characterized inductor library. The use of a library reduces designer effort and total inductor area when compared with previous resonant clock grids while still attaining 59% power reduction and competitive skew when compared to traditional buffered clock grids.
Xuchu Hu, Walter James Condley, Matthew R. Guthaus
DAC3
2012 Lithography-aware layout compaction
abstract
Optical Proximity Correction (OPC) tools can suffer if the origi- nal layout is inherently difficult to print. Most routing techniques are unaware of the lithographic process, but several have been proposed to make the layout easier for the OPC tool to correct. This paper proposes a generalized preprocess step for OPC that uses a modified 1D layout compactor to expand or shrink geometry to decrease the amount of OPC work needed. This lithography-aware compactor estimates the OPC effort required and then uses a non-linear formulation to adjust the layout geometry. We describe a method for performing this compaction on small layouts and then extend it to handle larger hierarchical standard cell layouts.
Curtis Andrus, Matthew R. Guthaus
ACM Great Lakes Symposium on VLSI2
2012 High-Performance, Low-Power Resonant Clocking: Embedded tutorial
abstract
Clock distribution networks consume a significant portion of on-chip power. Traditional buffered clock distribution power is limited by frequency, capacitance, and activity rates. Resonant clock distributions can reduce this power by "recycling" energy on-chip and reducing the overall clock power. This tutorial introduces recent techniques for distributed-LC, traveling wave, and standing wave resonant clock distributions. In particular, the tutorial discusses the recent developments and open research problems. The tutorial covers both circuits, computer-aided design algorithms and methodologies for resonant clocking.
Matthew R. Guthaus, Baris Taskin
ICCAD1
2012 Welcome from the general chair
Matthew R. Guthaus
VLSI-SoC1
2012 Dynamic voltage scaling for SEU-tolerance in low-power memories
Seokjoong Kim, Matthew R. Guthaus
VLSI-SoC2
2012 A single-VDD ultra-low energy sub-threshold FPGA
Rajsaktish Sankaranarayanan, Matthew R. Guthaus
VLSI-SoC2
2012 Harmonic resonant clocking
abstract
Distributed iuductor-ca p acitor (LC) resouaut c1ock iug is a receut, promisiug techuique to reduce the euergy cousum p tiou iu Clock Distributiou Networks (CDNs) by recycliug the euergy ou-chi p .Eveu though the majority of power is saved, resouaut clocks distribute a siuusoidal clock sigual with a 25 % slew which iucreases short-circuit power iu the sequeutial elemeuts com p ared to traditioual buffered clocks.Iu this work, we preseut the first harmouic resouant clock circuit that adds a third harmonic to the fundamental frequency in order to iucrease the slew rate of resonant clocks and reduce the short circuit p ower in the sequential elements.We p resent two different methods of tuning a secondary tank circuit in a Colpitts oscillator to minimize the slew: one based on the frequency res p onse of the circuit and the other based on matching an ideal square wave clock signal.Both methods provide benefits in slew reduction at a modest cost of components. I. INTRODUCTIONClock distribution networks (CDN) can consume between 30-70% of total chip power in high-performance processors.Many techniques have been proposed to minimize the dynamic power including minimal wirelength routing, buffer/wire siz ing, low-voltage-swing, dynamic voltage and frequency scal ing, and clock gating.Many of these techniques exploit inactivity or reduced workloads to reduce power, but the fundamental limit during active modes is a severe challenge in future highly-threaded many-core and GPU processors.Resonant clock distributions are another alternative to save power that can recycle the energy on-chip.Various resonant
Haven Blake Skinner, Xuchu Hu, Matthew R. Guthaus
VLSI-SoC3
2012 High-performance clock mesh optimization
abstract
Clock meshes are extremely effective at producing low-skew regional clock networks that are tolerant of environmental and process variations. For this reason, clock meshes are used in most high-performance designs, but this robustness consumes significant power. In this work, we present two techniques to optimize high-performance clock meshes. The first technique is a mesh perturbation methodology for nonuniform mesh routing. The second technique is a skew-aware buffer placement through iterative buffer deletion. We demonstrate how these optimizations can achieve significant power reductions and a near elimination of short-circuit power. In addition, the total wire length is decreased, the number of required buffers is decreased, and both skew and robustness are improved on average when variation is considered.
Matthew R. Guthaus, Xuchu Hu, Gustavo Wilke, Guilherme Flach, Ricardo Augusto da Luz Reis
ACM Trans. Design Autom. Electr. Syst.1
2011 Clock tree optimization for Electromagnetic Compatibility (EMC)
abstract
Electromagnetic Interference (EMI) generated by electronic systems is increasing with operating frequency and shrinking process technologies. The clock distribution network is one of the major causes of on-chip EMI. In this paper, we discuss the EMI problem in clock tree design. Spectrum analysis shows that slew rate of clock signal is the main parameter determining the high-frequency spectral content distribution. This is the first work to consider maximum and minimum buffer slew rates in clock tree synthesis to reduce EMI. In this paper, we propose a dynamic programming algorithm to optimize the clock tree considering both traditional metrics and Electromagnetic Compatibility (EMC). Our experimental results show that slew can be controlled in a feasible range and high-frequency spectrum contents can be reduced without sacrificing the traditional metrics such as power and skew. With the efficient optimization and pruning method, the biggest benchmark is able to complete in four minutes.
Xuchu Hu, Matthew R. Guthaus
ASP-DAC2
2011 Distributed Resonant clOCK grid Synthesis (ROCKS)
abstract
Clock distribution networks can consume 35-70% of total chip power in high-performance designs [13]. Resonant clocks can potentially reduce this power by recycling the energy using on-chip inductors. We propose the first automated Resonant clOCK Synthesis (ROCKS) algorithm. Experimental results show that with 10% inductor area, clock power can be reduced by 34%. With more inductor area, up to 90% power savings is shown feasible.
Xuchu Hu, Matthew R. Guthaus
DAC2
2011 Leakage-aware redundancy for reliable sub-threshold memories
abstract
In this work, we are the first to consider the optimization of sub-threshold stand-by VDD while simultaneously considering memory yield and redundant row/column usage. We propose a fast, optimal fault-repair analysis framework that is 200--600% faster than previous works and show that leakage can be reduced 10--14% using redundancy without sacrificing yield.
Seokjoong Kim, Matthew R. Guthaus
DAC2
2011 A methodology for local resonant clock synthesis using LC-assisted local clock buffers
abstract
Resonant clocking is a form of adiabatic clocking that retains much of the energy present in clock switching and recycles it into the following clock cycle. In this paper we present the first automated methodology using LC-assisted local clock buffers (LCLCB) for generating local resonant clocks. This uses a single-buffer single-inductor sector topology applied to non-uniform trees as found in most ASIC designs. We show that this form of adiabatic clocking can achieve power savings as much as 75% over traditional buffered clock networks.
Walter James Condley, Xuchu Hu, Matthew R. Guthaus
ICCAD3
2011 Low-power multiple-bit upset tolerant memory optimization
abstract
In this paper, we propose a framework for analyzing Soft Error Rates (SER) including Multiple-Bit Upsets (MBU). Then, using this framework, we optimize the soft error tolerant voltage (Vtol) and interleaving distance (ID) of low-power, error-tolerant memories. Experimental results show that the total power can be reduced by an average of 30.5% with Vtoloptimization and an average of 40.9% by simultaneously considering Vtoland ID together when compared to worst-case design practices.
Seokjoong Kim, Matthew R. Guthaus
ICCAD2
2011 Distributed LC resonant clock tree synthesis
abstract
Clock networks in high-performance designs are extremely power hungry. One potential method for reducing the power consumption is to use distributed LC tanks in which energy is conserved by shifting it between electrical and magnetic forms at the resonant frequency. However, no physical algorithms to physically synthesize resonant trees have been proposed. In order to utilize such techniques in ASICs, this work presents the first algorithm to synthesize resonant regional clock trees. Our results suggest that, on average, we can reduce clock power consumption by 41.7% at 2Ghz with no degredation to skew compared to a minimum buffer insertion algorithm.
Matthew R. Guthaus
ISCAS1
2011 SNM-aware power reduction and reliability improvement in 45nm SRAMs
abstract
Read stability and write ability are important factors in SRAM design. However, read stability and write ability of the 6T-cell are conflicting design requirements. It is increasingly more difficult to balance these requirements with conventional transistor sizing and Vthoptimization as technology scales. In this paper, we present a method to help reduce SRAM power consumption by deciding the optimal pull-up transistor size in 6T SRAM cell. Our experiments show that the proposed method can reduce the supply voltage as much as 7-8%. Furthermore, our method has an average 7.76% power reduction in both active and standby mode at 25°C, a 35% Vthvariation reduction, 7.78% SNM variation reduction and 11.3% NBTI degradation reduction with at most 1.95% area overhead.
Seokjoong Kim, Matthew R. Guthaus
VLSI-SoC2
2010 Non-uniform clock mesh optimization with linear programming buffer insertion
abstract
Clock meshes are extremely effective at filtering clock skew from environmental and process variations. For this reason, clock meshes are used in most high performance designs. However, this robustness costs power. In this work, we present a mesh edge displacement algorithm that is able to reduce mesh wire length by 7.6% and overall power by 10.5% with a small mean skew improvement. We also present the first non-greedy buffer placement and sizing technique using linear programming (LP) and iterative buffer removal. We show that compared to prior methods, we can obtain 41% power reduction and an 27ps mean skew reduction on average when variation is considered compared to prior algorithms.
Matthew R. Guthaus, Gustavo Wilke, Ricardo Augusto da Luz Reis
DAC1
2009 Measuring and modeling variabilityusing low-cost FPGAs
abstract
The focus of this paper is to measure and qualify high-level process variation models by measuring variability on FPGAs. Measurements are done with high spatial resolution and demonstrate how the high-resolution data matches two industry test cases. The benefit of such an approach is that several inexpensive FPGAs, which are normally on the leading edge of technologies compared to ASICs, obviate the need of fabricating many custom test chips. Specifically, our evaluation shows how measurements of an Altera Cyclone II FPGA can be used to derive variability models for several 90nm commercial designs such as the Sun Niagara and Intel Pentium D. Even though the FPGAs and commercial processors are produced by different fabs (TSMC, TI, and Intel, respectively), we find the FPGAs to be very useful for predicting variation in the commercial processors.
Cyrus Bazeghi, Matthew R. Guthaus, Jose Renau
FPGA3
2009 Fault-tolerant synthesis using non-uniform redundancy
abstract
As process technologies continue to scale into the nanometer regime, devices are becoming significantly more unreliable. Many forms of unreliability manifest as transient faults and can cause intermittent random logic upsets. These logic upsets are often caused by natural radiation (neutrons and alpha particles) or on-chip noise (cross-coupling, supply drop, or flicker noise). This research improves reliability by using non-uniform redundancy. Specifically, we present a dynamic programming algorithm that considers many possible topological redundancies, yet maintains a linear run-time due to efficient pruning of suboptimal solutions. Our algorithm provides designers with a Pareto-optimal set of solutions that trade reliability for area. Compared to existing triple modular redundancy (TMR), we see similar reliability with only 35% area overhead instead of 326%.
Keven L. Woo, Matthew R. Guthaus
ICCD2
2008 Clock tree synthesis with data-path sensitivity matching
abstract
This paper investigates methods for minimizing the impact of process variation on clock skew using buffer and wire sizing. While most papers on clock trees ignore data-path circuit variations and most papers on data-path circuit optimization disregard clock tree variation, we consider both. Using both clock and data-path variations together, we present a novel sensitivity-matching algorithm that allows clock tree skews to be intentionally correlated with data-path sensitivities to ameliorate timing violations due to variation. Our statistical tuning shows an improvement in terms of expected clock skew and clock skew variation over previously published robust algorithms.
Matthew R. Guthaus, Dennis Sylvester, Richard B. Brown
ASP-DAC1
2006 Process-induced skew reduction in nominal zero-skew clock trees
abstract
This work develops an analytic framework for clock tree analysis considering process variations that is shown to correspond well with Monte Carlo results. The analysis framework is used in a new algorithm that constructs deterministic nominal zero-skew clock trees that have reduced sensitivity to process variation. The new algorithm uses a sampling approach to perform route embedding during a bottom-up merging phase, but does not select the best embedding until the top-down phase. This results in clock trees that exhibit a mean skew reduction of 32.4% on average and a standard deviation reduction of 40.7% as verified by Monte Carlo. The average increase in total clock tree capacitance is less than 0.02%
Matthew R. Guthaus, Dennis Sylvester, Richard B. Brown
ASP-DAC1
2006 Clock buffer and wire sizing using sequential programming
abstract
This paper investigates methods for clock skew minimization using buffer and wire sizing. First, a technique that significantly improves solution quality and stability of sequential programming-based buffer/wire sizing is used. Then, a new formulation of clock skew minimization that uses quadratic programming and considers sub-critical skews in addition to the most critical skews is presented. The quality of results are verified to be more robust using Monte Carlo simulations to account for process sensitivity. For the same power budget, the sequential quadratic programming (SQP) method has better expected skew, standard deviation, and overall CPU time on average.
Matthew R. Guthaus, Dennis Sylvester, Richard B. Brown
DAC1
2005 Optimization objectives and models of variation for statistical gate sizing
abstract
This paper approaches statistical optimization by examining gate delay variation models and optimization objectives. Most previous work on statistical optimization has focused exclusively on the optimization algorithms without considering the effects of the variation models and objective functions. This work empirically derives a simple variation model that is then used to optimize for robustness. Optimal results from example circuits used to study the effect of the statistical objective function on parametric yield.
Matthew R. Guthaus, Natesan Venkateswaran, Vladimir Zolotov, Dennis Sylvester, Richard B. Brown
ACM Great Lakes Symposium on VLSI1
2005 Gate sizing using incremental parameterized statistical timing analysis
abstract
As technology scales into the sub-90 nm domain, manufacturing variations become an increasingly significant portion of circuit delay. As a result, delays must be modeled as statistical distributions during both analysis and optimization. This paper uses incremental, parametric statistical static timing analysis (SSTA) to perform gate sizing with a required yield target. Both correlated and uncorrelated process parameters are considered by using a first-order linear delay model with fitted process sensitivities. The fitted sensitivities are verified to be accurate with circuit simulations. Statistical information in the form of criticality probabilities are used to actively guide the optimization process which reduces run-time and improves area and performance. The gate sizing results show a significant improvement in worst slack at 99.86% yield over deterministic optimization.
Matthew R. Guthaus, Natesan Venkateswaran, Chandu Visweswariah, Vladimir Zolotov
ICCAD1
2005 Partitioning Variables across Register Windows to Reduce Spill Code in a Low-Power Processor
abstract
Low-power embedded processors utilize compact instruction encodings to achieve small code size. Such encodings place tight restrictions on the number of bits available to encode operand specifiers and, thus, on the number of architected registers. As a result, performance and power are often sacrificed as the burden of operand supply is shifted from the register file to the memory due to the limited number of registers. In this paper, we investigate the use of a windowed register file to address this problem by providing more registers than allowed in the encoding. The registers are organized as a set of identical register windows where, at each point in the execution, there is a single active window. Special window management instructions are used to change the active window and to transfer values between windows. This design gives the appearance of a large register file without compromising the instruction encoding. To support the windowed register file, we designed and implemented a graph partitioning-based compiler algorithm that partitions program variables and temporaries referenced within a procedure across multiple windows. On a 16-bit embedded processor, an average of 11 percent improvement in application performance and 25 percent reduction in system power was achieved as an 8-register design was scaled from one to two windows.
Rajiv A. Ravindran, Robert M. Senger, Eric D. Marsman, Ganesh S. Dasika, Matthew R. Guthaus, Scott A. Mahlke, Richard B. Brown
IEEE Trans. Computers5
2003 Increasing the number of effective registers in a low-power processor using a windowed register file
abstract
Low-power embedded processors utilize compact instruction encodings to achieve small code size. Instruction sizes of 8 to 16 bits are common. Such encodings place tight restrictions on the number of bits available to encode operand specifiers, and thus on the number of architected registers. The central problem with this approach is that performance and power are often sacrificed as the burden of operand supply is shifted from the register file to the memory due to the limited number of registers. In this paper, we investigate the use of a windowed register file to address this problem by providing more registers than allowed in the encoding. The registers are organized as a set of identical register windows where at each point in the execution there is a single active window. Special window management instructions are used to change the active window and to transfer values between windows. The goal of this design is to give the appearance of a large register file without compromising the instruction encoding. To support the windowed register file, we designed and implemented a novel graph partitioning based compiler algorithm that partitions virtual registers within a given procedure across multiple windows. On a 16-bit embedded processor with a parameterized register window, an average of 10% improvement in application performance and 7% reduction in system power was achieved as an eight-register design was scaled from one to four windows.
Rajiv A. Ravindran, Robert M. Senger, Eric D. Marsman, Ganesh S. Dasika, Matthew R. Guthaus, Scott A. Mahlke, Richard B. Brown
CASES5
2003 A 16-bit mixed-signal microsystem with integrated CMOS-MEMS clock reference
abstract
In this work, we report on an unprecedented design where digital, analog, and MEMS technologies are combined to realize a general-purpose single-chip CMOS microsystem. The convergence of these technologies has enabled the development of a low power, portable microinstrument ideally suited for controlling environmental and bio-implantable sensors.
Robert M. Senger, Eric D. Marsman, Michael S. McCorquodale, Fadi H. Gebara, Keith L. Kraver, Matthew R. Guthaus, Richard B. Brown
DAC6