EDBT 2026 Demo / reviewers in the wild / expert
Rob A. Rutenbar
dblp:r/RobARutenbar
· DBLP profile ↗
151ranked-venue papers
19as first author
2since 2021 · last 2024
0009-0006-6193-0537ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 132 · 17 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10Software engineering, systems software and programming languages · 9Applied, interdisciplinary, general and emerging computing · 6 · 2 first-authorArtificial intelligence and machine learning · 5Theory of computation · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
72 papers |
Electronic design automation · 59% Integrated circuit design · 13% Reconfigurable computing and FPGAs · 8% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computing education · 100% |
Topics — the 30 heaviest of 134, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Reconfigurable computing and FPGAs
FPGA accelerator |
0.6 | 5 | 2020 | Video-rate stereo matching using markov random field TRW-S inference on a hybrid CPU+FPGA computing platform · FPGA 2013 Studying the Potential of Automatic Optimizations in the Intel FPGA SDK for OpenCL · FPGA 2020 A Pixel-Parallel Virtual-Image Architecture for High Performance and Power Efficient Graph Cuts Inference · FPGA 2019 |
Electronic design automation
high-level synthesis |
0.4 | 3 | 2020 | Studying the Potential of Automatic Optimizations in the Intel FPGA SDK for OpenCL · FPGA 2020 OASYS: a framework for analog circuit synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1989 A Prototype Framework for Knowledge-Based Analog Circuit Synthesis · DAC 1987 |
Hardware accelerators and domain-specific architectures
vision accelerator |
0.4 | 1 | 2019 | A Pixel-Parallel Virtual-Image Architecture for High Performance and Power Efficient Graph Cuts Inference · FPGA 2019 |
Electronic design automation › timing analysis
statistical timing analysis |
0.4 | 5 | 2010 | Why Quasi-Monte Carlo is Better Than Monte Carlo or Latin Hypercube Sampling for Statistical Circuit Analysis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010 Probabilistic Interval-Valued Computation: Toward a Practical Surrogate for Statistics Inside CAD Tools · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 Digital Circuit Design Challenges and Opportunities in the Era of Nanoscale CMOS · Proc. IEEE 2008 |
Electronic design automation
physical design |
0.4 | 18 | 2005 | Timing-driven placement by grid-warping · DAC 2005 A synthesis flow toward fast parasitic closure for radio-frequency integrated circuits · DAC 2004 Large-scale placement by grid-warping · DAC 2004 |
Hardware reliability and fault tolerance
process variation |
0.3 | 3 | 2013 | Efficient Spatial Pattern Analysis for Variation Decomposition Via Robust Sparse Regression · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2013 Virtual Probe: A Statistical Framework for Low-Cost Silicon Characterization of Nanoscale Integrated Circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011 Beyond Low-Order Statistical Response Surfaces: Latent Variable Regression for Efficient, Highly Nonlinear Fitting · DAC 2007 |
Electronic design automation
design for manufacturability |
0.3 | 3 | 2013 | Automatic clustering of wafer spatial signatures · DAC 2013 Bayesian virtual probe: minimizing variation characterization cost for nanoscale IC technologies via Bayesian inference · DAC 2010 Efficient handling of operating range and manufacturing linevariations in analog cell synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2000 |
Electronic design automation
analog circuit synthesis |
0.3 | 12 | 2007 | Hierarchical Modeling, Optimization, and Synthesis for System-Level Analog and RF Designs · Proc. IEEE 2007 Remembrance of circuits past: macromodeling by data mining in large analog design spaces · DAC 2002 Anaconda: simulation-based synthesis of analog circuits viastochastic pattern search · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2000 |
Electronic design automation
yield analysis |
0.2 | 2 | 2013 | Automatic clustering of wafer spatial signatures · DAC 2013 Oil fields, hedge funds, and drugs · DAC 2009 |
Computing education › online education
massive open online courses |
0.2 | 1 | 2014 | The First EDA MOOC: Teaching Design Automation to Planet Earth · DAC 2014 |
Electronic design automation
design optimization |
0.2 | 2 | 2009 | Statistical Blockade: Very Fast Statistical Simulation and Modeling of Rare Circuit Events and Its Application to Memory Design · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2009 Probabilistic Interval-Valued Computation: Toward a Practical Surrogate for Statistics Inside CAD Tools · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 |
Hardware accelerators and domain-specific architectures › signal processing accelerator
speech recognition accelerator |
0.2 | 2 | 2009 | A multi-fpga 10x-real-time high-speed search engine for a 5000-word vocabulary speech recognizer · FPGA 2009 A 1000-word vocabulary, speaker-independent, continuous live-mode speech recognizer implemented in a single FPGA · FPGA 2007 |
Electronic design automation › physical design › routing
FPGA routing |
0.2 | 5 | 2004 | A Comparative Study of Two Boolean Formulations of FPGA Detailed Routing Constraints · IEEE Trans. Computers 2004 sub-SAT: a formulation for relaxed Boolean satisfiability with applications in routing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2003 A new FPGA detailed routing approach via search-based Booleansatisfiability · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2002 |
Mathematical optimization › statistical estimation › regression
sparse regression |
0.2 | 1 | 2013 | Efficient Spatial Pattern Analysis for Variation Decomposition Via Robust Sparse Regression · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2013 |
Performance modeling and evaluation
simulation |
0.1 | 3 | 2010 | Why Quasi-Monte Carlo is Better Than Monte Carlo or Latin Hypercube Sampling for Statistical Circuit Analysis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010 Interval-Valued Reduced-Order Statistical Interconnect Modeling · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 Fast Interval-Valued Statistical Modeling of Interconnect and Effective Capacitance · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 |
Electronic design automation › circuit simulation › reduced-order modeling
macromodeling |
0.1 | 3 | 2007 | Hierarchical Modeling, Optimization, and Synthesis for System-Level Analog and RF Designs · Proc. IEEE 2007 Remembrance of circuits past: macromodeling by data mining in large analog design spaces · DAC 2002 A case study of synthesis for industrial-scale analog IP: redesign of the equalizer/filter frontend for an ADSL CODEC · DAC 2000 |
Electronic design automation
interconnect modeling |
0.1 | 2 | 2007 | Interval-Valued Reduced-Order Statistical Interconnect Modeling · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 Fast Interval-Valued Statistical Modeling of Interconnect and Effective Capacitance · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 |
Electronic design automation › timing analysis
statistical timing and variation analysis |
0.1 | 2 | 2007 | Beyond Low-Order Statistical Response Surfaces: Latent Variable Regression for Efficient, Highly Nonlinear Fitting · DAC 2007 Probabilistic interval-valued computation: toward a practical surrogate for statistics inside CAD tools · DAC 2006 |
Integrated circuit design › analog and mixed-signal circuits
analog circuit design |
0.1 | 5 | 2006 | Generation of yield-aware Pareto surfaces for hierarchical circuit design space exploration · DAC 2006 Will Moore's Law rule in the land of analog? · DAC 2004 Anaconda: simulation-based synthesis of analog circuits viastochastic pattern search · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2000 |
Integrated circuit design
compressive sensing |
0.1 | 1 | 2011 | Virtual Probe: A Statistical Framework for Low-Cost Silicon Characterization of Nanoscale Integrated Circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011 |
Electronic design automation › hardware verification and test › VLSI testing
manufacturing test |
0.1 | 1 | 2011 | Virtual Probe: A Statistical Framework for Low-Cost Silicon Characterization of Nanoscale Integrated Circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011 |
Electronic design automation › hardware verification and test
hardware verification |
0.1 | 4 | 2008 | Verifying really complex systems: on earth and beyond · DAC 2008 Satisfiability-Based Layout Revisited: Detailed Routing of Complex FPGAs vis Search-Based Boolean SAT · FPGA 1999 A Scanline Data Structure Processor for VLSI Geometry Checking · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1987 |
Electronic design automation › physical design
placement |
0.1 | 4 | 2005 | Timing-driven placement by grid-warping · DAC 2005 Large-scale placement by grid-warping · DAC 2004 Placement by Simulated Annealing on a Multiprocessor · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1987 |
Integrated circuit design › memory circuit design
SRAM design |
0.1 | 1 | 2010 | Two Fast Methods for Estimating the Minimum Standby Supply Voltage for Large SRAMs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010 |
Electronic design automation › circuit analysis
statistical circuit analysis |
0.1 | 1 | 2010 | Why Quasi-Monte Carlo is Better Than Monte Carlo or Latin Hypercube Sampling for Statistical Circuit Analysis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010 |
Performance modeling and evaluation › statistical analysis
statistical modeling |
0.1 | 1 | 2010 | Bayesian virtual probe: minimizing variation characterization cost for nanoscale IC technologies via Bayesian inference · DAC 2010 |
Integrated circuit design › variation-aware design
variability analysis |
0.1 | 1 | 2010 | Bayesian virtual probe: minimizing variation characterization cost for nanoscale IC technologies via Bayesian inference · DAC 2010 |
Performance modeling and evaluation › simulation
variance reduction |
0.1 | 1 | 2010 | Why Quasi-Monte Carlo is Better Than Monte Carlo or Latin Hypercube Sampling for Statistical Circuit Analysis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010 |
Electronic design automation › physical design › placement
timing-driven placement |
0.1 | 2 | 2005 | Timing-driven placement by grid-warping · DAC 2005 A synthesis flow toward fast parasitic closure for radio-frequency integrated circuits · DAC 2004 |
Performance modeling and evaluation › simulation
monte carlo methods |
0.1 | 1 | 2009 | Oil fields, hedge funds, and drugs · DAC 2009 |
Methods — techniques the papers use, named apart from their topics
sparse regression · 0.5monte carlo simulation · 0.5pragma-directed optimization · 0.4OpenCL · 0.4push-relabel algorithm · 0.4viterbi search · 0.3statistical blockade · 0.2affine arithmetic · 0.2robust regression · 0.2numerical algorithm · 0.2markov random field · 0.2l-method · 0.2hierarchical clustering · 0.2TRW-S inference · 0.2beam search · 0.1acoustic modeling · 0.1thresholding and counting transformation · 0.0linear relaxation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Area-Efficient Iterative Logarithmic Approximate Multipliers for IEEE 754 and Posit NumbersabstractThe IEEE 754 standard for floating-point (FP) arithmetic is widely used for real numbers. Recently, a variant called posit was proposed to improve the precision around 1 and −1. Since FP multiplication requires high computational complexity, various algorithmic approaches and hardware accelerator solutions have been explored. In this context, this article proposes a novel area-efficient logarithmic multiplier architecture for different real number formats, which also provides a significant and useful accuracy/latency tradeoff at runtime. To reduce the logic area in field-programmable gate arrays (FPGAs), this article offers two innovations: applying logarithm to only a single operand and mitigating the accuracy drop caused by this modification with advanced error converging and operand selection schemes. Our multiplier design for single-precision FP (SPFP) numbers uses 58% fewer hardware resources than the iterative Mitchell’s multiplier (IMM) design of Babić et al. extended for SPFP numbers. The error falls within 0.5% when the number of iterations reaches 5. In JPEG, our SPFP multiplier with four iterations produces nearly identical image quality results to the conventional exact multiplier. We further show how to merge two SPFP multipliers for double-precision FP (DPFP) multiplication. This DPFP multiplier design reduces the hardware resources of the IMM design extended for DPFP numbers by 60%. Finally, we demonstrate how our SPFP multiplier design can be slightly modified for 32-bit posit multiplication. It achieves a significantly higher accuracy by increasing the number of iterations compared to state-of-the-art approximate posit multiplier designs. Sunwoong Kim, Cameron James Norris, James Oelund, Rob A. Rutenbar |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2022 | HPVM2FPGA: Enabling True Hardware-Agnostic FPGA ProgrammingabstractCurrent FPGA programming tools require extensive hardware-specific manual code tuning to achieve performance, which is intractable for most software application teams. We present HPVM2FPGA, a novel end-to-end compiler and auto-tuning system that can automatically tune hardware-agnostic programs for FPGAs. HPVM2FPGA uses a hardware-agnostic abstraction of parallelism as an intermediate representation (IR) to represent hardware-agnostic programs. HPVM2FPGA's powerful optimization framework uses sophisticated compiler optimizations and design space exploration (DSE) to automatically tune a hardware-agnostic program for a given FPGA. HPVM2FPGA is able to support software programmers by shifting the burden of performing hardware-specific optimizations to the compiler and DSE. We show that HPVM2FPGA can achieve up to 33×speedup compared to unoptimized baselines and can match the performance of hand-tuned HLS code for three of four benchmarks. We have designed HPVM2FPGA to be a modular and extensible framework, and we expect it to match hand-tuned code for most programs as the system matures with more optimizations. Overall, we believe that it constitutes a solid step closer to fully hardware-agnostic FPGA programming, making it a suitable cornerstone for future FPGA compiler research. Adel Ejjeh, Leon Medvinsky, Aaron Councilman, Hemang Nehra, Suraj Sharma, Vikram S. Adve, Luigi Nardi, Eriko Nurvitadhi, Rob A. Rutenbar |
ASAP | 9 |
| 2020 | Hardware Architecture of a Number Theoretic Transform for a Bootstrappable RNS-based Homomorphic Encryption SchemeabstractHomomorphic encryption (HE) is one of the most promising solutions to secure cloud computing. The number theoretic transform (NTT) that is widely used for convolution operations in HE requires a large amount of computation and has high parallelism, and therefore it has been a good candidate for hardware acceleration. Nevertheless, prior NTT hardware solutions for HE-based applications are impractical in most applications because they do not seriously consider the critical bootstrapping procedure that allows unlimited homomorphic operations on encrypted data. In this paper, we suggest practical bootstrappable parameters, specifically for an established residue number system (RNS)based HE scheme, and apply them to our NTT hardware design. In addition, to limit the size of internal memory for roots of unity increased by the bootstrappable parameters, only a few roots of unity are stored and others are generated on the fly. In our NTT hardware architecture, multiple NTT butterfly units (BUs) are efficiently deployed for high throughput and high resource utilization. In particular, several groups of BUs for respective moduli work in a parallel and pipelined manner, which is effective in an RNS-based HE scheme with a number of moduli. Our implementation on a Xilinx UltraScale FPGA with the bootstrappable parameters achieves a $118 \times$ faster processing speed than a software implementation, and it further provides various trade-off choices such as the number of DSP slices against BRAMs based on available FPGA resources. Sunwoong Kim, Keewoo Lee, Wonhee Cho 0001, Yujin Nam, Jung Hee Cheon, Rob A. Rutenbar |
FCCM | 6 |
| 2020 | Studying the Potential of Automatic Optimizations in the Intel FPGA SDK for OpenCLabstractHigh Level Synthesis (HLS) tools, like the Intel FPGA SDK for OpenCL, improve hardware design productivity and enable efficient design space exploration, by providing simple program directives (pragmas) and/or API calls that allow hardware programmers to use higher-level languages (like HLS-C or OpenCL). However, modern HLS tools sometimes miss important optimizations that are necessary for high performance. In this poster, we present a study of the tradeoffs in HLS optimizations, and the potential of a modern HLS tool in automatically optimizing an application. We perform the study on a generic, 5-stage camera ISP pipeline using the Intel FPGA SDK for OpenCL and an Arria 10 FPGA Dev Kit. We show that automatic optimizations in the HLS tool are valuable, achieving up to 2.7x speedup over equivalent CPU execution. With further hand tuning, however, we can achieve up to 36.5x speedup over CPU. We draw several specific lessons about the effectiveness of automatic optimizations guided by simple directives and about the nature of manual rewriting required for high performance. Finally, we conclude that there is a gap in the current potential of HLS tools which needs to be filled by next-gen research. Adel Ejjeh, Vikram S. Adve, Rob A. Rutenbar |
FPGA | 3 |
| 2020 | A Scalable Bayesian Inference Accelerator for Unsupervised LearningabstractThis article consists only of a collection of slides from the author's conference presentation. Glenn G. Ko, Yuji Chai, Marco Donato, Paul N. Whatmough, Thierry Tambe, Rob A. Rutenbar, Gu-Yeon Wei, David Brooks 0001 |
Hot Chips Symposium | 6 |
| 2019 | A Virtual Image Accelerator for Graph Cuts Inference on FPGAabstractGraph Cuts is a popular technique for Maximum A Posteriori inference in computer vision. It transforms a Markov Random Field problem into a network flow problem, solved via the Push-Relabel algorithm. While attractively simple, the large size of a typical image and the large number of necessary pixel-level iterations render the technique computationally expensive. Prior accelerator attempts have been reported with GPUs and FPGAs. In 2017, we demonstrated the first pixel-parallel architecture on FPGA, but limited to only 256-pixel images. This paper extends this pixel-parallel concept and proposed a Virtual-Image architecture which solves the size limitation. We demonstrate the first working virtual-image Graph Cuts accelerator, implemented on a state of the art FPGA, applied to standard benchmark images for a background segmentation task. The design is 11-13× faster than other FPGA designs, and slightly faster than a modern GPU benchmark by about 30%. Tianqi Gao, Rob A. Rutenbar |
ASAP | 2 |
| 2019 | FlexGibbs: Reconfigurable Parallel Gibbs Sampling Accelerator for Structured GraphsabstractMany consider one of the key components to the success of deep learning as its compatibility with existing accelerators, mainly GPU. While GPUs are great at handling linear algebra kernels commonly found in deep learning, they are not the optimal architecture for handling unsupervised learning methods such as Bayesian models and inference. As a step towards, achieving better understanding of architectures for probabilistic models, Gibbs sampling, one of the most commonly used algorithms for Bayesian inference, is studied with a focus on parallelism that converges to the target distribution and parameterized components. We propose FlexGibbs, a reconfigurable parallel Gibbs sampling inference accelerator for structured graphs. We designed an architecture optimal for solving Markov Random Field tasks using an array of parallel Gibbs samplers, enabled by chromatic scheduling. We show that for sound source separation application, FlexGibbs configured on the FPGA fabric of Xilinx Zync CPU-FPGA SoC achieved Gibbs sampling inference speedup of 1048x and 99.85% reduction in energy over running it on ARM Cortex-A53. Glenn G. Ko, Yuji Chai, Rob A. Rutenbar, David Brooks 0001, Gu-Yeon Wei |
FCCM | 3 |
| 2019 | A Pixel-Parallel Virtual-Image Architecture for High Performance and Power Efficient Graph Cuts InferenceabstractA Pixel-Parallel Virtual-Image Architecture for High Performance and Power Efficient Graph Cuts Inference Tianqi Gao, University of Illinois Urbana Champaign Rob A. Rutenbar, University of Pittsburgh Contact: [email protected] Graph Cuts is a popular technique for Maximum A Posteriori inference in computer vision. It transforms a Markov Random Field problem into a network flow problem, solved via the Push-Relabel algorithm. While attractively simple, the large size of a typical image and the large number of necessary pixel-level iterations render the technique computationally expensive. Prior accelerator attempts have been reported with GPUs and FPGAs. In [1], the first pixel-parallel architecture was demonstrated in FPGA, but limited to only 256-pixel images. This paper extends this pixel-parallel concept and makes following contributions: a "Virtual Image" architecture solves the size limitation: large images are decomposed into "tiles", and "stacked" on the physical processor array; appropriate addressing mechanisms handle virtual pixels and a range of tile edge effects; scaling up the processor array to 1536 pixels; a novel and hardware-friendly heuristic shortens the convergence. We demonstrate the first working virtual-image Graph Cuts accelerator, applied to standard 640x480 images. Scaling up the hardware and the new heuristic bring 6.5x and 1.65x speedups respectively compared with [1]. The design is 7-20x faster than prior FPGA designs, and roughly comparable in speed to a modern GPU benchmark. However, the architecture also offers significant performance-per-unit-power advantages. Formulating a figure of merit particularly for Graph Cuts inference - Graph Cuts per second per Watt - our architecture is about 4 times better than other implementations. Keywords: FPGA; Machine learning; Hardware Acceleration; Computer Vision DOI: https://doi.org/10.1145/3289602.3293948 Tianqi Gao, Rob A. Rutenbar |
FPGA | 2 |
| 2019 | Accelerating Bayesian Inference on Structured Graphs Using Parallel Gibbs SamplingabstractBayesian models and inference is a class of machine learning that is useful for solving problems where the amount of data is scarce and prior knowledge about the application allows you to draw better conclusions. However, Bayesian models often requires computing high-dimensional integrals and finding the posterior distribution can be intractable. One of the most commonly used approximate methods for Bayesian inference is Gibbs sampling, which is a Markov chain Monte Carlo (MCMC) technique to estimate target stationary distribution. The idea in Gibbs sampling is to generate posterior samples by iterating through each of the variables to sample from its conditional given all the other variables fixed. While Gibbs sampling is a popular method for probabilistic graphical models such as Markov Random Field (MRF), the plain algorithm is slow as it goes through each of the variables sequentially. In this work, we describe a binary label MRF Gibbs sampling inference architecture and extend it to 64-label version capable of running multiple perceptual applications, such as sound source separation and stereo matching. The described accelerator employs a chromatic scheduling of variables to parallelize all the conditionally independent variables to 257 samplers, implemented on the FPGA portion of a CPU-FPGA SoC. For real-time streaming sound source separation task, we show the hybrid CPU-FPGA implementation is 230x faster than a commercial mobile processor, while maintaining a recommended latency under 50 ms. The 64-label version showed 137x and 679x speedups for binary label MRF Gibbs sampling inference and 64 labels, respectively. Glenn G. Ko, Yuji Chai, Rob A. Rutenbar, David Brooks 0001, Gu-Yeon Wei |
FPL | 3 |
| 2019 | An Area-Efficient Iterative Single-Precision Floating-Point Multiplier Architecture for FPGAabstractApproximate multipliers have been widely used in critical applications, such as machine learning and multimedia, which are tolerant to approximation errors. This paper proposes a novel single-precision floating-point (SPFP) multiplication algorithm and its architecture. The proposed work approximates only one of the operands to reduce the number of logic blocks and iteratively compensates the approximation error to achieve acceptable error ranges in applications. To reduce the accuracy degradation by the single operand approximation, a rounding scheme and an operand selection scheme are additionally introduced. Compared with the widely-known previous iterative Mitchell design, our proposed SPFP multiplier design decreases the numbers of look up tables (LUTs) and flip flops (FFs) by 55% and 59% respectively, and shows two cycles shorter latency. The accuracy of our design becomes close to that of the iterative Mitchell design as the number of iterations increases, and it always meets the error tolerance of 1% when the number of iterations is four. Sunwoong Kim, Rob A. Rutenbar |
ACM Great Lakes Symposium on VLSI | 2 |
| 2018 | Accelerator Design with Effective Resource Utilization for Binary Convolutional Neural Networks on an FPGAabstractIn binary convolutional neural networks (BCNN), arithmetic operations are replaced by bitwise operations and the required memory size is greatly reduced, which is a good opportunity to accelerate training or inference on FPGAs. This paper proposes a BCNN architecture with a single engine that achieves high resource utilization. The proposed design deploys a large number of processing elements in parallel to increase throughput, and a forwarding scheme to increase resource utilization on the existing engine. In addition, we demonstrate a novel reuse scheme to make fully-connected layers exploit the same engine. The proposed design is combined with an inference environment for comparison and implemented on a Xilinx XCVU190 FPGA. The implemented design uses 61k look-up tables (LUTs), 45k flip-flops (FFs), and 13.9Mbit block RAM (BRAM). In addition, it achieves 61.6 GOPS/kLUT at 240MHz, which is 1.16 times higher than that of the best prior BCNN design, even though it uses a single engine without optimal configurations on each layer. Sunwoong Kim, Rob A. Rutenbar |
FCCM | 2 |
| 2018 | Real-Time and Low-Power Streaming Source Separation Using Markov Random FieldabstractMachine learning (ML) has revolutionized a wide range of recognition tasks, ranging from text analysis to speech to vision, most notably in cloud deployments. However, mobile deployment of these ideas involves a very different category of design problems. In this article, we develop a hardware architecture for a sound source separation task, intended for deployment on a mobile phone. We focus on a novel Markov random field (MRF) sound source separation algorithm that uses expectation-maximization and Gibbs sampling to learn MRF parameters on the fly and infer the best separation of sources. The intrinsically iterative algorithm suggests challenges for both speed and power. A real-time streaming FPGA implementation runs at 150MHz with 207KB RAM, achieves a speed-up of 22× over a software reference, performs with an SDR of up to 7.021dB with 1.601ms latency, and exhibits excellent perceived audio quality. A 45nm CMOS ASIC virtual prototype simulated at 20MHz shows that this architecture is small (<10 million gates) and consumes only 70mW, which is less than 2% of the power of an ARM Cortex-A9 software version. To the best of our knowledge, this is the first Gibbs sampling inference accelerator designed in conventional FPGA/ASIC technology that targets a realistic mobile perceptual application. Glenn G. Ko, Rob A. Rutenbar |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2017 | Toward a pixel-parallel architecture for graph cuts inference on FPGAabstractThe method of Graph Cuts converts a Maximum a Posteriori (MAP) inference problem on a Markov Random Field (MRF) into a network flow, which can be solved efficiently. Many computer vision problems can be conveniently cast as an inference task to find most likely labels for pixels. The method is widely used, but computationally burdensome. Prior accelerator attempts have failed to exploit the problem's attractive, maximum available parallelism: push-relabel flow solvers can run in parallel across every pixel. This paper describes the design and implementation of the first pixel-parallel Graph Cuts inference engine. Our prototype implements a 256-pixel tile of an image, implemented as 256 locally-connected pixel processors. A checkerboard scheduling scheme allows for maximum parallelism while avoiding critical data dependencies. A 150MHz implementation on an FPGA can solve a segmentation task in 6 microseconds. We also discuss strategies for extending our prototype to larger "virtual" images that span more than the physical extent of the inference tile. Our model suggests 2-40× speedups compared with previous accelerator experiments. To the best of our knowledge, this is the first fully functional, pixelparallel accelerator demonstration for Graph Cuts inference. Tianqi Gao, Jungwook Choi, Shang-nien Tsai, Rob A. Rutenbar |
FPL | 4 |
| 2017 | A case study of machine learning hardware: Real-time source separation using Markov Random Fields via sampling-based inferenceabstractWe explore sound source separation to isolate human voice from background noise on mobile phones, e.g. talking on your cell phone in an airport. The challenges involved are real-time execution and power constraints. As a solution, we present a novel hardware-based sound source separation implementation capable of real-time streaming performance. The implementation uses a recently introduced Markov Random Field (MRF) inference formulation of foreground/background separation, and targets voice separation on mobile phones with two microphones. We demonstrate a real-time streaming FPGA implementation running at 150 MHz with total of 207 KB RAM. Our implementation achieves a speedup of 20× over a conventional software implementation, achieves an SDR of 6.655 dB with 1.601 ms latency, and exhibits excellent perceived audio quality. A virtual ASIC design shows that this architecture is quite small (less than 10M gates), consumes only 69.977 mW running at 20 MHz (52× less than an ARM Cortex-A9 software reference), and appears amenable to additional optimization for power. Glenn G. Ko, Rob A. Rutenbar |
ICASSP | 2 |
| 2016 | Configurable and scalable belief propagation accelerator for computer visionabstractWe demonstrate a novel FPGA-based accelerator architecture that can tackle a range of standard computer vision (CV) problems, with scalable performance and attractive speedups. The architecture relies on multiple pipelined processing elements (PEs) that can be configured to support various belief propagation (BP) settings for different CV tasks. Inside each PE, innovative implementation of Jump Flooding for efficient computation of BP solves the core configurability challenge. A novel block-parallel memory interface supports parallelization by distributing BP inference workloads across the PEs. Experimental results demonstrate that our accelerator achieves scalable performance with 11-41× speedup over standard sequential CPU implementations across a subset of well-known Middlebury and OpenGM benchmarks, with no compromise in quality of inference results. To the best of our knowledge, this is the first FPGA hardware implementation of BP capable of running a range of standard CV benchmarks with significant speedups. Jungwook Choi, Rob A. Rutenbar |
FPL | 2 |
| 2016 | Analysis of error resiliency of belief propagation in computer visionabstractProbabilistic inference is a versatile tool to solve a large variety of pixel-labeling problems in computer vision such as stereo matching and image denoising. Belief Propagation (BP) is an effective method for such inference tasks, and has also shown attractive error-resilience properties—the ability to converge to usable solutions in the presence of low-level hardware errors. This is of increasing interest, as the looming end of Moore's Law scaling brings with it a vast increase in the statistical variability of nanoscale circuit fabrics. In this work we seek to understand why certain combinations of BP and error-resilience mechanisms work so well in practice. We focus on Algorithmic Noise Tolerance (ANT) techniques for the resilience mechanisms, and Max-Product BP for inference. We analyze the error characteristics of BP in this hardware context, derive novel asymptotic error bounds, and provide theoretical reasoning to explain why ANT works well in this BP context. Experimental results from detailed resilient-BP simulations for various stereo matching tasks offer empirical support for this analysis. Jungwook Choi, Ameya Patil 0001, Rob A. Rutenbar, Naresh R. Shanbhag |
ICASSP | 3 |
| 2016 | Keynote address Wednesday: Hardware inference accelerators for machine learningabstractMachine learning (ML) technologies have revolutionized the ways in which we interact with large-scale, imperfect, real-world data. As a result, there is rising interest in opportunities to implement ML efficiently in custom hardware. We have designed hardware for one broad class of ML techniques: Inference on Probabilistic Graphical Models (PGMs). In these graphs, labels on nodes encode what we know and “how much” we believe it; edges encode belief relationships among labels; statistical inference answers questions such as “if we observe some of the labels in the graph, what are most likely labels on the remainder?” These problems are interesting because they can be very large (e.g., every pixel in an image is one graph node) and because we need answers very fast (e.g., at video frame rates). Inference done as iterative Belief Propagation (BP) can be efficiently implemented in hardware, and we demonstrate several examples from current FPGA prototypes. We have the first configurable, scalable parallel architecture capable of running a range of standard vision benchmarks, with speedups up to 40X over conventional software. We also show that BP hardware can be made remarkably tolerant to the low-level statistical upsets expected in end-of-Moore's-Law nanoscale silicon and post-silicon circuit fabrics, and summarize some effective resilience mechanisms in our prototypes. Rob A. Rutenbar |
ITC | 1 |
| 2016 | Video-Rate Stereo Matching Using Markov Random Field TRW-S Inference on a Hybrid CPU+FPGA Computing PlatformabstractWe demonstrate a video-rate stereo matching system implemented on a hybrid CPU+field-programmable gate array (FPGA) platform (Convey HC-1). Stereo matching is a fundamental problem of computer vision, and emerging applications, such as 3-D gesture recognition and automotive navigation, demand fast and high-quality stereo matching. Markov random field (MRF)-based approaches are widely used, but conventional software solvers are slow. Belief propagation (BP) solvers, which use patterns of local message passing on MRFs, have been studied in hardware, but their performance is unreliable. We show how a superior method, sequential tree-reweighted message passing (TRW-S), can be rendered in hardware. TRW-S has reliable convergence, guaranteed by its so-called sequential computation. Analysis reveals many opportunities for TRW-S hardware acceleration. Starting from the core architecture for streaming TRW-S, we describe the end-to-end system engineering for full video stereo matching. We partition the stereo matching procedure across the CPU and the FPGAs and apply frame level optimizations, such as message reuse based on scene change detection, frame level parallelization, and function level pipelining. The experimental results show that our system achieves a speed of 22.8 frames/s for a challenging QVGA video stereo matching task. We noticed that our system is significantly faster than several recent GPU/ASIC implementations of similar stereo inference methods based on BP. Jungwook Choi, Rob A. Rutenbar |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Error Resilient and Energy Efficient MRF Message-Passing-Based Stereo MatchingabstractMessage-passing-based inference algorithms have immense importance in real-world applications. In this paper, error resiliency of a message passing based Markov random field (MRF) stereo matching hardware is explored and enhanced through the application of statistical error compensation. Error resiliency is of particular interest for subnanometer and postsilicon devices. The inherent robustness of iteration-based MRF inference algorithms is explored and shows that small errors are tolerable, while large errors degrade the performance significantly. Based on these error characteristics, algorithmic noise tolerance (ANT) has been applied at the arithmetic, iteration, and system levels. Introducing timing errors via voltage overscaling, at the arithmetic level, results show that the ANT-based hardware can tolerate an error rate of 21.3%, with performance degradation of only 3.5% at an overhead of 97.4%, compared with an error-free hardware with an energy savings of 39.7%. To reduce compensation complexity, iteration and system-level compensation was explored. Results show that, compared with arithmetic level, system-level compensation reduces overhead to 59%, while maintaining stereo matching performance with only 2.5% degradation with 16% additional power savings. These results are verified via FPGA emulation with timing errors induced within the message passing unit via relaxed synthesis. Eric P. Kim, Jungwook Choi, Naresh R. Shanbhag, Rob A. Rutenbar |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | Fast hierarchical implementation of sequential tree-reweighted belief propagation for probabilistic inferenceabstractMaximum a posteriori probability (MAP) inference on Markov random fields (MRF) is the basis of many computer vision applications. Sequential tree-reweighted belief propagation (TRW-S) has been shown to provide very good inference quality and strong convergence properties. However, software TRW-S solvers are slow due to the algorithm's high computational requirements. A state-of-the-art FPGA implementation has been developed recently, which delivers substantial speedup over software. In this paper, we improve upon the TRW-S algorithm by using a multi-level hierarchical MRF formulation. We demonstrate the benefits of Hierarchical-TRW-S over TRW-S, and incorporate the proposed improvements on a Convey HC-1 CPU-FPGA hybrid platform. Results using four Middlebury stereo vision benchmarks show a 21% to 53% reduction in inference time compared with the state-of-the-art TRW-S FPGA implementation. To the best of our knowledge, this is the fastest hardware implementation of TRW-S reported so far. Skand Hurkat, Jungwook Choi, Eriko Nurvitadhi, José F. Martínez, Rob A. Rutenbar |
FPL | 5 |
| 2015 | Analog Circuit and Layout Synthesis RevisitedabstractIn the first decade of the twenty first century, the first generation of analog synthesis tools - for circuit sizing and optimization and for physical design - moved from academic research projects, to startups, to integration on various standard platforms. In addition, we began to see some concerted efforts to formulate the verification problem in formal terms, albeit on rather simplified version of designs. In this invited talk I will try to summarize what we got right in that first generation effort, but also, what we got wrong. The "right" part is the idea that all design is optimization: continuous, combinatorial, geometric, etc. Formulating tough analog circuit and layout problems as the "right" optimization problem(s) got us the first generation of tools that did anything right. The "wrong" part was an incomplete appreciation of the importance of designer use models - how real people do real designs in this business. Closing that gap is one remaining challenge. Another is the leap to non-planar end-of-roadmap CMOS technologies, where lithographic and manufacturability concerns combine to create some difficult problems, and new opportunities for tool innovation. Rob A. Rutenbar |
ISPD | 1 |
| 2014 | Efficiently Enforcing Diversity in Multi-Output Structured PredictionabstractThis paper proposes a novel method for efficiently generating multiple diverse predictions for structured prediction problems. Existing methods like SDPPs or DivMBest work by making a series of predictions where each prediction is made after considering the predictions that came before it. Such approaches are inherently sequential and computationally expensive. In contrast, our method, Diverse Multiple Choice Learning, learns a set of models to make multiple independent, yet diverse, predictions at testtime. We achieve this by including a diversity encouraging term in the loss function used for training the models. This approach encourages diversity in the predictions while preserving computational efficiency at test-time. Experimental results on a number of challenging problems show that our method learns models that not only predict more diverse results than competing methods, but are also able to generalize better and produce results with high test accuracy. Abner Guzmán-Rivera, Pushmeet Kohli, Dhruv Batra, Rob A. Rutenbar |
AISTATS | 4 |
| 2014 | The First EDA MOOC: Teaching Design Automation to Planet EarthabstractMassive Open Online Courses (MOOCs) can deliver advanced course material at planetary scale, combining internet-based video content delivery, and cloud-based assignments. From March to May 2013, I taught the world's first EDA MOOC, entitled VLSI CAD: Logic to Layout, based on roughly 20 years of experience teaching electronic design automation in a conventional face-to-face classroom setting. Over 17,000 participants registered for this MOOC. This paper summarizes my experience with teaching EDA at planetary scale: how we covered ASIC synthesis, verification, layout, and timing; how we built cloud resources to enable students to experiment with open-source tools; how we designed software projects and deployed cloud-based auto-graders to support realistic EDA tool projects. The paper also discusses what MOOCs could mean to the dynamism of the EDA community. Rob A. Rutenbar |
DAC | 1 |
| 2014 | A robust message passing based stereo matching kernel via system-level error resiliencyabstractIn this paper, we present an error resilient Markov random field (MRF) message passing based stereo matching hardware (HW) architecture. Previously, algorithmic noise tolerance (ANT) has been applied at the arithmetic level of the reparameterize unit and showed greatly enhanced robustness of message passing inference based architectures. In this work, correction was targeted at the system level to reduce correction overhead while maintaining performance. An erroneous FPGA based accelerator was employed as our emulation platform. Through relaxed synthesis, we show that timing errors occur within the message passing unit, and are successfully compensated. Error correction has been implemented at several hierarchical levels, including end of iteration, and the final depth map output. Significant enhancement in robustness is achieved with minimal correction overhead. Compared to HW error compensation at the arithmetic level, system level error compensation reduces overhead by more than 50 %, while maintaining stereo matching performance with only 3.8 % degradation. Eric P. Kim, Jungwook Choi, Naresh R. Shanbhag, Rob A. Rutenbar |
ICASSP | 4 |
| 2013 | Automatic clustering of wafer spatial signaturesabstractIn this paper, we propose a methodology based on unsupervised learning for automatic clustering of wafer spatial signatures to aid yield improvement. Our proposed methodology is based on three steps. First, we apply sparse regression to automatically capture wafer spatial signatures by a small number of features. Next, we apply an unsupervised hierarchical clustering algorithm to divide wafers into a few clusters where all wafers within the same cluster are similar. Finally, we develop a modified L-method to determine the appropriate number of clusters from the hierarchical clustering result. The accuracy of the proposed methodology is demonstrated by several industrial data sets of silicon measurements. Wangyang Zhang, Xin Li 0001, Sharad Saxena, Andrzej J. Strojwas, Rob A. Rutenbar |
DAC | 5 |
| 2013 | Video-rate stereo matching using markov random field TRW-S inference on a hybrid CPU+FPGA computing platformabstractWe demonstrate a video-rate stereo matching system implemented on a hybrid CPU+FPGA platform (Convey HC-1). Emerging applications such as 3D gesture recognition and automotive navigation demand fast and high quality stereo vision. We describe a custom hardware-accelerated Markov Random Field inference system for this task. Starting from a core architecture for streaming tree-reweighted message passing (TRW-S) inference, we describe the end-to-end system engineering needed to move from this single frame message update to full stereo video. We partition the stereo matching procedure across the CPU and the FPGAs, and apply both function-level pipelining and frame-level parallelism to achieve the required speed. Experimental results show that our system achieves a speed of 12 frames per second for challenging video stereo matching tasks. We note that this appears to be the first implementation of TRW-S inference at video rates, and that our system is also significantly faster than several recent GPU implementations of similar stereo inference methods based on belief propagation (BP). Jungwook Choi, Rob A. Rutenbar |
FPGA | 2 |
| 2013 | EMERALD: Characterization of emerging applications and algorithms for low-power devicesabstractCompute-intensive applications are emerging in intelligent home, retail store and automotive industries. These applications are becoming more sophisticated with new features rich in audio, video, image, and machine learning capabilities that demand heavy computations. We present the EMERALD (EMERging Applications and algorithms for Low power Device) workload suite. We profile the workloads to show the hotspot functions that are candidates for hardware accelerators. Chuanjun Zhang, Glenn G. Ko, Jungwook Choi, Shang-nien Tsai, Minje Kim 0001, Abner Guzmán-Rivera, Rob A. Rutenbar, Paris Smaragdis, Mi Sun Park, Narayanan Vijaykrishnan, Hongyi Xin, Onur Mutlu, Bin Li 0018, Li Zhao 0002 |
ISPASS | 7 |
| 2013 | FPGA acceleration of Markov Random Field TRW-S inference for stereo matching
Jungwook Choi, Rob A. Rutenbar |
MEMOCODE | 2 |
| 2013 | Efficient Spatial Pattern Analysis for Variation Decomposition Via Robust Sparse RegressionabstractIn this paper, we propose a new technique to achieve accurate decomposition of process variation by efficiently performing spatial pattern analysis. We demonstrate that the spatially correlated systematic variation can be accurately represented by the linear combination of a small number of templates. Based on this observation, an efficient sparse regression algorithm is developed to accurately extract the most adequate templates to represent spatially correlated variation. In addition, a robust sparse regression algorithm is proposed to automatically remove measurement outliers. We further develop a fast numerical algorithm that may reduce the computational time by several orders of magnitude over the traditional direct implementation. Our experimental results based on both synthetic and silicon data demonstrate that the proposed sparse regression technique can capture spatially correlated variation patterns with high accuracy and efficiency. Wangyang Zhang, Karthik Balakrishnan, Xin Li 0001, Duane S. Boning, Sharad Saxena, Andrzej J. Strojwas, Rob A. Rutenbar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2012 | A High-Rate, Low-Power, ASIC Speech Decoder Using Finite State TransducersabstractThe use of Finite State Transducers in speech recognition has been increasing in recent years. Their application in speech decoding allows for a tradeoff between larger memory requirements and less run-time computation. We believe that this paradigm is especially well suited for a highspeed, energy-efficient hardware solution where customized caching, reduced bit widths, and prefetching can be used to mitigate the effect of the increased model size. We present a virtual silicon prototype for a novel hardware architecture that, using these optimizations and running at 556MHz, is capable of performing recognition on the Wall Street Journal 60K- word speech model with 92.3 percent accuracy a speed 127 times faster than real time, while consuming less than 0.5 watts. Jeffrey R. Johnston, Rob A. Rutenbar |
ASAP | 2 |
| 2012 | Hardware implementation of MRF map inference on an FPGA platformabstractIn this paper, we describe hardware for inference computations on Markov Random Fields (MRFs). MRFs are widely used in applications like computer vision, but conventional software solvers are slow. Belief Propagation (BP) solvers, which use patterns of local message passing on MRFs, have been studied in hardware, but their performance is unreliable. We show how a superior method-Sequential Tree-Reweighted message passing (TRW-S)-can be rendered in hardware. TRW-S has reliable convergence, guaranteed by its so-called “sequential” computation. Analysis reveals many opportunities for TRW-S hardware acceleration. We show how to implement TRW-S in FPGA hardware so that it exploits significant parallelism and memory bandwidth. Our implementation is capable of running a standard stereo vision benchmark at rates approaching 40 frames/sec; this represents the first time TRW-S methods have been accelerated to these speeds on an FPGA platform. Jungwook Choi, Rob A. Rutenbar |
FPL | 2 |
| 2011 | Toward efficient spatial variation decomposition via sparse regressionabstractIn this paper, we propose a new technique to accurately decompose process variation into two different components: (1) spatially correlated variation, and (2) uncorrelated random variation. Such variation decomposition is important to identify systematic variation patterns at wafer and/or chip level for process modeling, control and diagnosis. We demonstrate that spatially correlated variation carries a unique sparse signature in frequency domain. Based upon this observation, an efficient sparse regression algorithm is applied to accurately separate spatially correlated variation from uncorrelated random variation. An important contribution of this paper is to develop a fast numerical algorithm that reduces the computational time of sparse regression by several orders of magnitude over the traditional implementation. Our experimental results based on silicon measurement data demonstrate that the proposed sparse regression technique can capture spatially correlated variation patterns with high accuracy. The estimation error is reduced by more than 3.5× compared to other traditional methods. Wangyang Zhang, Karthik Balakrishnan, Xin Li 0001, Duane S. Boning, Rob A. Rutenbar |
ICCAD | 5 |
| 2011 | Virtual Probe: A Statistical Framework for Low-Cost Silicon Characterization of Nanoscale Integrated CircuitsabstractIn this paper, we propose a new technique, referred to as virtual probe (VP), to efficiently measure, characterize, and monitor spatially-correlated inter-die and/or intra-die variations in nanoscale manufacturing process. VP exploits recent breakthroughs in compressed sensing to accurately predict spatial variations from an exceptionally small set of measurement data, thereby reducing the cost of silicon characterization. By exploring the underlying sparse pattern in spatial frequency domain, VP achieves substantially lower sampling frequency than the well-known Nyquist rate. In addition, VP is formulated as a linear programming problem and, therefore, can be solved both robustly and efficiently. Our industrial measurement data demonstrate the superior accuracy of VP over several traditional methods, including 2-D interpolation, Kriging prediction, and k-LSE estimation. Wangyang Zhang, Xin Li 0001, Frank Liu 0001, Emrah Acar, Rob A. Rutenbar, R. D. (Shawn) Blanton |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2010 | Bayesian virtual probe: minimizing variation characterization cost for nanoscale IC technologies via Bayesian inferenceabstractThe expensive cost of testing and characterizing parametric variations is one of the most critical issues for today's nanoscale manufacturing process. In this paper, we propose a new technique, referred to as Bayesian Virtual Probe (BVP), to efficiently measure, characterize and monitor spatial variations posed by manufacturing uncertainties. In particular, the proposed BVP method borrows the idea of Bayesian inference and information theory from statistics to determine an optimal set of sampling locations where test structures should be deployed and measured to monitor spatial variations with maximum accuracy. Our industrial examples with silicon measurement data demonstrate that the proposed BVP method offers superior accuracy (1.5x error reduction) over the VP approach that was recently developed in [12]. Wangyang Zhang, Xin Li 0001, Rob A. Rutenbar |
DAC | 3 |
| 2010 | Multi-Wafer Virtual Probe: Minimum-cost variation characterization by exploring wafer-to-wafer correlationabstractIn this paper, we propose a new technique, referred to as Multi-Wafer Virtual Probe (MVP) to efficiently model wafer-level spatial variations for nanoscale integrated circuits. Towards this goal, a novel Bayesian inference is derived to extract a shared model template to explore the wafer-to-wafer correlation information within the same lot. In addition, a robust regression algorithm is proposed to automatically detect and remove outliers (i.e., abnormal measurement data with large error) so that they do not bias the modeling results. The proposed MVP method is extensively tested for silicon measurement data collected from 200 wafers at an advanced technology node. Our experimental results demonstrate that MVP offers superior accuracy over other traditional approaches such as VP and EM, if a limited number of measurement data are available. Wangyang Zhang, Xin Li 0001, Emrah Acar, Frank Liu 0001, Rob A. Rutenbar |
ICCAD | 5 |
| 2010 | Analog layout synthesis: what's missing?abstractIn the first decade of the twenty first century, the first generation of analog synthesis tools -- for sizing and for physical design -- moved from academic research projects, to startups, to integration on various standard platforms. In this talk I'll try to summarize what we got right in that first generation effort, but more importantly, what we got wrong. I'll talk in particular about usage models -- how real people do real layouts in this business -- and highlight what I think are the big gaps, the big problems, and the big opportunities. Rob A. Rutenbar |
ISPD | 1 |
| 2010 | Why Quasi-Monte Carlo is Better Than Monte Carlo or Latin Hypercube Sampling for Statistical Circuit AnalysisabstractAt the nanoscale, no circuit parameters are truly deterministic; most quantities of practical interest present themselves as probability distributions. Thus, Monte Carlo techniques comprise the strategy of choice for statistical circuit analysis. There are many challenges in applying these techniques efficiently: circuit size, nonlinearity, simulation time, and required accuracy often conspire to make Monte Carlo analysis expensive and slow. Are we-the integrated circuit community-alone in facing such problems? As it turns out, the answer is “no.” Problems in computational finance share many of these characteristics: high dimensionality, profound nonlinearity, stringent accuracy requirements, and expensive sample evaluation. We perform a detailed experimental study of how one celebrated technique from that domain-quasi-Monte Carlo (QMC) simulation-can be adapted effectively for fast statistical circuit analysis. In contrast to traditional pseudorandom Monte Carlo sampling, QMC uses a (shorter) sequence of deterministically chosen sample points. We perform rigorous comparisons with both Monte Carlo and Latin hypercube sampling across a set of digital and analog circuits, in 90 and 45 nm technologies, varying in size from 30 to 400 devices. We consistently see superior performance from QMC, giving 2× to 8× speedup over conventional Monte Carlo for roughly 1% accuracy levels. We present rigorous theoretical arguments that support and explain this superior performance of QMC. The arguments also reveal insights regarding the (low) latent dimensionality of these circuit problems; for example, we observe that over half of the variance in our test circuits is from unidimensional behavior. This analysis provides quantitative support for recent enthusiasm in dimensionality reduction of circuit problems. Amith Singhee, Rob A. Rutenbar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2010 | Two Fast Methods for Estimating the Minimum Standby Supply Voltage for Large SRAMsabstractThe data retention voltage (DRV) defines the minimum supply voltage for an SRAM cell to hold its state. Intra-die variation causes a statistical distribution of DRV for individual cells in a memory array. We present two fast and accurate methods to estimate the tail of the DRV distribution. The first method uses a new analytical model based on the relationship between DRV and static noise margin. The second method extends the statistical blockade technique to a recursive formulation. It uses conditional sampling for rapid statistical simulation and fits the results to a generalized Pareto distribution (GPD) model. Both the analytical DRV model and the generic GPD model show a good match with Monte Carlo simulation results and offer speedups of up to four or five orders of magnitude over Monte Carlo at the 6σ point. In addition, the two models show a very close agreement with each other at the tail up to 8σ. For error within 5% with a confidence of 95%, the analytical DRV model and the GPD model can predict DRV quantiles out to 8σ and 6.6σ respectively; and for the mean of the estimate, both models offer within 1% error relative to Monte Carlo at the 4σ point. Jiajing Wang, Amith Singhee, Rob A. Rutenbar, Benton H. Calhoun |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2009 | Oil fields, hedge funds, and drugsabstractStatistical analysis is a fundamental method in analysis, design, and optimization of large systems with uncertainties. It is being applied in drug development, analyzing financial markets, search for new oil fields, and many more areas. As the silicon process technology scales to its limits, uncertainty plays an increasing role and has made its way into tools such as yield analysis, statistical time analysis, etc. In this educational panel, we will start with a short tutorial on Monte Carlo methods including their use in EDA. Three experts from distinct application fields will then follow and discuss their experience in using Monte Carlo methods for solving large-scale problems in their domain. It may come as a surprise to some attendees of DAC to learn how much commonality there is between methods used in EDA and other field that seem far-fetched. Patrick Groeneveld, Rob A. Rutenbar, Jed W. Pitera, Erik C. Carlson |
DAC | 2 |
| 2009 | A multi-fpga 10x-real-time high-speed search engine for a 5000-word vocabulary speech recognizerabstractToday's best quality speech recognition systems are implemented in software. These systems fully occupy the resources of a high-end server to deliver results at real-time speed: each hour of audio requires a significant fraction of an hour of computation for recognition. This is profoundly limiting for applications that require extreme recognition speed, for example, high-volume tasks such as video indexing (e.g., YouTube), or high-speed tasks such as triage of homeland security intelligence. We describe the architecture and implementation of one critical component -- the backend search stage -- of a high-speed, large-vocabulary recognizer. Implemented on a multi-FPGA Berkeley Emulation Engine 2 (BEE2) platform, we handle a standard 5000-word Wall Street Journal speech benchmark. Our backend search engine can decode on average 10 times faster than real-time running at 100 MHz, i.e, 10x faster than real-time, with negligible degradation in accuracy, running at a clock rate ~ 30x slower than a conventional server. To the best of our knowledge, this is both the most complex, and the fastest recognizer ever to be realized in a hardware form. Edward C. Lin, Rob A. Rutenbar |
FPGA | 2 |
| 2009 | Virtual probe: A statistically optimal framework for minimum-cost silicon characterization of nanoscale integrated circuitsabstractIn this paper, we propose a new technique, referred to as virtual probe (VP), to efficiently measure, characterize and monitor both inter-die and spatially-correlated intra-die variations in nanoscale manufacturing process. VP exploits recent breakthroughs in compressed sensing [15]-[17] to accurately predict spatial variations from an exceptionally small set of measurement data, thereby reducing the cost of silicon characterization. By exploring the underlying sparse structure in (spatial) frequency domain, VP achieves substantially lower sampling frequency than the well-known (spatial) Nyquist rate. In addition, VP is formulated as a linear programming problem and, therefore, can be solved both robustly and efficiently. Our industrial measurement data demonstrate that by testing the delay of just 50 chips on a wafer, VP accurately predicts the delay of the other 219 chips on the same wafer. In this example, VP reduces the estimation error by up to 10× compared to other traditional methods. Categories and Subject Descriptors B.7.2 [Integrated Circuits]: Design Aids — Verification General Terms Algorithms Xin Li 0001, Rob A. Rutenbar, R. D. (Shawn) Blanton |
ICCAD | 2 |
| 2009 | Profiling large-vocabulary continuous speech recognition on embedded devices: a hardware resource sensitivity analysis
Rob A. Rutenbar |
INTERSPEECH | 2 |
| 2009 | Statistical Blockade: Very Fast Statistical Simulation and Modeling of Rare Circuit Events and Its Application to Memory DesignabstractCircuit reliability under random parametric variation is an area of growing concern. For highly replicated circuits, e.g., static random access memories (SRAMs), a rare statistical event for one circuit may induce a not-so-rare system failure. Existing techniques perform poorly when tasked to generate both efficient sampling and sound statistics for these rare events. Statistical blockade is a novel Monte Carlo technique that allows us to efficiently filter-to block-unwanted samples that are insufficiently rare in the tail distributions we seek. The method synthesizes ideas from data mining and extreme value theory and, for the challenging application of SRAM yield analysis, shows speedups of 10 - 100 times over standard Monte Carlo. Amith Singhee, Rob A. Rutenbar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2008 | Verifying really complex systems: on earth and beyondabstractFunctional verification is a major part of the effort to design electronic systems. Over the years, EDA has developed a suite of tools and methods to address the verification challenges by a patchwork of approaches. However, as the system complexity continues to increase, traditional methods may not be adequate to ensure flawless behavior. In this educational panel, we will explore how complex systems are validated in other areas. Four speakers will cover the verification challenges in airplane design, complex Mars exploration missions, modeling and rendering of movie animations, and the design of continent-wide national power grids. Using real-life examples, each speaker will introduce the general topic, outline the specific verification challenges, and discuss how they are approached in their specific domain. The following discussion will analyze commonalities and differences between the areas and explore lessons to be learned from them for EDA. Andreas Kuehlmann, Anjan Bose, David E. Corman, Rob A. Rutenbar, Robert M. Manning, Anna Newman |
DAC | 4 |
| 2008 | Exploiting Correlation Kernels for Efficient Handling of Intra-Die Spatial Correlation, with Application to Statistical TimingabstractIntra-die manufacturing variations are unavoidable in nanoscale processes. These variations often exhibit strong spatial correlation. Standard grid-based models assume model parameters (grid-size, regularity) in an ad hoc manner and can have high measurement cost. The random Leld model overcomes these issues. However, no general algorithm has been proposed for the practical use of this model in statistical CAD tools. In this paper, we propose a robust and efficient numerical method, based on the Galerkin technique and Karhunen Loeve Expansion, that enables effective use of the model. We test the effectiveness of the technique using a Monte Carlo-based Statistical Static Timing Analysis algorithm, and see errors less than 0.7%, while reducing the number of random variables from thousands to 25, resulting in speedups of up to 100 x. Amith Singhee, Sonia Singhal, Rob A. Rutenbar |
DATE | 3 |
| 2008 | Practical, fast Monte Carlo statistical static timing analysis: why and howabstractStatistical static timing analysis (SSTA) has emerged as an essential tool for nanoscale designs. Monte Carlo methods are universally employed to validate the accuracy of the approximations made in all SSTA tools, but Monte Carlo itself is never employed as a strategy for practical SSTA. It is widely believed to be ldquotoo slowrdquo - despite an uncomfortable lack of rigorous studies to support this belief. We offer the first large-scale study to refute this belief. We synthesize recent results from fast quasi-Monte Carlo (QMC) deterministic sampling and efficient Karhunen-Loeve expansion (KLE) models of spatial correlation to show that Monte Carlo SSTA need not be slow. Indeed, we show for the ISCAS89 circuits, a few hundred, well-chosen sample points can achieve errors within 5%, with no assumptions on gate models, wire models, or the core STA engine, with runtimes less than 90 s. Amith Singhee, Sonia Singhal, Rob A. Rutenbar |
ICCAD | 3 |
| 2008 | A low-power hardware search architecture for speech recognitionabstractHigh-performance speech recognition is extremely computationally expensive, limiting its use in the mobile domain. We therefore propose a low-power hardware speech recognition architecture for mobile applications, exploiting the orders-of-magnitude efficiency improvements dedicated hardware can offer. Our system is based on the Sphinx 3.0 software recognizer developed at Carnegie Mellon University, capable of large-vocabulary, speaker-independent, continuous, real-time speech recognition. We show through cycle-accurate simulation that our hardware, targeting the backend search stage of recognition, is capable of recognizing speech from a 5,000 word vocabulary 1.3 times faster than real-time, within a 196mW power budget. Patrick J. Bourke, Rob A. Rutenbar |
INTERSPEECH | 2 |
| 2008 | Digital Circuit Design Challenges and Opportunities in the Era of Nanoscale CMOSabstractWell-designed circuits are one key ldquoinsulatingrdquo layer between the increasingly unruly behavior of scaled complementary metal-oxide-semiconductor devices and the systems we seek to construct from them. As we move forward into the nanoscale regime, circuit design is burdened to ldquohiderdquo more of the problems intrinsic to deeply scaled devices. How this is being accomplished is the subject of this paper. We discuss new techniques for logic circuits and interconnect, for memory, and for clock and power distribution. We survey work to build accurate simulation models for nanoscale devices. We discuss the unique problems posed by nanoscale lithography and the role of geometrically regular circuits as one promising solution. Finally, we look at recent computer-aided design efforts in modeling, analysis, and optimization for nanoscale designs with ever increasing amounts of statistical variation. Benton H. Calhoun, Yu Cao 0001, Xin Li 0001, Ken Mai, Lawrence T. Pileggi, Rob A. Rutenbar, Kenneth L. Shepard |
Proc. IEEE | 6 |
| 2008 | Probabilistic Interval-Valued Computation: Toward a Practical Surrogate for Statistics Inside CAD ToolsabstractInterval methods offer a general fine-grain strategy for modeling correlated range uncertainties in numerical algorithms. We present a new improved interval algebra that extends the classical affine form to a more rigorous statistical foundation. Range uncertainties now take the form of confidence intervals. In place of pessimistic interval bounds, we minimize the probability of numerical "escape"; this can tighten interval bounds by an order of magnitude while yielding 10-100 times speedups over Monte Carlo. The formulation relies on the following three critical ideas: liberating the affine model from the assumption of symmetric intervals; a unifying optimization formulation; and a concrete probabilistic model. We refer to these as probabilistic intervals for brevity. Our goal is to understand where we might use these as a surrogate for expensive explicit statistical computations. Results from sparse matrices and graph delay algorithms demonstrate the utility of the approach and the remaining challenges. Amith Singhee, Claire Fang Fang, James D. Ma, Rob A. Rutenbar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2007 | Next-Generation Design and EDA Challenges: Small Physics, Big Systems, and Tall Tool-Chains
Rob A. Rutenbar |
ASP-DAC | 1 |
| 2007 | Beyond Low-Order Statistical Response Surfaces: Latent Variable Regression for Efficient, Highly Nonlinear FittingabstractThe number and magnitude of process variation sources are increasing as we scale further into the nano regime. Today's most successful response surface methods limit us to low-order forms -- linear, quadratic -- to make the fitting tractable. Unfortunately, not all variation-al scenarios are well modeled with low-order surfaces. We show how to exploit latent variable regression ideas to support efficient extraction of arbitrarily nonlinear statistical response surfaces. An implementation of these ideas called SiLVR, applied to a range of analog and digital circuits, in technologies from 90 to 45nm, shows significant improvements in prediction, with errors reduced by up to 21X, with very reasonable runtime costs. Amith Singhee, Rob A. Rutenbar |
DAC | 2 |
| 2007 | Statistical blockade: a novel method for very fast Monte Carlo simulation of rare circuit events, and its application
Amith Singhee, Rob A. Rutenbar |
DATE | 2 |
| 2007 | A 1000-word vocabulary, speaker-independent, continuous live-mode speech recognizer implemented in a single FPGAabstractThe Carnegie Mellon In Silico Vox project seeks to move best-quality speech recognition technology from its current software-only form into a range of efficient all-hardware implementations. The central thesis is that, like graphics chips, the application is simply too performance hungry, and too power sensitive, to stay as a large software application. As a first step in this direction, we describe the design and implementation of a fully functional speech-to-text recognizer on a single Xilinx XUP platform. The design recognizes a 1000 word vocabulary, is speaker-independent, recognizes continuous (connected) speech, and is a "live mode" engine, wherein recognition can start as soon as speech input appears. To the best of our knowledge, this is the most complex recognizer architecture ever fully committed to a hardware-only form. The implementation is extraordinarily small, and achieves the same accuracy as state-of-the-art software recognizers, while running at a fraction of the clock speed. Edward C. Lin, Rob A. Rutenbar, Tsuhan Chen |
FPGA | 3 |
| 2007 | Generating small, accurate acoustic models with a modified Bayesian information criterionabstractAlthough Gaussian mixture models are commonly used in acoustic models for speech recognition, there is no standard method for determining the number of mixture components. Most models arbitrarily assign the number of mixture components with little justification. While model selection techniques with a mathematical derivation, such as the Bayesian information criterion (BIC), have been applied, these criteria focus on properly modeling the true distribution of individual tied-states (senones) without considering the entire acoustic model; this leads to suboptimal speech recognition performance. In this paper we present a method to generate statistically-justified acoustic models that consider inter-senone effects by modifying the BIC. Experimental results in the CMU Communicator domain show that in contrast to previous strategies, the new method generates not only attractively smaller acoustic models, but also ones with lower word error rate. Index Terms: acoustic model training, model selection, BIC, Gaussian mixture models Rob A. Rutenbar |
INTERSPEECH | 2 |
| 2007 | Mixed-size placement with fixed macrocells using grid-warpingabstractGrid-warping is a placement strategy based on a novel physical analogy: rather than move the gates to optimize their location, it elastically deforms a model of the 2-D chip surface on which the gates have been coarsely placed via a standard quadratic solve. Although the original warping idea works well for cell-based placement, it works poorly for mixed-size placements with large, fixed macrocells. The new problem is how to avoid elastically deforming gates into illegal overlaps with these background objects. We develop a new lightweight mechanism called "geometric hashing" which relocates gates to avoid these overlaps, but is efficient enough to embed directly in the nonlinear warping optimization. Results from a new placer (WARP3) running on the ISPD 2005 benchmark suite show both good quality and scalability. Zhong Xiu, Rob A. Rutenbar |
ISPD | 2 |
| 2007 | Hierarchical Modeling, Optimization, and Synthesis for System-Level Analog and RF DesignsabstractThe paper describes the recent state of the art in hierarchical analog synthesis, with a strong emphasis on associated techniques for computer-aided model generation and optimization. Over the past decade, analog design automation has progressed to the point where there are industrially useful and commercially available tools at the cell level-tools for analog components with 10-100 devices. Automated techniques for device sizing, for layout, and for basic statistical centering have been successfully deployed. However, successful component-level tools do not scale trivially to system-level applications. While a typical analog circuit may require only 100 devices, a typical system such as a phase-locked loop, data converter, or RF front-end might assemble a few hundred such circuits, and comprise 10 000 devices or more. And unlike purely digital systems, mixed-signal designs typically need to optimize dozens of competing continuous-valued performance specifications, which depend on the circuit designer's abilities to successfully exploit a range of nonlinear behaviors across levels of abstraction from devices to circuits to systems. For purposes of synthesis or verification, these designs are not tractable when considered "flat." These designs must be approached with hierarchical tools that deal with the system's intrinsic design hierarchy. This paper surveys recent advances in analog design tools that specifically deal with the hierarchical nature of practical analog and RF systems. We begin with a detailed survey of algorithmic techniques for automatically extracting a suitable nonlinear macromodel from a device-level circuit. Such techniques are critical to both verification and synthesis activities for complex systems. We then survey recent ideas in hierarchical synthesis for analog systems and focus in particular on numerical techniques for handling the large number of degrees of freedom in these designs and for exploring the space of performance tradeoffs early in the design process. Finally, we briefly touch on recent ideas for accommodating models of statistical manufacturing variations in these tools and flows Rob A. Rutenbar, Georges Gielen, Jaijeet S. Roychowdhury |
Proc. IEEE | 1 |
| 2007 | Interval-Valued Reduced-Order Statistical Interconnect ModelingabstractWe show how advances in the handling of correlated interval representations of range uncertainty can be used to approximate the mass of a probability density function as it moves through numerical operations and, in particular, to predict the impact of statistical manufacturing variations on linear interconnect. We represent correlated statistical variations in resistance-inductance-capacitance parameters as sets of correlated intervals and show how classical model-order reduction methods - asymptotic waveform evaluation and passive reduced-order interconnect macromodeling algorithm - can be retargeted to compute interval-valued, rather than scalar-valued, reductions. By applying a simple statistical interpretation and sampling to the resulting compact interval-valued model, we can efficiently estimate the impact of variations on the original circuit. Results show that the technique can predict mean delay and standard deviation with errors between 5% and 10% for correlated parameter variations up to 35%. James D. Ma, Rob A. Rutenbar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | Probabilistic interval-valued computation: toward a practical surrogate for statistics inside CAD toolsabstractInterval methods offer a general, fine-grain strategy for modeling correlated range uncertainties in numerical algorithms. We present a new, improved interval algebra that extends the classical affine form to a more rigorous statistical foundation. Range uncertainties now take the form of confidence intervals. In place of pessimistic interval bounds, we minimize the probability of numerical "escape"; this can tighten interval bounds by 10X, while yielding 10-100X speedups over Monte Carlo. The formulation relies on three critical ideas: liberating the affine model from the assumption of symmetric intervals; a unifying optimization formulation; and a concrete probabilistic model. We refer to these as probabilistic intervals, for brevity. Our goal is to understand where we might use these as a surrogate for expensive, explicit statistical computations. Results from sparse matrices and graph delay algorithms demonstrate the utility of the approach, and the remaining challenges. Amith Singhee, Claire Fang Fang, James D. Ma, Rob A. Rutenbar |
DAC | 4 |
| 2006 | Generation of yield-aware Pareto surfaces for hierarchical circuit design space explorationabstractPareto surfaces in the performance space determine the range of feasible performance values for a circuit topology in a given technology. We present a non-dominated sorting based global optimization algorithm to generate the nominal pareto front efficiently using a simulator-in-a-loop approach. The solutions on this pareto front combined with efficient Monte Carlo approximation ideas are then used to compute the yield-aware pareto fronts. We show experimental results for both the nominal and yield-aware pareto fronts for power and phase noise for a voltage controlled oscillator (VCO) circuit. The presented methodology computes yield-aware pareto fronts in approximately 5-6 times the time required for a single circuit synthesis run and is thus practically efficient. We also show applications of yield-aware paretos to find the optimal VCO circuit to meet the system level specifications of a phase locked loop. Saurabh K. Tiwary, Pragati K. Tiwary, Rob A. Rutenbar |
DAC | 3 |
| 2006 | Verifying analog oscillator circuits using forward/backward abstraction refinementabstractProperties of analog circuits can be verified formally by partitioning the continuous state space and applying hybrid system verification techniques to the resulting abstraction. To verify properties of oscillator circuits, cyclic invariants need to be computed. Methods based on forward reachability have proven to be inefficient and in some cases inadequate in constructing these invariant sets. In this paper, we propose a novel approach combining forward- and backward-reachability while iteratively refining partitions at each step. The technique can yield dramatic memory and runtime reductions. We illustrate the effectiveness by verifying, for the first time, the limit cycle oscillation behavior of a third-order model of a differential VCO circuit Goran Frehse, Bruce H. Krogh, Rob A. Rutenbar |
DATE | 3 |
| 2006 | In silico vox: Towards speech recognition in silicon
Edward C. Lin, Rob A. Rutenbar, Tsuhan Chen |
Hot Chips Symposium | 3 |
| 2006 | Design automation for analog: the next generation of tool challengesabstractThe decade of the 1990s saw the first wave of practical "post-SPICE" tools for analog designs. A range of synthesis, optimization, layout and modeling techniques made their way from academic prototypes to first-generation commercial offerings. We offer some pragmatic prognostications for what the next wave might (or, more bluntly, should) focus on next, as pressure to improve AMS design productivity grows. Rob A. Rutenbar |
ICCAD | 1 |
| 2006 | Faster, parametric trajectory-based macromodels via localized linear reductionsabstractTrajectory-based methods offer an attractive methodology for automated, on-demand generation of macromodels for custom circuits. These models are generated by sampling the state trajectory of a circuit as it simulates in the time domain, and building macromodels by reducing and interpolating among the linearizations created at a suitably spaced subset of the time points visited during training simulations. However, a weak point in conventional trajectory models is the reliance on a single, global reduction matrix for the state space. We develop a new, faster method that generates and weaves together a larger set of smaller localized linearizations for the trajectory samples. The method not only improves speedups to 30X over SPICE, but as a side benefit also provides a platform for parametric small-signal simulation of circuits with variational device/process parameters, at a speedup of roughly 200X over SPICE. Saurabh K. Tiwary, Rob A. Rutenbar |
ICCAD | 2 |
| 2006 | Moving speech recognition from software to silicon: the in silico vox projectabstractTo achieve much faster decoding, or much lower power consumption, we need to liberate speech recognition from the artificial constraints of its current software-only form, and move the essential computations directly into silicon. There are vast efficiencies waiting to be unlocked in this application – we need the proper architecture to do so. We report results from a firstgeneration hardware architecture simulated at bit-level, and a complete, working FPGA-based prototype. Simulation results show that rather modest hardware designs, running 10-20X slower than conventional processors, can already decode at 0.6 xRT, running the standard 5K Wall Street Journal benchmark. Edward C. Lin, Rob A. Rutenbar, Tsuhan Chen |
INTERSPEECH | 3 |
| 2006 | Fast Interval-Valued Statistical Modeling of Interconnect and Effective CapacitanceabstractCorrelated interval representations of range uncertainty offer an attractive solution to approximating computations on statistical quantities. The key idea is to use finite intervals to approximate the essential mass of a probability density function (pdf) as it moves through numerical operators; the resulting compact interval-valued solution can be easily interpreted as a statistical distribution and efficiently sampled. This paper first describes improved interval-valued algorithms for asymptotic wave evaluation (AWE)/passive reduced-order interconnect macromodeling algorithm (PRIMA) model order reduction for tree-structured interconnect circuits with correlated resistance, inductance, and capacitance (RLC) parameter variations. By moving to a much faster interval-valued linear solver based on path-tracing ideas, and making more optimal tradeoffs between interval- and scalar-valued computations, the delay statistics roughly 10/spl times/ faster than classical Monte Carlo (MC) simulation, with accuracy to within 5% can be extracted. This improved interval analysis strategy is further applied in order to build statistical effective capacitance (C/sub eff/) models for variational interconnect, and show how to extract statistics of C/sub eff/ over 100/spl times/ faster than classical MC simulation, with errors less than 4%. James D. Ma, Rob A. Rutenbar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2005 | Scalable trajectory methods for on-demand analog macromodel extractionabstractTrajectory methods sample the state trajectory of a circuit as it simulates in the time domain, and build macromodels by reducing and interpolating among the linearizations created at a suitably spaced subset of the time points visited during training simulations. Unfortunately, moving from simple to industrial circuits requires more extensive training, which creates models too large to interpolate efficiently. To make trajectory methods practical, we describe a scalable interpolation architecture, and the first implementation of a complete trajectory "infrastructure" inside a full SPICE engine. The approach supports arbitrarily large training runs, automatically prunes redundant trajectory samples, supports limited hierarchy, enables incremental macromodel updates, and gives 3-10X speedups for larger circuits. Saurabh K. Tiwary, Rob A. Rutenbar |
DAC | 2 |
| 2005 | Timing-driven placement by grid-warpingabstractGrid-warping is a recent placement strategy based on a novel physical analogy: rather than move the gates to optimize their location, it elastically deforms a model of the 2-D chip surface on which the gates have been coarsely placed via a standard quadratic solve. In this paper, we introduce a timing-driven grid-warping formulation that incorporates slack-sensitivity-based net weighting. Given inevitable concerns about wirelength and runtime degradation in any timingdriven scheme, we also incorporate a more efficient net model and an integrated local improvement ("rewarping") step. An implementation of these ideas, WARP2, can improve worst-case negative slack by 37% on average, with very modest increases in wirelength and runtime. Zhong Xiu, Rob A. Rutenbar |
DAC | 2 |
| 2005 | Designer-Driven Topology Optimization for Pipelined Analog to Digital ConvertersabstractThe paper suggests a practical "hybrid" synthesis methodology which integrates designer-derived analytical models for system-level description with simulation-based models at the circuit level. We show how to optimize stage-resolution to minimize the power in a pipelined ADC. Exploration (via detailed synthesis) of several ADC configurations is used to show that a 4-3-2... resolution distribution uses the least power for a 13-bit 40 MSPS converter in a 0.25 /spl mu/m CMOS process. Yu-Tsun Chien, Jea-Hong Lou, Gin-Kou Ma, Rob A. Rutenbar, Tamal Mukherjee |
DATE | 5 |
| 2005 | Interval-valued statistical modeling of oxide chemical-mechanical polishingabstractTechnology-oriented tools provide the raw data needed to optimize the fabrication process itself, and to predict problematic variational impacts on silicon design. Unfortunately, even in these physics-oriented tools, statistically uncertain quantities appear as crucial inputs. To date, Monte Carlo techniques have been the dominant solution method. We suggest an alternative in which uncertainties are represented as correlated intervals, and interval-valued computations replace the standard scalar operations in the numerical algorithm for the tool. We use an oxide chemical-mechanical polishing tool as an example, and show how to "retrofit" workable statistical models on top of the original algorithm. Accuracies to within /spl sim/1-10% of Monte Carlo simulation, and speedups of /spl sim/10-100X can be achieved, depending on whether we choose a formulation which emphasizes accuracy, or efficiency. James D. Ma, Claire Fang Fang, Rob A. Rutenbar, Xiaolin Xie, Duane S. Boning |
ICCAD | 3 |
| 2005 | Fast interval-valued statistical interconnect modeling and reductionabstractCorrelated interval representations of range uncertainty offer an attractive solution for approximating computations on statistical quantities. The key idea is to use finite intervals to approximate the essential mass of a pdf as it moves through numerical operators; the resulting compact interval-valued solution can be easily interpreted as a statistical distribution and efficiently sampled. This paper describes improved interval-valued algorithms for AWE/PRIMA model order reduction for tree-structured interconnect with correlated $RLC$ parameter variations. By moving to a faster interval-valued linear solver based on path-tracing ideas, and making more optimal trade-offs between interval- and scalar-valued computations, we can extract delay statistics roughly 10X faster than a classical Monte Carlo simulation loop, with accuracy to within 5%. James D. Z. Ma, Rob A. Rutenbar |
ISPD | 2 |
| 2005 | Early research experience with OpenAccess gear: an open source development environment for physical designabstractPhysical design EDA research in academia has historically been based on infrastructure developed independently by individual contributors. This has led to fragmentation in the community, where interaction, data interchange and comparison of results between tools are difficult. We discuss our early experience with the OpenAccess Gear system, an open source software initiative intended to provide pieces of the critical integration and analysis infrastructure that are taken for granted in proprietary tools, but often wholly absent in research tools. Built on top of the widely available OpenAccess database, OA Gear provides components such as industrial-strength static timing analysis and extensible layout and netlist visualization. We discuss preliminary results from two on-going research efforts that have adopted OA Gear as their infrastructure: retrofitting the University of Michigan Capo placer into this environment, and the addition of a timing-driven capability to the Carnegie Mellon Warp placer. Zhong Xiu, David A. Papa, Philip Chong, Christoph Albrecht, Andreas Kuehlmann, Rob A. Rutenbar, Igor L. Markov |
ISPD | 6 |
| 2004 | Will Moore's Law rule in the land of analog?abstractOnce upon a time there was a wise and benevolent ruler whose Law multiplied his subjects' wealth and happiness--about 2X, every couple of years, but the kingdom was divided. Those in the happy hamlet of Digital got fatter (and faster), year after year. Not so the talented artisans in the town of Analog complained constantly about "voltage headroom", "variability", "noise", "matching", "kT/C limits", the rising costs of supporting their neighbors' insatiable addiction to shrinking transistors, and how the grass looked greener just over the border, in Silicon-Germania. So, what's a King to do? Will we see billion transistor chips with integrated RF made from transistors that are 25 atoms wide? Or will the peasants in the land of Analog really revolt. Rob A. Rutenbar, Anthony R. Bonaccio, Teresa H. Meng, Ernesto Perea, Robert Pitts, Charles G. Sodini, James B. Wieser |
DAC | 1 |
| 2004 | Large-scale placement by grid-warpingabstractGrid-warping is a new placement algorithm based on a strikingly simple idea: rather than move the gates to optimize their location, we elastically deform a model of the 2-D chip surface on which the gates have been roughly placed, "stretching" it until the gates arrange themselves to our liking. Put simply: we move the grid, not the gates. Deforming the elastic grid is a surprisingly simple, low-dimensional nonlinear optimization, and augments a traditional quadratic formulation. A preliminary implementation, WARP1, is already competitive with most recently published placers, e.g., placements that average 4% better wirelength, 40% faster than GORDIAN-L-DOMINO. Zhong Xiu, James D. Z. Ma, Suzanne M. Fowler, Rob A. Rutenbar |
DAC | 4 |
| 2004 | A synthesis flow toward fast parasitic closure for radio-frequency integrated circuitsabstractAn electrical and physical synthesis flow for high-speed analog and radio-frequency circuits is presented in this paper. Novel techniques aiming at fast parasitic closure are employed throughout the flow. Parasitic corners generated based on the earlier placement statistics are included for circuit resizing to enable parasitic robust designs. A performance-driven placement with simultaneous fast incremental global routing is proposed to achieve accurate parasitic estimation. Device tuning is utilized during layout to compensate for layout induced performance degradations. This methodology allows sophisticated macromodels of performances versus device variables and parasitics to be used during layout synthesis to make it truly performance-driven. Experimental results of a 4GHz LNA and a mixer demonstrate fast parasitic closure with this methodology. E. Aykut Dengi, Ronald A. Rohrer, Rob A. Rutenbar, L. Richard Carley |
DAC | 4 |
| 2004 | Towards formal verification of analog designsabstractWe show how model checking methods developed for hybrid dynamic systems may be usefully applied for analog circuit verification. Finite-state abstractions of the continuous analog behavior are automatically constructed using polyhedral outer approximations to the flows of the underlying continuous differential and difference equations. In contrast to previous approaches, we do not discretize the entire continuous state space, and our abstraction captures the relevant behaviors for verification in terms of the transitions between "states" (regions of the continuous state space) as a finite state machine in the hybrid system model. The approach is illustrated for two circuits, a standard oscillator benchmark, and a much larger and more realistic delta-sigma (AI) modulator. Smriti Gupta, Bruce H. Krogh, Rob A. Rutenbar |
ICCAD | 3 |
| 2004 | Interval-valued reduced order statistical interconnect modelingabstractWe show how recent advances in the handling of correlated interval representations of range uncertainty can be used to predict the impact of statistical manufacturing variations on linear interconnect. We represent correlated statistical variations in RLC parameters as sets of correlated intervals, and show how classical model order reduction methods - AWE and PRIMA - can be re-targeted to compute interval-valued, rather than scalar-valued reductions. By applying a statistical interpretation and sampling to the resulting compact interval-valued model, we can efficiently estimate the impact of variations on the original circuit. Results show the technique can predict mean delay with errors between 5-10%, for correlated RLC parameter variations up to 35%. James D. Ma, Rob A. Rutenbar |
ICCAD | 2 |
| 2004 | Static statistical timing analysis for latch-based pipeline designsabstractA latch-based timing analyzer is an essential tool for developing high-speed pipeline designs. As process variations increasingly influence the timing characteristics of DSM designs, a timing analyzer capable of handling process-induced timing variations for latch-based pipeline designs becomes in demand. In this work, we present a static statistical timing analyzer, STAP, for latch-based pipeline designs. Our analyzer propagates statistical worst-case delays as well as critical probabilities across the pipeline stages. We present an efficient method to handle correlations due to re-convergent fanouts. We also demonstrate the impact of not including the analysis of reconvergent fanouts in latch-based pipeline designs. Comparing to a Monte-Carlo based timing analyzer, our experiments show that STAP can accurately evaluate the critical probability that a design violates the timing constraints under a given statistical timing model. The runtime comparison further demonstrates the efficiency of our STAP. Rob A. Rutenbar, Li-C. Wang, Kwang-Ting Cheng, Sandip Kundu |
ICCAD | 1 |
| 2004 | A Comparative Study of Two Boolean Formulations of FPGA Detailed Routing ConstraintsabstractWe present empirical analyses of two Boolean satisfiability (SAT) formulations of FPGA (field programmable gate array) detailed routing constraints. Boolean SAT-based routing transforms a routing problem into a Boolean SAT instance by rendering geometric routing constraints as an atomic Boolean function. The generated Boolean function is satisfiable if and only if the corresponding routing is possible. Two different Boolean SAT-based routing models are analyzed: the track-based and the route-based routing constraint model. The track-based routing model transforms a routing task into a net-to-track assignment problem, whereas the route-based routing model reduces it into a routability-checking problem with explicitly enumerated set of detailed routes for nets. In both models, routing constraints are represented as CNF Boolean satisfiability clauses. Through comparative experiments, we demonstrate that the route-based formulation yields an easier-to-evaluate and more scalable routability Boolean function than the track-based method. This is empirical evidence that a smart/efficient Boolean formulation can achieve significant performance improvement in real-world applications. Gi-Joon Nam, Fadi A. Aloul, Karem A. Sakallah, Rob A. Rutenbar |
IEEE Trans. Computers | 4 |
| 2003 | Toward efficient static analysis of finite-precision effects in DSP applications via affine arithmetic modelingabstractWe introduce a static error analysis technique, based on smart interval methods from affine arithmetic, to help designers translate DSP codes from full-precision floating-point to smaller finite-precision formats. The technique gives results for numerical error estimation comparable to detailed simulation, but achieves speedups of three orders of magnitude by avoiding actual bit-level simulation. We show results for experiments mapping common DSP transform algorithms to implementations using small custom floating point formats. Claire Fang Fang, Rob A. Rutenbar, Markus Püschel, Tsuhan Chen |
DAC | 2 |
| 2003 | Mixed signals on mixed-signal: the right next technologyabstractCMOS dominates digital microelectronics. However, wireless applications require RF circuits at 1-5GHz, and exotic higher frequency applications are on the horizon. Silicon-Germanium (SiGe) is a growing choice for these designs. But is it "the" answer? Some argue that scaled CMOS will handle all tomorrow's RF ICs. Others argue that one-chip SoC solutions will never be the winning strategy for these highly heterogeneous designs, and place their bets on system-in-package (SiP) technologies. Is there a right answer here? Is CMOS the "only" way, or just "another" way? Rob A. Rutenbar, David L. Harame, Kurt Johnson, Paul Kempf, Teresa H. Meng, Reza Rofougaran, James Spoto |
DAC | 1 |
| 2003 | Floating-point error analysis based on affine arithmeticabstractDuring the development of floating-point signal processing systems, an efficient error analysis method is needed to guarantee the output quality. We present a novel approach to floating-point error bound analysis based on affine arithmetic. The proposed method not only provides a tighter bound than the conventional approach, but also is applicable to any arithmetic operation. The error estimation accuracy is evaluated across several different applications which cover linear operations, nonlinear operations, and feedback systems. The accuracy decreases with the depth of computation path and also is affected by the linearity of the floating-point operations. Claire Fang Fang, Tsuhan Chen, Rob A. Rutenbar |
ICASSP (2) | 3 |
| 2003 | Fast, Accurate Static Analysis for Fixed-Point Finite-Precision Effects in DSP Designs
Claire Fang Fang, Rob A. Rutenbar, Tsuhan Chen |
ICCAD | 2 |
| 2003 | sub-SAT: a formulation for relaxed Boolean satisfiability with applications in routingabstractAdvances in methods for solving Boolean satisfiability (SAT) for large problems have motivated recent attempts to recast physical design problems as Boolean SAT problems. One persistent criticism of these approaches is their inability to supply partial solutions, i.e., to satisfy most but not all of the constraints cast in the SAT style. In this paper, we present a formulation for "subset satisfiable" Boolean SAT: we transform a "strict" SAT problem with N constraints into a new, "relaxed" SAT problem which is satisfiable just if not more than k/spl Lt/N of these constraints cannot be satisfied in the original problem. We describe a transformation based on explicit thresholding and counting for the necessary SAT relaxation. Examples from field-programmable gate-array routing show how we can determine efficiently when we can satisfy "almost all" of our geometric constraints. Hui Xu 0001, Rob A. Rutenbar, Karem A. Sakallah |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2002 | Remembrance of circuits past: macromodeling by data mining in large analog design spacesabstractThe introduction of simulation-based analog synthesis tools creates a new challenge for analog modeling. These tools routinely visit 103 to 105 fully simulated circuit solution candidates. What might we do with all this circuit data? We show how to adapt recent ideas from large-scale data mining to build models that capture significant regions of this visited performance space, parameterized by variables manipulated by synthesis, trained by the data points visited during synthesis. Experimental results show that we can automatically build useful nonlinear regression models for large analog design spaces. Hongzhou Liu, Amith Singhee, Rob A. Rutenbar, L. Richard Carley |
DAC | 3 |
| 2002 | Hybrid Routing for FPGAs by Integrating Boolean Satisfiability with Geometric Search
Gi-Joon Nam, Karem A. Sakallah, Rob A. Rutenbar |
FPL | 3 |
| 2002 | Floating-point bit-width optimization for low-power signal processing applicationsabstractTo enable floating-point (FP) signal processing applications in low-power mobile devices, we propose a lightweight FP design flow that can optimize the bit-width configuration. The optimization considers both the hardware cost and the numerical precision. Variable grouping is used to reduce the complexity of optimization by connecting software description and hardware implementation. The optimization algorithm is able to avoid local optima, and multiple-phase optimization helps to reduce the cost further. We apply the proposed design flow to the design of inverse discrete cosine transform (IDCT), and show that the power consumption of our lightweight FP IDCT is comparable to an optimized fixed-point design. In addition, promising results on some real-world applications such as video coding and speech recognition demonstrate that lightweight FP signal processing will find more and more applications in low-power devices. Claire Fang Fang, Tsuhan Chen, Rob A. Rutenbar |
ICASSP | 3 |
| 2002 | sub-SAT: a formulation for relaxed boolean satisfiability with applications in routingabstractAdvances in methods for solving Boolean satisfiability (SAT) for large problems have motivated recent attempts to recast physical design problems as Boolean SAT problems. One persistent criticism of these approaches is their inability to supply partial solutions, i.e, to satisfy most but not all of the constraints cast in the SAT style. In this paper we present a formulation for "subset satisfiable" Boolean SAT: we transform a "strict" SAT problem with N constraints into a new, "relaxed" SAT problem which is satisfiable just if not more than k << N of these constraints cannot be satisfied in the original problem. We describe a transformation based on explicit thresholding and counting for the necessary SAT relaxation. Examples from FPGA routing show how we can determine efficiently when we can satisfy "almost all" of our geometric constraints. Hui Xu 0001, Rob A. Rutenbar, Karem A. Sakallah |
ISPD | 2 |
| 2002 | A new FPGA detailed routing approach via search-based BooleansatisfiabilityabstractBoolean-based routing methods transform the geometric FPGA routing task into a large but atomic Boolean function with the property that any assignment of input variables that satisfies the function specifies a valid routing solution. We present a new search-based satisfiability (SAT) FPGA detailed routing formulation that handles all channels in an FPGA simultaneously. The formulation has the virtue that it considers all nets concurrently allowing higher degrees of freedom for each net, in contrast to the classical one-net-at-a-time approaches and is able to prove the unroutability of a given circuit by demonstrating the absence of a satisfying assignment to the routing Boolean function. To demonstrate the effectiveness of this method, we first present comparative experimental results between integer linear programming (ILP)-based routing, which is an alternative concurrent method, and SAT-based routing. We also present the. first comparisons of search-based Boolean SAT routing results to other conventional routers and offer the first evidence that SAT methods can actually demonstrate the unroutability of a layout. Preliminary experimental results suggest that our approach compares very favorably with both the ILP-based approach and conventional FPGA routers. Gi-Joon Nam, Karem A. Sakallah, Rob A. Rutenbar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2001 | Panel: (When) Will FPGAs Kill ASICs?abstractThere was a time - in the dim historical past - when foundries actually made ASICs with only 5000 to 50,000 logic gates. But FPGAs and CPLDs conquered those markets and pushed ASIC silicon toward opportunities with more logic, volume, and speed. Today's largest FPGAs approach the few-million-gate size of a typical ASIC design, and continue to sprout embedded cores, such as CPUs, memories, and interfaces. And given the risks of nonworking nanometer silicon, FPGA costs and time-to-market are looking awfully attractive. So, will FPGAs kill ASICs? ASIC technologists certainly think not. ASICs are themselves sprouting patches of programmable FPGA fabric, and pushing new realms of size and especially speed. New tools claim to have tamed the convergence problems of older ASIC flows. Is the future to be found in a market full of FPGAs with ASIC-like cores? ASICs with FPGA cores? Other exotic hybrids? Our panelists will share their disagreements on these prognostications. Rob A. Rutenbar, Max Baron, Thomas Daniel, Rajeev Jayaraman, Zvi Or-Bach, Jonathan Rose, Carl Sechen |
DAC | 1 |
| 2001 | A boolean satisfiability-based incremental rerouting approach with application to FPGAsabstractIncremental redesign is an increasingly essential step in any complex design. Late changes or corrections in functional specifications (so-called "engineering change orders" or ECOs) force us to search for a minimal perturbation that achieves the desired repair. In reconfigurable design scenarios, these incremental repairs may be in response to physical faults: the goal is to "design around" the fault. For FPGAs, incremental rerouting is an essential component of this repair problem. We have developed a new incremental rerouting algorithm for FPGAs using techniques from Boolean Satisfiability (SAT). In this application, these techniques have the twin virtues that they (1) represent all possible routing (and rerouting) constraints simultaneously and exactly, and (2) search for rerouting solutions by perturbing all nets concurrently. Preliminary results are promising. For several FPGA benchmarks, we were able to reroute fault reconfigurations that perturb up to 5.74% of all nets for a small number of fault sets (one to four faults) with only 1.55 track overhead per channel on average, with CPU time 0.76 to 4.91 seconds/fault. Gi-Joon Nam, Karem A. Sakallah, Rob A. Rutenbar |
DATE | 3 |
| 2001 | Direct Transistor-Level Layout for Digital BlocksabstractWe present a complete transistor-level layout flow, from logic netlist to final shapes, for blocks of combinational logic up to a few thousand transistors in size. The direct transistor-level attack easily accommodates the demands for careful custom sizing necessary in high-speed design, and is also significantly denser than a comparable cell-based layout. The key algorithmic innovations are (a) early identification of essential diffusion-merged MOS device groups called clusters, but (b) deferred binding of clusters to a specific shape-level layout until the very end of a multi-phase placement strategy. A global placer arranges uncommitted clusters; a detailed placer optimizes clusters at shape level for density and for overall routability. A commercial router completes the flow. Experiments comparing to a commercial standard cell-level layout flow show that, when flattened to transistors, our tool consistently achieves 100% routed layouts that average 23% less area. Prakash Gopalakrishnan, Rob A. Rutenbar |
ICCAD | 2 |
| 2001 | ASF: A Practical Simulation-Based Methodology for the Synthesis of Custom Analog CircuitsabstractThis paper describes ASF, a novel cell-level analog synthesis framework that can size and bias a given circuit topology subject to a set of performance objectives and a manufacturing process. To manage complexity and time-to-market, SoC designs require a high level of automation and reuse. Digital methodologies are inapplicable to analog IP, which relies on tight control of low-level device and circuit properties that vary widely across manufacturing processes. This analog synthesis solution automates these tedious, technology specific aspects of analog design. Unlike previously proposed approaches, ASF extends the prevalent "schematic and SPICE" methodology used to design analog and mixed-signal circuits. ASF is topology and technology independent and can be easily integrated into a commercial schematic capture design environment. Furthermore, ASF employs a novel numerical optimization formulation that incorporates classical downhill techniques into stochastic search. ASF consistently produces results comparable to expert manual design with 10/spl times/ fewer candidate solution evaluations than previously published approaches that rely on traditional stochastic optimization methods. Michael Krasnicki, Rodney Phelps, James R. Hellums, Mark McClung, Rob A. Rutenbar, L. Richard Carley |
ICCAD | 5 |
| 2001 | Embedded Tutorial: CAD Solutions and Outstanding Challenges for Mixed-Signal and RF IC DesignabstractAddresses the problems and solutions that are posed by the design of mixed-signal integrated systems on chip (SoC). These include problems in mixed-signal design methodologies and flows, problems in analog design productivity, as well as open problems in analog, mixed-signal and RF design, modeling and verification tools. The tutorial explains the problems that are posed by these mixed-signal/RF SoC designs, describes the solutions and their underlying methods that exist today and outlines the challenges that still remain to be solved at present. In the first part the design of analog and mixed-signal circuits is addressed, while the second part focuses on the specific problems raised by RF wireless circuits. Domine Leenaerts, Rob A. Rutenbar, Georges Gielen |
ICCAD | 2 |
| 2001 | Automatic Hierarchical Design: Fantasy or Reality? (Panel)
Rob A. Rutenbar, Olivier Coudert, Patrick Groeneveld, Jürgen Koehl, Scott Peterson, Vivek Raghavan, Naresh Soni |
ICCAD | 1 |
| 2001 | Low-power technology mapping for mixed-swing logicabstractArticle Share on Low-power technology mapping for mixed-swing logic Authors: Rob A. Rutenbar Department of Electrical and Computer Engineering, Carnegie Mellon University, Pittsburgh, PA Department of Electrical and Computer Engineering, Carnegie Mellon University, Pittsburgh, PAView Profile , L. Richard Carley Department of Electrical and Computer Engineering, Carnegie Mellon University, Pittsburgh, PA Department of Electrical and Computer Engineering, Carnegie Mellon University, Pittsburgh, PAView Profile , Roberto Zafalon STMicroelectronics, Agrate (MI), Italy STMicroelectronics, Agrate (MI), ItalyView Profile , Nicola Dragone Department of Electrical and Computer Engineering, Carnegie Mellon University, Pittsburgh, PA and PDF Solutions, Desenzano (BS), Italy Department of Electrical and Computer Engineering, Carnegie Mellon University, Pittsburgh, PA and PDF Solutions, Desenzano (BS), ItalyView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 291–294https://doi.org/10.1145/383082.383171Online:06 August 2001Publication History 2citation209DownloadsMetricsTotal Citations2Total Downloads209Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Rob A. Rutenbar, L. Richard Carley, Roberto Zafalon, Nicola Dragone |
ISLPED | 1 |
| 2001 | A comparative study of two Boolean formulations of FPGA detailed routing constraintsabstractA Boolean-based router expresses the routing constraints as a Bool?ean function which is satisfiable if and only if the layout is routable. Compared to traditional routers, Boolean-based routers offer two unique features: (1) simultaneous embedding of all nets regardless of net ordering, and (2) ability to demonstrate routing infeasibility by proving the unsatisfiability of the generated routing constraint Boolean function. In this paper, we introduce a new Boolean-based FPGA detailed routing formulation that yields an easy-to-evaluate and more scalable routability Boolean function than the previous methods. The routability constraints are expressed in terms of a set of route variables each of which designating a specific detailed route for a given net. Experimental results clearly show the superi?ority of this formulation over an earlier formulation that expressed the constraints in terms of track variables. Gi-Joon Nam, Fadi A. Aloul, Karem A. Sakallah, Rob A. Rutenbar |
ISPD | 4 |
| 2001 | Wire packing - a strong formulation of crosstalk-aware chip-leveltrack/layer assignment with an efficient integer programming solutionabstractBy focusing on chip-wide slices of the global routing grid, making a few mild geometric assumptions about layer use, and suitably abstracting pin details, we derive an efficient integer linear programming formulation for track/layer assignment. The key technical insight is to model all constraints both geometric and crosstalk - as cliques in an appropriate conflict graph; these cliques can be extracted quickly from the interval structure of wires in a slice. We develop a "strong" linear relaxation of this problem that almost always yields the integral optimum; this solution gives us directly the maximum number of wires that can be packed legally without crosstalk risk. Experiments on synthetic netlists that match statistics of wire layouts from industrial 0.25-/spl mu/m designs demonstrate that we can pack 100-1000 wires optimally or, at worst, a very few overflows in seconds. Rony Kay, Rob A. Rutenbar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2000 | Case studies: Chip design on the bleeding edge (panel session abstract)abstractOften, the most interesting tools, methodologies, and insights come from designs that push hard on the “leading edge” of technology--what the survivors commonly call the “bleeding edge” of design. In this session, we collect three such on-the-edge designs, each done in a different style, each aimed at a very different market, each with its own unique set of challenges and solutions. All designers worry about complexity, about time-to-market, about correctness. But these designs go past the sorts of problems many of us have today, and offer some glimpses of problems and solutions we may all be facing tomorrow. John M. Cohn, Rob A. Rutenbar, Steve J. Young, Chris Malachowsky, Luis Aldaz |
DAC | 2 |
| 2000 | Survival strategies for mixed-signal systems-on-chip (panel session)abstractMore and more large ASICs require analog subsystems to interface to the real world--to wireless and wired networks, to sensors and transducers in embedded applications, to electrically complex high-speed interconnect. This is a major problem, since these analog subsystems break almost every assumption we know and love about digital systems. Analog circuits interact tightly with the technology, exploit rather than hide the physics of the fab, and manipulate precise electrical quantities rather than friendly binary abstractions. With respect to today's logic-centric CAD flows, analog blocks fit poorly and abstract badly. The digital side of SoC designs is addressed via a mix of cell-based logical and physical synthesis, commercially available soft and hard IP, and company-specific reuse methodologies for migrating complex functional blocks. On the analog side, “reuse” usually means hoping you still employ the person who designed the legacy analog block you are desperately trying to update. Stephan Ohr, Rob A. Rutenbar, Henry Chang, Georges Gielen, Rudolf Koch, Roy McGuffin, K. C. Murphy |
DAC | 2 |
| 2000 | A case study of synthesis for industrial-scale analog IP: redesign of the equalizer/filter frontend for an ADSL CODECabstractA persistent criticism of analog synthesis techniques is that they cannot cope with the complexity of realistic industrial designs, especially system-level designs. We show how recent advances in simulation-based synthesis can be augmented, via appropriate macromodeling, to attack complex analog blocks. To support this claim, we resynthesize from scratch, in several different styles, a complex equalizer/filter block from the frontend of a commercial ADSL CODEC, and verify by full simulation that it matches its original design specifications. As a result, we argue that synthesis has significant potential in both custom and analog IP reuse scenarios. Rodney Phelps, Michael Krasnicki, Rob A. Rutenbar, L. Richard Carley, James R. Hellums |
DAC | 3 |
| 2000 | Life at the end of CMOS scaling (and beyond) (panel session) (abstract only)abstractIt is clear by now that scaling for CMOS will ultimately hit a roadblock, and require a radical change in fundamental device technology. And yet--we continue to make progress in making impressively small MOSFET devices, at 70nm, 50nm, 20nm.Suppose we can actually get to a 70nm or 50nm or smaller technology, where one device is only a few hundred atoms across. What might life be like down here? How (and why) do the lab versions of these devices work today? And what obstacles exist to using them in real circuits, on real chips? Can we really make wires and pins and other interconnect? Can we manufacture them reliably?And what happens out beyond this inevitable end-of-scaling barrier? What options have we for new device paradigms.Our three invited speakers will fearlessly speculate on what might lie ahead on this wild frontier from three different technical perspectives: highly scaled devices themselves, issues with highly scaled circuits and interconnect, and devices that might actually work out beyond the scaling limit. Rob A. Rutenbar, Cheming Hu, Mark Horowitz, Stephen Y. Chow |
DAC | 1 |
| 2000 | Wire packing: a strong formulation of crosstalk-aware chip-level track/layer assignment with an efficient integer programming solutionabstractBy focusing on chip-wide slices of the global routing grid, making a few mild geometric assumptions about layer use, and suitably abstracting pin details, we derive an extremely effcient integer linear programming (ILP)formulationfor track/layer assignment.They key technical insight is to model all constraints-both geometric and crosstalk--as cliques in an appropriate conflict graph; these cliques can be extracted quickly from the interval structure of wires in a slice.We develop a "strong" linear relanation of this problem that almost always yields the integral optimum; this solution gives us directly the maximum number of wires that can pack legally without crosstalk risk.Experiments on synthetic netlists that match statistics of wire layouts from industrial 0.25um designs demonstrate that we can pack 100 -1000 wires optimally, or with at worst a very few over$ows, in seconch. Rony Kay, Rob A. Rutenbar |
ISPD | 2 |
| 2000 | Layout tools for analog ICs and mixed-signal SoCs: a surveyabstractArticle Free Access Share on Layout tools for analog ICs and mixed-signal SoCs: a survey Authors: Rob A. Rutenbar Dept. of ECE, Carnegie Mellon University, Pittsburgh, Pennsylvania Dept. of ECE, Carnegie Mellon University, Pittsburgh, PennsylvaniaView Profile , John M. Cohn IBM, Essex Junction, Vermont IBM, Essex Junction, VermontView Profile Authors Info & Claims ISPD '00: Proceedings of the 2000 international symposium on Physical designMay 2000 Pages 76–83https://doi.org/10.1145/332357.332378Online:01 May 2000Publication History 32citation1,208DownloadsMetricsTotal Citations32Total Downloads1,208Last 12 Months23Last 6 weeks4 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Rob A. Rutenbar, John M. Cohn |
ISPD | 1 |
| 2000 | Computer-aided design of analog and mixed-signal integrated circuitsabstractThis survey presents an overview of recent advances in the state of the art for computer-aided design (CAD) tools for analog and mixed-signal integrated circuits (ICs). Analog blocks typically constitute only a small fraction of the components on mixed-signal ICs and emerging systems-on-a-chip (SoC) designs. But due to the increasing levels of integration available in silicon technology and the growing requirement for digital systems to communicate with the continuous-valued external world, there is a growing need for CAD tools that increase the design productivity and improve the quality of analog integrated circuits. This paper describes the motivation and evolution of these tools and outlines progress on the various design problems involved: simulation and modeling, symbolic analysis, synthesis and optimization, layout generation, yield analysis and design centering, and test. This paper summarizes the problems for which viable solutions are emerging and those which are still unsolved. Georges Gielen, Rob A. Rutenbar |
Proc. IEEE | 2 |
| 2000 | Efficient handling of operating range and manufacturing linevariations in analog cell synthesisabstractWe describe a synthesis system that takes operating range constraints and inter and intracircuit parametric manufacturing variations into account while designing a sized and biased analog circuit. Previous approaches to computer-aided design for analog circuit synthesis have concentrated on nominal analog circuit design, and subsequent optimization of these circuits for statistical fluctuations and operating point ranges. Our approach simultaneously synthesizes and optimizes for operating and manufacturing variations by mapping the circuit design problem into an infinite programming problem and solving it using an annealing within annealing formulation. We present circuits designed by this integrated synthesis system, and show that they indeed meet their operating range and parametric manufacturing constraints. And finally, we show that our consideration of variations during the initial optimization-based circuit synthesis leads to better starting points for post-synthesis yield optimization than a classical nominal synthesis approach. Tamal Mukherjee, L. Richard Carley, Rob A. Rutenbar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2000 | Anaconda: simulation-based synthesis of analog circuits viastochastic pattern searchabstractAnalog synthesis tools have traditionally traded quality for speed, substituting simplified circuit evaluation methods for full simulation in order to accelerate the numerical search for solution candidates. As a result, these tools have failed to migrate into mainstream use primarily because of difficulties in reconciling the simplified models required for synthesis with the industrial-strength simulation environments required for validation. We argue that for synthesis to be practical, it is essential to synthesize a circuit using the same simulation environment created to validate the circuit. In this paper, we develop a new numerical search algorithm efficient enough to allow full circuit simulation of each circuit candidate, and robust enough to find good solutions for difficult circuits. The method combines the population-of-solutions ideas from evolutionary algorithms with a novel variant of pattern search, and supports transparent network parallelism. Comparison of several synthesized cell-level circuits against manual industrial designs demonstrates the utility of the approach. Rodney Phelps, Michael Krasnicki, Rob A. Rutenbar, L. Richard Carley, James R. Hellums |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2000 | Reducing power by optimizing the necessary precision/range of floating-point arithmeticabstractLow-power systems often find the power cost of floating-point (FP) hardware prohibitively expensive. This paper explores ways of reducing FP power consumption by minimizing the bitwidth representation of FP data. Analysis of several FP programs that manipulate low-resolution human sensory data shows that these programs suffer no loss of accuracy even with a significant reduction in bitwidth. Most FP programs in our benchmark suite maintain the same output even when the mantissa bitwidth is reduced by half. This FP bitwidth reduction can deliver a significant power saving through the use of a variable bitwidth FP unit. Our results show that up to 66% reduction in multiplier energy/operation can be achieved in the FP unit by this bitwidth reduction technique without sacrificing any program accuracy. J. Y. F. Tong, David Nagle, Rob A. Rutenbar |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1999 | MAELSTROM: Efficient Simulation-Based Synthesis for Custom Analog CellsabstractAnalog synthesis tools have failed to migrate into mainstream use primarily because of difficulties in reconciling the simplified models required for synthesis with the industrial-strength simulation environments required for validation. MAELSTROM is a new approach that synthesizes a circuit using the same simulation environment created to validate the circuit. We introduce a novel genetic/annealing optimizer, and leverage network parallelism to achieve efficient simulator-in-the-loop analog synthesis. Michael Krasnicki, Rodney Phelps, Rob A. Rutenbar, L. Richard Carley |
DAC | 3 |
| 1999 | Satisfiability-Based Layout Revisited: Detailed Routing of Complex FPGAs vis Search-Based Boolean SATabstractlier BDD-based methods.Boolean-based routing transforms the geometric FPGA routing task into a single, large Boolean equation with the property that any assignment of input variables that "satisfies" the equation (that renders equation identically "1") specifies a valid routing.The formulation has the virtue that it considers all nets simultaneously, and the absence of a satisfying assignment implies that the layout is unroutable.Initial Boolean-based approaches to routing used Binary Decision Diagrams (BDDs) to represent and solve the layout problem.BDDs, however, limit the size and complexity of the FPGAs that can be routed, leading these approaches to concentrate only on individual FPGA channels.In this paper, we present a new search-based Satisfiability (SAT) formulation that can handle entire FPGAs, routing all nets concurrently.The approach relies on a recently developed SAT engine (GRASP) that uses systematic search with conflict-directed non-chronological backtracking, capable of handling very large SAT instances.We present the first comparisons of search-based SAT routing results to other routers, and offer the first evidence that SAT methods can actually demonstrate the unroutability of a layout.Preliminary experimental results suggest that this approach to FPGA routing is more viable than ear-1.1 Gi-Joon Nam, Karem A. Sakallah, Rob A. Rutenbar |
FPGA | 3 |
| 1999 | Inverse polarity techniques for high-speed/low-power multipliersabstractVarious high-speed techniques have been developed for multipliers, but with the increasing popularity of mobile computing, a recent goal has been to minimize power dissipation. A popular delay-reduction technique applied to adder circuits is polarity inversion of bits. As this optimization reduces transistor count, it also has the potential for lowering power dissipation, and can be effectively applied to Wallace tree partial product reduction stages. We illustrate how this technique reduces power, interconnect capacitance, and chip area. Power reduction of up to 25 % is achieved. 1.1 Keywords Multiplier, low power, inverse polarity. Pascal C. H. Meier, Rob A. Rutenbar, L. Richard Carley |
ISLPED | 2 |
| 1999 | Device-level early floorplanning algorithms for RF circuitsabstractHigh-frequency circuits are notoriously difficult to lay out because of the tight coupling between device-level placement and wiring. Given that successful electrical performance requires careful control of the lowest-level geometric features-wire bends, precise length, planarity, etc., we suggest a new layout strategy for these circuits: early floorplanning at the device level. This paper develops a floorplanner for radio-frequency circuits based on a genetic algorithm (GA) that supports fully simultaneous placement and routing. The GA evolves slicing-style floorplans comprising devices and planned areas for wire meanders. Each floorplan candidate is fully routed with a gridless, detailed maze-router which can dynamically resize the floorplan as necessary. Experimental results demonstrate the ability of this approach to successfully optimize for wire planarity, realize multiple constraints on net lengths or phases, and achieve reasonable area in modest CPU times. Mehmet Aktuna, Rob A. Rutenbar, L. Richard Carley |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1998 | Device-level early floorplanning algorithms for RF circuitsabstractHigh-frequency circuits are notoriously difficult to lay out because of the tight coupling between device-level placement and wiring. Given that successful electrical performance requires careful control of the lowest-level geometric features—wire bends, precise length, proximity, planarity, etc.—we suggest a new layout strategy for these circuits: early floorplanning at the device level. This paper develops a floorplanner for RF circuits based on a genetic algorithm (GA) that supports fully simultaneous placement and routing. The GA evolves slicing-style floorplans comprising devices and planned areas for wire meanders. Each floorplan candidate is fully routed with a gridless, detailed maze-router which can dynamically resize the floorplan as necessary. Experimental results demonstrate the ability of this approach to successfully optimize for wire planarity, realize multiple constraints on net lengths or phases, and achieve reasonable area in modest CPU times. Mehmet Aktuna, Rob A. Rutenbar, L. Richard Carley |
ISPD | 2 |
| 1998 | Performance-driven simultaneous placement and routing for FPGA'sabstractSequential place and route tools for field programmable gate arrays (FPGA's) are inherently weak at addressing both wirability and timing optimizations. This is primarily due to the difficulty of accurately predicting wirability and delay during placement. A set of new performance-driven simultaneous placement/routing techniques has been developed for both row-based and island-style FPGA designs. These techniques rely on an iterative improvement placement algorithm augmented with fast, complete routing heuristics in the placement loop. For row-based designs, this new layout strategy yielded up to 28% improvements in timing and 33% in wirability for several MCNC benchmarks when compared to a traditional sequential place and route system in use at Texas Instruments. On a set of industrial designs for Xilinx 4000-series island-style FPGA's, our scheme produced 100% routed designs with 8-15% improvement in delay when compared to the Xilinx XACT5.0 place and route system. Sudip Nag, Rob A. Rutenbar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1998 | FPGA routing and routability estimation via Boolean satisfiabilityabstractGuaranteeing or even estimating the routability of a portion of a placed field programmable gate array (FPGA) remains difficult or impossible in most practical applications. In this paper, we develop a novel formulation of both routing and routability estimation that relies on a rendering of the routing constraints as a single large Boolean equation. Any satisfying assignment to this equation specifies a complete detailed routing. By representing the equation as a binary decision diagram (BDD), we represent all possible routes for all nets simultaneously. Routability estimation is transformed to Boolean satisfiability, which is trivial for BDD's. We use the technique in the context of a perfect routability estimator for a global router. Experimental results from a standard FPGA benchmark suite suggest the technique is feasible for realistic circuits, but refinements are needed for very large designs. R. Glenn Wood, Rob A. Rutenbar |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1997 | FPGA Routing and Routability Estimation via Boolean SatisfiabilityabstractGuaranteeing or even estimating the routability of a portion of a placed FPGA remains difficult or impossible in most practical applications. In this paper we develop a novel formulation of both routing and routability estimation that relies on a rendering of the routing constraints as a single large Boolean equation. Any satisfying assignment to this equation specifies a complete detailed routing. By representing the equation as a Binary Decision Diagram (BDD), we represent all possible routes for allnets simultaneously. Routability estimation is transformed to Boolean satisfiability, which is trivial for BDDs. We use the technique in the context of a perfect routability estimator for a global router. Experimental results from a standard FPGA benchmark suite suggest the technique is feasible for realistic circuits, but refinements are needed for very large designs. R. Glenn Wood, Rob A. Rutenbar |
FPGA | 2 |
| 1997 | A hierarchical decomposition methodology for multistage clock circuitsabstractThis paper describes a novel methodology to automate the design of the interconnect distribution for multistage clock circuits. We introduce two key ideas. First, a hierarchical decomposition of the layout divides the problem into a set of local Steiner-wired latch clusters (to minimize and balance local capacitance) fed globally by a balanced binary tree (to maximize performance). Second, we recast the global clock distribution problem as a simultaneous optimization of clock topology, clock segment routing, wire sizing and buffering. The hierarchical decomposition reduces the problem complexity and allows use of more aggressive optimization techniques. Integration of the geometric and electrical optimizations likewise allows more aggressive performance goals. Experiments with an industrial design comprising over 16,000 latches demonstrate the efficiency of the approach: a complete clock distribution solution met a 200-MHz cycle time specification with only 310 ps of skew, met strict current density constraints, exhibited good delay matching across uniform wire width and device variations, and was completed in under 10 CPU hours. Gary Ellis, Lawrence T. Pileggi, Rob A. Rutenbar |
ICCAD | 3 |
| 1996 | An O(n) Algorithm for Transistor Stacking with Performance ConstraintsabstractWe describe a new constraint-driven stacking algorithm for diffusion area minimization of CMOS circuits. It employs an Eulerian trail finding algorithm that can satisfy analog-specific performance constraints. Our technique is superior to other published approaches both ill terms of its time complexity and in the optimality of the stacks it produces. For a circuit with n transistors. The time complexity is O(n). All performance constraints are satisfied and, for a certain class of circuits, optimum stacking is guaranteed. Bülent Basaran, Rob A. Rutenbar |
DAC | 2 |
| 1996 | Synthesis Tools for Mixed-Signal ICs: Progress on Frontend and Backend StrategiesabstractArticle Free Access Share on Synthesis tools for mixed-signal ICs: progress on frontend and backend strategies Authors: L. Richard Carley Electrical and Computer Engineering, Carnegie Mellon University, Pittsburgh, PA Electrical and Computer Engineering, Carnegie Mellon University, Pittsburgh, PAView Profile , Georges G. E. Gielen Electrical Engineering, Katholieke Universiteit Leuven, Leuven, Belgium Electrical Engineering, Katholieke Universiteit Leuven, Leuven, BelgiumView Profile , Rob A. Rutenbar Electrical and Computer Engineering, Carnegie Mellon University, Pittsburgh, PA Electrical and Computer Engineering, Carnegie Mellon University, Pittsburgh, PAView Profile , Willy M. C. Sansen Electrical Engineering, Katholieke Universiteit Leuven, Leuven, Belgium Electrical Engineering, Katholieke Universiteit Leuven, Leuven, BelgiumView Profile Authors Info & Claims DAC '96: Proceedings of the 33rd annual Design Automation ConferenceJune 1996 Pages 298–303https://doi.org/10.1145/240518.240573Online:01 June 1996Publication History 36citation523DownloadsMetricsTotal Citations36Total Downloads523Last 12 Months21Last 6 weeks5 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF L. Richard Carley, Georges Gielen, Rob A. Rutenbar, Willy M. C. Sansen |
DAC | 3 |
| 1996 | Synthesis of high-performance analog circuits in ASTRX/OBLXabstractWe present a new synthesis strategy that can automate fully the path from an analog circuit topology and performance specifications to a sized circuit schematic. This strategy relies on asymptotic waveform evaluation to predict circuit performance and simulated annealing to solve a novel unconstrained optimization formulation of the circuit synthesis problem. We have implemented this strategy in a pair of tools called ASTRX and OBLX. To show the generality of our new approach, we have used this system to resynthesize essentially all the analog synthesis benchmarks published in the past decade; ASTRX/OBLX has resynthesized circuits in an afternoon that, for some prior approaches, had required months. To show the viability of the approach on difficult circuits, we have resynthesized a recently published (and patented), high-performance operational amplifier; ASTRX/OBLX achieved performance comparable to the expert manual design. And finally, to test the limits of the approach on industrial-sized problems, we have synthesized the component cells of a pipelined A/D converter; ASTRX/OBLX successfully generated cells 2-3/spl times/ more complex than those published previously. Emil S. Ochotta, Rob A. Rutenbar, L. Richard Carley |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1995 | Performance-driven simultaneous place and route for island-style FPGAsabstractSequential place and route tools for FPGAs are inherently weak at addressing both wirability and timing optimizations. This is primarily due to the difficulty of accurately predicting wirability and delay during placement. A new performance-driven simultaneous placement/routing technique has been developed for island-style FPGA designs. On a set of industrial designs for Xilinx 4000-series FPGAs, our scheme produces 100% routed designs with 8%-15% improvement in delay when compared to the Xilinx XACT5.0 place and route system. Sudip K. Nag, Rob A. Rutenbar |
ICCAD | 2 |
| 1995 | Reengineering the curriculum: design and analysis of a new undergraduate Electrical and Computer Engineering degree at Carnegie Mellon UniversityabstractIn the Fall of 1991, after approximately two years of development, the department of Electrical and Computer Engineering (ECE) at Carnegie Mellon University (CMU) implemented a new curriculum that differed radically from its predecessor. Key features of this curriculum include: Engineering in the Freshman year, a small core of required classes, area requirements in place of most specific course requirements, mandated breadth, depth, design, and coverage across ECE technical areas, a relatively large fraction of free electives, and a single integrated Bachelor of Science degree in Electrical and Computer Engineering. In this paper we review the design of this curriculum, including a taxonomy of problems we needed to address, and a set of general principles we evolved to address them. The new curriculum is described in detail, including new data from an ongoing analysis of its impact on students' curricula choices.> Stephen W. Director, Pradeep K. Khosla, Ronald A. Rohrer, Rob A. Rutenbar |
Proc. IEEE | 4 |
| 1995 | Integer programming based topology selection of cell-level analog circuitsabstractA new approach to cell-level analog circuit synthesis is presented. This approach formulates analog synthesis as a Mixed-Integer Nonlinear Programming (MINLP) problem in order to allow simultaneous topology and parameter selection. Topology choices are represented as binary integer variables and design parameters (e.g., device sizes and bias voltages) as continuous variables. Examples using a Branch and Bound method to efficiently solve the MINLP problem for CMOS two-stage op amps are given.> Prabir C. Maulik, L. Richard Carley, Rob A. Rutenbar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1994 | Performance-Driven Simultaneous Place and Route for Row-Based FPGAsabstractSequential place and route tools for FPGAs are inherently weak at addressing both wirability and timing optimizations. This is primarily due to the difficulty in predicting these at the placement level. A new performance-driven simultaneous placement / routing technique has been developed for row-based designs. Up to 28 % improvements in timing and 33 % in wirability have been achieved over a traditional sequential place and route system in use at Texas Instruments for several MCNC benchmark examples. 1 Sudip Nag, Rob A. Rutenbar |
DAC | 2 |
| 1994 | ASTRX/OBLX: Tools for Rapid Synthesis of High-Performance Analog CircuitsabstractWe describe ASTRX/OBLX, a synthesis system that can size high-performance analog circuit topologies to meet usersupplied linear performance specifications without designer-supplied equations.We present synthesis results for a large suite of circuit benchmarks and show that, when compared to prior approaches, ASTRX/OBLX can synthesize high-performance circuits with up to 3 orders of magnitude less initial design effort. Emil S. Ochotta, Rob A. Rutenbar, L. Richard Carley |
DAC | 2 |
| 1994 | Synthesis of manufacturable analog circuitsabstractWe describe a synthesis system that takes operating range constraints and inter- and intra-circuit parametric manufacturing variations into account while designing a sized and biased analog circuit. Previous approaches to CAD for analog circuit synthesis have concentrated on nominal analog circuit design, and subsequent optimization of these circuits for statistical fluctuations and operating point ranges. Our approach simultaneously synthesizes and optimizes for operating and manufacturing variations by mapping the circuit design problem into an Infinite Programming problem and solving it using an annealing within annealing formulation. We present circuits designed by this integrated synthesis system, and show that they indeed meet their operating range and parametric manufacturing constraints. Tamal Mukherjee, L. Richard Carley, Rob A. Rutenbar |
ICCAD | 3 |
| 1993 | Latchup-aware placement and parasitic-bounded routing of custom analog cellsabstractThis paper presents new results in constraint-directed placement and routing of device-level analog cells. We describe the first algorithm for latchup-aware device placement that guarantees sufficient placement of well/substrate contacts by simultaneous placement of latchup protection geometry and devices. A novel cost-to-target predictor for cost-based A/sup */ routing, and a new scheme for pruning evolving paths that violate user-supplied constrained routing of analog signals in dense placements are described. Experimental results suggest the strategy avoids the artifacts of length-for-crosstalk trade-offs seen in earlier algorithms, and allows users more fine-grain control of the routing of defense, high-performance cells. An implementation of these ideas in the tool set KA III is used to derive improved layouts from several analog cells. Bülent Basaran, Rob A. Rutenbar, L. Richard Carley |
ICCAD | 2 |
| 1992 | A Mixed-Integer Nonlinear Programming Approach to Analog Circuit Synthesis
Prabir C. Maulik, L. Richard Carley, Rob A. Rutenbar |
DAC | 3 |
| 1992 | System-level routing of mixed-signal ASICs in WRENabstractTechniques for global and detailed routing of the macrocell-style analog core of a mixed-signal ASIC are discussed. A comparatively simple geometric model of the problem is combined with an aggressive simulated annealing formulation that selects paths while accommodating numerous signal-integrity constraints. Experimental results demonstrate that it is critical to attack such constraints both globally (system-level) and locally (channel-level) to meet designer-specified performance targets.> Sujoy Mitra, Sudip Nag, Rob A. Rutenbar, L. Richard Carley |
ICCAD | 3 |
| 1992 | Knowledge Representation and Reasoning in a Software Synthesis ArchitectureabstractThe knowledge representation and reasoning strategies in an automatic program synthesis architecture called ELF are described. ELF synthesizes computer-aided design (CAD) tools that automatically route wires in VLSI circuits. The design space ELF confronts, requires it to understand various physical technologies, to select an appropriate procedure-level decomposition, to choose algorithms and data structures, to manage any interdependencies, and to generate efficient code. ELF manages the design space using a variety of knowledge sources, including domain-specific knowledge. The manner in which knowledge is used determines the representation method of choice. The effectiveness of these ideas is illustrated via a tour through the synthesis steps for a specific routing tool, and a brief discussion of the performance of the resulting synthetic router as measured against an industrial tool.> Dorothy E. Setliff, Rob A. Rutenbar |
IEEE Trans. Software Eng. | 2 |
| 1991 | Techniques for Simultaneous Placement and Routing of Custom Analog Cells in KOAN/ANAGRAM IIabstractThe authors describe novel techniques for simultaneous device placement and detailed routing of analog cells. Both nets and devices are treated as placeable, malleable objects in a common simulated-annealing framework. A detailed routing abstraction called a k-bend net limits the complexity of each net's topology, and allows simple incremental net reshaping during annealing. Analog layouts in which the critical interacting nets are simultaneously embedded during device placement prove to be superior, in terms of performance-limiting crosstalk violations, to sequentially placed and routed layouts.> John M. Cohn, David J. Garrod, Rob A. Rutenbar, L. Richard Carley |
ICCAD | 3 |
| 1991 | A Parallel Steiner Heuristic for Wirelength Estimation of Large Net PopulationsabstractThe authors discuss techniques to produce Steiner trees for a large population of nets, e.g., 1000 to 5000 nets, in parallel. This problem arises in iterative-improvement layout strategies that perturb not just a single placed object, but a few thousand objects simultaneously. Such strategies are the focus of work on mapping large placement problems onto massively parallel computers. The authors present a new heuristic that computes Steiner trees for an arbitrary number of nets, each with an arbitrary number of terminals, in essentially constant time given sufficient data-parallel machine resources and a constant distribution of net sizes. Experiments on a Connection Machine demonstrate that it is possible to create good Steiner trees for a few thousand nets in a few hundred milliseconds.> Rajeev Jayaraman, Rob A. Rutenbar |
ICCAD | 2 |
| 1991 | Massively parallel switch-level simulation: a feasibility studyabstractThe feasibility of mapping the COSMOS switch-level simulator onto a computer with thousands of simple processors is addressed. COSMOS preprocesses transistor networks into Boolean behavioral models, capturing the switch-level behavior of a circuit in a set of Boolean formulas. A class of massively parallel computers and a mapping of COSMOS onto these computers are described. The factors affecting the performance of such a massively parallel simulator are discussed, including: the amount of parallelism in the simulation model, performance measures for massively parallel machines, and the impact of event scheduling on simulator performance. Compilation tools that automatically map a MOS circuit onto a massively parallel computer have been developed. Techniques for restructuring Boolean expressions for greater parallelism and mapping Boolean expressions for evaluation on massively parallel machines are described. Massively parallel switch-level simulation is illustrated by a pilot implementation on a 32k-processor Thinking Machines Connection Machine system.> Saul A. Kravitz, Randal E. Bryant, Rob A. Rutenbar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1991 | On the feasibility of synthesizing CAD software from specifications: generating maze router tools in ELFabstractThe application of program synthesis techniques to the generation of technology-sensitive VLSI physical design tools is described. The architecture and implementation of a particular software generator (called ELF) targeted at the generation of maze routing software is described. ELF strives to meet the demands of the target technology by automatically generating maze router implementations to match the application requirements. ELF has three key features. First, a very high level language, lacking data structure implementation specifications, is used to describe algorithm design styles. Second, application-specific expertise about routing and application independent code synthesis techniques are used to guide search among alternative design styles for algorithms and data structures. Third, code generation is used to transform the resulting abstract descriptions of selected algorithms and data structures into final, executable code. Code generation is an incremental, stepwise refinement process. Experimental results are presented covering several correct. fully functional routers synthesized by ELF from varying high-level specifications. Results from synthetic and industrial benchmarks are examined to illustrate ELF's capabilities.> Dorothy E. Setliff, Rob A. Rutenbar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1990 | Design and Performance Evaluation of New Massively Parallel VLSI Mask Verification Algorithms in JIGSAWabstractThis paper describes JIGSAW, the massively parallel mask checking system that has evolved from our earlier feasibility study on large-scale, fine-grain parallelism in simple mask checking tasks [1]. Unlike previous systems, JIGSAW parallelizes all phases of the checking process. We describe new techniques to handle all-angle geometry, the first massively parallel mask flattening and multi-layer netlist extraction algorithms, and measurements made comparing JIGSAW, running on a Connection Machine, against industry-standard tools. End-to-end speedups, (i.e., from CIF to errors) range from 19 to 58 over DRACULA, with larger masks producing larger speedups. Erik C. Carlson, Rob A. Rutenbar |
DAC | 2 |
| 1989 | Massively Parallel Switch-Level Simulation: A Feasibility StudyabstractThis work addresses the feasibility of mapping the COSMOS switch-level simulator onto a computer with thousands of simple processors. COSMOS preprocesses transistor networks into Boolean behavioral models, capturing the switch-level behavior of a circuit in a set of Boolean formulas. We describe a class of massively parallel computers and a mapping of COSMOS onto these computers. We discuss the factors affecting the performance of such a massively parallel simulator including: the amount of parallelism in the simulation model, performance measures for massively parallel machines, and the impact of event scheduling on simulator performance. We have developed compilation tools which automatically map a MOS circuit onto a massively parallel computer. Massively parallel switch-level simulation is illustrated by describing our pilot implementation on a 32k processor Thinking Machines Connection Machine System. Saul A. Kravitz, Randal E. Bryant, Rob A. Rutenbar |
DAC | 3 |
| 1989 | ELF: A Tool for Automatic Synthesis of Custom Physical CAD SoftwareabstractThis paper describes how program synthesis techniques can be applied to the generation of technology-sensitive VLSI design tools. We present results from ELF, a synthesis tool for wire-routing software. The ELF synthesis architecture has three key features. First, a very high level language, lacking data structure implementation specifications is used to describe algorithm design styles. Second, routing domain knowledge and generic program synthesis knowledge are used to guide search among candidate design styles for all necessary component algorithms, and to deduce compatible data structure implementations for these components. Third, code generation is used to transform the resulting abstract descriptions of selected algorithms and data structures into final, executable code. Code generation is an incremental, stepwise refinement process. We present experimental results from several correct, fully-functional routers synthesized by ELF from varying high-level specifications. Dorothy E. Setliff, Rob A. Rutenbar |
DAC | 2 |
| 1989 | Logic Simulation on Massively Parallel ArchitecturesabstractThis work examines the mapping of logic simulation onto massively parallel computer architectures. We discuss alternative communication primitives for a massively parallel instruction set architecture and the impact of the choice of communication primitives on logic simulation. We have developed compilation tools to automatically map the simulation of an MOS transistor circuit onto a massively parallel computer. We analyze the efficiency of this mapping as a function of the available communication primitives. The compilation process is illustrated by describing our pilot implementation on a 32k processor Connection Machine. Saul A. Kravitz, Randal E. Bryant, Rob A. Rutenbar |
ISCA | 3 |
| 1989 | OASYS: a framework for analog circuit synthesisabstractA hierarchically structured framework for analog circuit synthesis is described. This hierarchical structure has two important features: it decomposes the design task into a sequence of smaller tasks with uniform structure, and it simplifies the reuse of design knowledge. Mechanisms are described that select from among alternate design styles and translate performance specifications from one level in the hierarchy to the next lower, more concrete level. A prototype implementation, OASYS, synthesizes sized transistor schematics for CMOS operational amplifiers from performance specifications and process parameters. Measurements from detailed circuit simulation and from actual fabricated analog ICs based on OASYS-synthesized designs demonstrate that OASYS is capable of synthesizing functional circuits.> Ramesh Harjani, Rob A. Rutenbar, L. Richard Carley |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1988 | Mask Verification on the Connection Machine
Erik C. Carlson, Rob A. Rutenbar |
DAC | 2 |
| 1988 | Automatic layout of custom analog cells in ANAGRAMabstractANAGRAM models cell layout in the style of a macrocell place-and-route problem. Individual cell primitives (transistor-level objects of widely varying sizes) are the macrocells. Module generation techniques are used to generate these internal primitives and to preserve critical matching and symmetries. An annealing-based placement algorithm then places these primitives. This is followed by a novel line-expansion signal router, which includes mechanisms to avoid noise coupling due to internodal capacitances between the signal wires and shared parasitic resistances in the DC supply wiring and operates in an iterative improvement fashion to eliminate such violations. Layouts for several custom CMOS cells have been successfully generated. Circuit-simulation results based on cell extractions demonstrate the effectiveness of the crosstalk-avoidance mechanisms.> David J. Garrod, Rob A. Rutenbar, L. Richard Carley |
ICCAD | 2 |
| 1988 | Analog circuit synthesis for performance in OASYSabstractMechanisms needed to meet stringent performance demands in a hierarchically structured analog circuit synthesis tool are described. Experiences with adding a high-speed comparator design style to the OASYS synthesis tool are discussed. It is argued that design iteration (the process of making a heuristic design choice, following it through to possible failure, then diagnosing the failure and modifying the overall plan of attack for the synthesis) is essential to meet such performance demands. Examples of high-speed comparators automatically synthesized by OASYS are presented. Designs competitive in quality with manual expert designs, e.g. with response time of 6 ns and input drive of 1 mV, can be synthesized in under 5 seconds on a workstation.> Ramesh Harjani, Rob A. Rutenbar, L. Richard Carley |
ICCAD | 2 |
| 1988 | Analog circuit synthesis and exploration in OASYSabstractExperimental results obtained with OASYS, a behavior-to-structure synthesis tool for analog circuits, are described. In particular, measurements from fabricated analog ICs based on OASYS-synthesized designs are presented, and used to verify that OASYS is capable of producing real, functional circuits. Possibilities for automatically exploring the space of designable analog circuits, an ability made possible by a fast, automatic synthesis tool such as OASYS, are also described. Examples of using OASYS to explore tradeoffs among process and performance specifications are presented.> Ramesh Harjani, Rob A. Rutenbar, L. Richard Carley |
ICCD | 2 |
| 1988 | Systolic routing hardware: performance evaluation and optimizationabstractThe performance of maze-routing algorithms mapped onto linear systolic array hardware is examined. Cell expansions in the wavefront-expansion phase of maze routing are performed in parallel in each processing stage of the hardware as the routing grid streams through the processor array. The authors concentrate on optimizing the performance of single-net routing problems with respect to a given systolic hardware configuration. A heuristic called constant-increment framing is introduced as a simple method for scheduling all the required wavefront expansion steps on a pipeline of processors. One-layer and two-layer routers using this heuristic have been implemented on a prototype systolic processor. Experimental and theoretical comparisons suggest that the constant-increment heuristic exhibits performance within a factor of two of optimal over a range of hardware configurations, and is substantially easier to compute than the optimal solution.> Rob A. Rutenbar, Daniel E. Atkins |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1987 | A Prototype Framework for Knowledge-Based Analog Circuit SynthesisabstractAn organization for a knowledge-based analog circuit synthesis tool is described. Analog circuit topologies are represented as a hierarchy of functional blocks; a planning mechanism is introduced to translate performance specifications between levels in this circuit hierarchy. A prototype implementation, OASYS, synthesizes sized transistor schematics for simple CMOS operational amplifiers from performance specifications and process parameters, and demonstrates the workability of the approach. Ramesh Harjani, Rob A. Rutenbar, L. Richard Carley |
DAC | 2 |
| 1987 | A Scanline Data Structure Processor for VLSI Geometry CheckingabstractThis paper proposes an architecture to support VLSI geometry checking tasks based on scanline algorithms. Rather than recast the entire verification task in hardware, we identify primitives around which geometry checking tools can be built, and examine the feasibility of accelerating two of these critical primitives. We focus on the operations of Boolean combinations of mask layers, and region numbering within a mask layer. Unlike previous proposals for special hardware (e.g., bit map processors), this architecture operates on a more realistic representation of masks: a sorted stream of possibly oblique edges. The architecture can be viewed as directly interpreting the operators that manipulate the relevant scanline data structures. We show how the edge computations in these two algorithms can be restructured into the form of a single, shared hardware pipeline. Data from a simulation of this processor suggests that, relative to the specific software functions it is intended to replace, the scanline processor can reduce computation time significantly. In particular, simulations of one possible implementation for this processor yield speedups of three orders of magnitude for Manhattan mask data, degrading gracefully to speedups of two orders of magnitude for highly oblique mask data. Erik C. Carlson, Rob A. Rutenbar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1987 | Placement by Simulated Annealing on a MultiprocessorabstractPhysical design tools based on simulated annealing algorithms have been shown to produce results of extremely high quality, but typically at a very high cost in execution time. This paper selects a representative annealing application--standard cell placement--and develops multiprocessor-based annealing algorithms for placement. A taxonomy of possible multiprocessor decompositions of annealing algorithms is presented which divides decomposition schemes into two broad classes: those which divide individual moves into subtasks and distribute them across cooperating processors, and those which perform complete moves in parallel. It is shown that the choice of multiprocessor annealing strategy is influenced by temperature; in particular, the paper introduces the idea of adaptive strategies that dynamically change the parallel decomposition scheme to achieve maximum speedup as the annealing task progresses through each temperature regime. Implementations of three parallel placement strategies are described for an experimental shared-memory multiprocessor. Practical speedups are achieved over a serial version of the algorithm, and it is shown that an adaptive strategy which switches between two parallel decompositions at the optimal temperature yields speedup significantly better than any single strategy approach. Models are developed to account for the observed performance, and to predict the crossover points for switching strategies. Saul A. Kravitz, Rob A. Rutenbar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1986 | Multiprocessor-based placement by simulated annealing
Saul A. Kravitz, Rob A. Rutenbar |
DAC | 2 |
| 1985 | Future directions for DA machine research (panel session)abstractNo abstract available. Rob A. Rutenbar |
DAC | 1 |
| 1984 | A Class of Cellular Architectures to Support Physical Design AutomationabstractSpecial-purpose hardware has been proposed as a solution to several increasingly complex problems in design automation. This paper examines a class of cellular architectures called raster pipeline subarrays--RPS architectures--applicable to problems in physical DA that are (1) representable on a cellular grid, and (2) characterized by local functional dependencies among grid cells. Machines with this architecture first evolved in conventional cellular applications that exhibit similarities to grid-based DA problems. To analyze the properties of the RPS organization in context, machines designed for cellular applications are reviewed, and it is shown that many DA machines proposed/constructed for grid-based problems fit naturally into a taxonomy of cellular machines. The implementation of DA algorithms on RPS hardware is partitioned into local issues that involve the processing of individual cell neighborhoods, and global issues that involve strategies for handling complete grids in a pipeline environment. Design rule checking and routing algorithms are examined in an RPS environment with respect to these issues. Experimental measurements for such algorithms running on an existing RPS machine exhibit significant speedups. From these studies are derived the necessary performance characteristics of RPS hardware optimized specifically for grid-based DA. Finally, the practical merits of such an architecture are evaluated. Rob A. Rutenbar, Trevor N. Mudge, Daniel E. Atkins |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1982 | Cellular image processing techniques for VLSI circuit layout validation and routingabstractThe architecture of the Cytocomputer?, an existing special-purpose, pipelined cellular image processor, is described. A formalism used to express cellular operations on images is then given. Cellular image processing algorithms are then developed that perform (1) design rule checks (DRC's) on VLSI circuit layouts, and (2) Lee-type wire routing. Two sets of cellular image processing transformations for checking the Mead and Conway design rules and for Lee-routing have been defined and used to program the Cytocomputer. Some experimental results are shown for these cellular implementations. Trevor N. Mudge, Rob A. Rutenbar, Robert M. Lougheed, Daniel E. Atkins |
DAC | 2 |
| 1981 | Case study of a VLSI design project: A simple inner product machineabstractWe present a case study of the application of recently evolved structured VLSI design methodologies to the design and implementation of a simple VLSI quasi-serial inner product machine. Rob A. Rutenbar, Y. E. Park |
IEEE Symposium on Computer Arithmetic | 1 |