VLDB 2026 Research / reviewers in the wild / expert
Richard B. Brown
dblp:61/2613
· DBLP profile ↗
45ranked-venue papers
1as first author
0since 2021 · last 2016
0000-0003-2539-2728ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 44 · 1 first-authorSoftware engineering, systems software and programming languages · 6Applied, interdisciplinary, general and emerging computing · 6
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
13 papers |
Electronic design automation · 43% Processor architecture and microarchitecture · 20% Memory systems · 14% | |
| Software engineering, system software, and programming languages
4 papers |
Compilers and program optimization · 89% Operating systems · 11% |
Topics — the 30 heaviest of 41, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
physical design |
0.1 | 2 | 2006 | Clock buffer and wire sizing using sequential programming · DAC 2006 Congestion Driven Quadratic Placement · DAC 1998 |
Electronic design automation › physical design › interconnect optimization
buffer insertion and wire sizing |
0.1 | 1 | 2006 | Clock buffer and wire sizing using sequential programming · DAC 2006 |
Electronic design automation › physical design › clock network synthesis
clock network optimization |
0.1 | 1 | 2006 | Clock buffer and wire sizing using sequential programming · DAC 2006 |
Electronic design automation › physical design › clock network synthesis
clock skew optimization |
0.1 | 1 | 2006 | Clock buffer and wire sizing using sequential programming · DAC 2006 |
Compilers and program optimization
register allocation |
0.1 | 1 | 2005 | Partitioning Variables across Register Windows to Reduce Spill Code in a Low-Power Processor · IEEE Trans. Computers 2005 |
Compilers and program optimization › register allocation
spill code minimization |
0.1 | 1 | 2005 | Partitioning Variables across Register Windows to Reduce Spill Code in a Low-Power Processor · IEEE Trans. Computers 2005 |
Processor architecture and microarchitecture
register file |
0.1 | 1 | 2005 | Partitioning Variables across Register Windows to Reduce Spill Code in a Low-Power Processor · IEEE Trans. Computers 2005 |
Processor architecture and microarchitecture › register file
register window |
0.1 | 1 | 2005 | Partitioning Variables across Register Windows to Reduce Spill Code in a Low-Power Processor · IEEE Trans. Computers 2005 |
Integrated circuit design › analog and mixed-signal circuits
mixed-signal circuit design |
0.0 | 1 | 2003 | A 16-bit mixed-signal microsystem with integrated CMOS-MEMS clock reference · DAC 2003 |
Performance modeling and evaluation › simulation › discrete-event simulation
trace-driven simulation |
0.0 | 2 | 1997 | Multilevel Optimization of Pipelined Caches · IEEE Trans. Computers 1997 Resource Allocation in a High Clock Rate Microprocessor · ASPLOS 1994 |
Integrated circuit design
digital circuit design |
0.0 | 1 | 2000 | CGaAs PowerPC FXU · DAC 2000 |
Integrated circuit design
low-power circuit design |
0.0 | 1 | 2000 | CGaAs PowerPC FXU · DAC 2000 |
Memory systems › memory management › virtual memory › address translation
TLB |
0.0 | 2 | 1994 | Design Tradeoffs for Software-Managed TLBs · ACM Trans. Comput. Syst. 1994 Design Tradeoffs for Software-Managed TLBs · ISCA 1993 |
Electronic design automation › physical design › routing
congestion prediction |
0.0 | 1 | 1998 | Congestion Driven Quadratic Placement · DAC 1998 |
Electronic design automation › physical design
placement |
0.0 | 1 | 1998 | Congestion Driven Quadratic Placement · DAC 1998 |
Electronic design automation › physical design › placement
routability-driven placement |
0.0 | 1 | 1998 | Congestion Driven Quadratic Placement · DAC 1998 |
Electronic design automation › physical design
routing |
0.0 | 1 | 1998 | Congestion Driven Quadratic Placement · DAC 1998 |
Processor architecture and microarchitecture › pipelining
pipeline hazard |
0.0 | 2 | 1993 | A microarchitectural performance evaluation of a 3.2 Gbyte/s microprocessor bus · MICRO 1993 Performance Optimization of Pipelined Primary Caches · ISCA 1992 |
Performance modeling and evaluation › simulation
architectural simulation |
0.0 | 1 | 1997 | Multilevel Optimization of Pipelined Caches · IEEE Trans. Computers 1997 |
Memory systems › cache management › cache capacity management
cache capacity optimization |
0.0 | 1 | 1997 | Multilevel Optimization of Pipelined Caches · IEEE Trans. Computers 1997 |
Memory systems
cache design |
0.0 | 1 | 1997 | Multilevel Optimization of Pipelined Caches · IEEE Trans. Computers 1997 |
Memory systems › memory hierarchy
cache hierarchy |
0.0 | 1 | 1997 | Multilevel Optimization of Pipelined Caches · IEEE Trans. Computers 1997 |
Memory systems
cache |
0.0 | 2 | 1992 | Performance Optimization of Pipelined Primary Caches · ISCA 1992 Implementing a Cache for a High-Performance GaAs Microprocessor · ISCA 1991 |
Embedded and real-time systems › embedded processor
low-power embedded processor |
0.0 | 1 | 2005 | Partitioning Variables across Register Windows to Reduce Spill Code in a Low-Power Processor · IEEE Trans. Computers 2005 |
Operating systems › resource management › memory management
virtual memory |
0.0 | 2 | 1994 | Design Tradeoffs for Software-Managed TLBs · ISCA 1993 Design Tradeoffs for Software-Managed TLBs · ACM Trans. Comput. Syst. 1994 |
Electronic design automation
high-level synthesis |
0.0 | 1 | 1995 | The Aurora RAM Compiler · DAC 1995 |
Cloud and datacenter computing
resource allocation |
0.0 | 1 | 1994 | Resource Allocation in a High Clock Rate Microprocessor · ASPLOS 1994 |
Memory systems › virtual memory management
software TLB |
0.0 | 1 | 1994 | Design Tradeoffs for Software-Managed TLBs · ACM Trans. Comput. Syst. 1994 |
Memory systems › memory management
virtual memory |
0.0 | 1 | 1994 | Design Tradeoffs for Software-Managed TLBs · ACM Trans. Comput. Syst. 1994 |
Processor architecture and microarchitecture › pipelining
instruction pipeline |
0.0 | 1 | 1993 | A microarchitectural performance evaluation of a 3.2 Gbyte/s microprocessor bus · MICRO 1993 |
Methods — techniques the papers use, named apart from their topics
graph partitioning · 0.1sequential quadratic programming · 0.1monte carlo simulation · 0.1trace-driven simulation · 0.0simulation · 0.0hardware monitoring · 0.0quadratic placement · 0.0line-probe heuristics · 0.0a* router · 0.0timing analysis · 0.0microarchitectural simulation · 0.0dependency classification · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | An empirical model of UWB large-scale signal fading in neocortical researchabstractIn this paper, we present an investigation of the suitability of the impulse ultra-wideband (I-UWB) technology for neural recording systems. A fully-digital I-UWB transmitter, implemented in a 65 nm CMOS, is connected to a miniature antenna and packaged in bio-compatible material. Two experimental telemetry configurations are considered. The head of a domestic pig is used to mimic a neural activity monitoring system that is conventionally implanted in swine. Three main objectives are accomplished (1) the strength of the UWB signal after propagating from the cranial cavity through the pig's skull and skin is measured for several separation distances; (2) a statistical model of large-scale signal fading is developed; and (3) the developed model is used to evaluate the system performance with recently published UWB receivers. Ondrej Novák, Richard B. Brown |
ISCAS | 2 |
| 2013 | Slew-Rate Monitoring Circuit for On-Chip Process Variation DetectionabstractThe need for efficient and accurate detection schemes to assess the impact of process variations on the parametric yield of integrated circuits has increased in the nanometer design era. In this paper, the difference of rise and fall slew is presented as another process-variation metric along with the delay in determining the relative mismatch between the drive strengths of nMOS and pMOS devices. The importance of considering both of these metrics is illustrated, and a new slew-rate monitoring circuit is presented for measuring the difference of rise and fall slew of a signal on the critical path of a circuit. Sensitivity analysis with multiple pulses as input has also been investigated. Bias generator circuits that track nMOS and pMOS threshold voltages have been incorporated, which makes the design less susceptible to process variation. Design considerations, simulation results, and characteristics of the slew-rate monitor circuitry in a 65-nm IBM CMOS process are presented, and a sensitivity of 50 MHz/50 ps for single pulse input is achieved. The measurement sensitivity of a fabricated slew-rate monitor in a 65-nm IBM CMOS technology is 0.11 V/μs, with 1089 pF as the output load of the slew-rate monitor. Amlan Ghosh, Rahul M. Rao, Jae-Joon Kim, Ching-Te Chuang, Richard B. Brown |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2010 | Dynamically Pulsed MTCMOS With Bus Encoding for Reduction of Total Power and Crosstalk NoiseabstractIncreased buffer insertion along on-chip global lines and growing amounts of leakage power have resulted in buffer-based leakage emerging as one of the chief contributors to system leakage power. In this paper, a bus system prototype is implemented in an industrial 65-nm SOI technology and measured results show up to a 45% reduction in total bus system power and an average reduction of 2.4× in standby mode leakage power. Harmander Singh, Rahul M. Rao, Kanak Agarwal 0001, Dennis Sylvester, Richard B. Brown |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2009 | A Precise Negative Bias Temperature Instability Sensor using Slew-rate Monitor CircuitryabstractNegative bias temperature instability (NBTI) has become an important cause of degradation in scaled PMOS devices, affecting power, performance, yield and reliability of circuits. This paper proposes a scheme to detect PMOS threshold voltage (VTH) degradation using on-chip slew-rate monitor circuitry. The degradation in the PMOS threshold voltage is determined with high resolution by sensing the change in rise time in a stressed ring oscillator. Simulations in IBM's 65 nm PD/SOI CMOS technology demonstrate good linearity and an output sensitivity of 0.25 mV/mV using the proposed scheme. Amlan Ghosh, Richard B. Brown, Rahul M. Rao, Ching-Te Chuang |
ISCAS | 2 |
| 2009 | A centralized supply voltage and local body bias-based compensation approach to mitigate within-die process variationabstractWith the scaling of MOSFET dimensions and the enhancements introduced to boost its performance, variation in semiconductor manufacturing has increased. The manufactured designs are usually shifted from the intended operating point, degrading the parametric yield. In this paper, we partition the chip into multiple regions with localized sensors and introduce a centralized control system with region-specific bias control to mitigate the impact of within-die (WID) process variation. An algorithm for determining the minimum required global supply voltage across all the regions and optimal body-biasing voltages for the individual regions is illustrated. This system ensures the desired frequency of operation for the chip under optimal power conditions for each of the regions. Design considerations, simulation results and power-performance characteristics of this fine-grain body biasing compensation technique are presented based on simulations of the IBM 65 nm technology. This method achieves an average reduction of 7.2% in total power dissipated across process corners while bringing the critical path delay in all modules within the desired +/− 3% of nominal delay. Amlan Ghosh, Rahul M. Rao, Richard B. Brown |
ISLPED | 3 |
| 2008 | Clock tree synthesis with data-path sensitivity matchingabstractThis paper investigates methods for minimizing the impact of process variation on clock skew using buffer and wire sizing. While most papers on clock trees ignore data-path circuit variations and most papers on data-path circuit optimization disregard clock tree variation, we consider both. Using both clock and data-path variations together, we present a novel sensitivity-matching algorithm that allows clock tree skews to be intentionally correlated with data-path sensitivities to ameliorate timing violations due to variation. Our statistical tuning shows an improvement in terms of expected clock skew and clock skew variation over previously published robust algorithms. Matthew R. Guthaus, Dennis Sylvester, Richard B. Brown |
ASP-DAC | 3 |
| 2008 | A 25MHz all-CMOS reference clock generator for XO-replacement in serial wire interfacesabstractA 25 MHz all-CMOS clock generator is demonstrated where measured performance makes it suitable for direct replacement of the reference crystal oscillator (XO) for serial wire interfaces. Fabricated in a 0.25 mum 1P5M logic CMOS process, and with no external components, the developed clock generator dissipates 59.4 mW while exhibiting plusmn152 ppm frequency error over process, plusmn10% variation in the power supply voltage and from -5-75degC. Nominal period jitter and power-on start-up latency are 3.93 psrmsand 268 mus respectively. Michael S. McCorquodale, Scott M. Pernia, Sundus Kubba, Gordy A. Carichner, Justin D. O'Day, Eric D. Marsman, Jonathan J. Kuhn, Richard B. Brown |
ISCAS | 8 |
| 2007 | Parametric Yield Analysis and Optimization in Leakage Dominated TechnologiesabstractParametric yield loss has become a serious concern in nanometer technologies. In this paper, we propose a methodology to estimate and optimize the parametric yield of a design in the presence of process variations. We discuss the impact of leakage on parametric yield given that leakage causes the parametric yield window to shrink by imposing a two-sided constraint in conjunction with performance targets on the yield window. We present a mathematical framework for yield estimation under process variation for a given power and frequency constraints. The model is validated against Monte Carlo SPICE simulations in a 90-nm CMOS process and is shown to have a typical error of less than 5%. We then demonstrate the importance of optimal supply and threshold voltage selection for yield maximization. Our results show that parametric yield is highly sensitive to supply voltage with only a 5% change in the supply voltage potentially leading to nearly 15% yield degradation. We also investigate the sensitivity of parametric yield to required frequency and power constraints. Finally, we apply the proposed framework to the problem of maximizing the shipping frequency in the presence of given yield and power constraints. Kanak Agarwal 0001, Rahul M. Rao, Dennis Sylvester, Richard B. Brown |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2006 | Process-induced skew reduction in nominal zero-skew clock treesabstractThis work develops an analytic framework for clock tree analysis considering process variations that is shown to correspond well with Monte Carlo results. The analysis framework is used in a new algorithm that constructs deterministic nominal zero-skew clock trees that have reduced sensitivity to process variation. The new algorithm uses a sampling approach to perform route embedding during a bottom-up merging phase, but does not select the best embedding until the top-down phase. This results in clock trees that exhibit a mean skew reduction of 32.4% on average and a standard deviation reduction of 40.7% as verified by Monte Carlo. The average increase in total clock tree capacitance is less than 0.02% Matthew R. Guthaus, Dennis Sylvester, Richard B. Brown |
ASP-DAC | 3 |
| 2006 | Newton: a library-based analytical synthesis tool for RF-MEMS resonatorsabstractNewton is a library-based CAD tool with an analytical synthesis engine which has been developed to support the direct synthesis of the physical design and an electromechanically equivalent model of RF-MEMS resonators based on process parameters and performance metrics. Newton provides accuracy comparable to finite element analysis while requiring a fraction of the computation and design time. A comparison of results from synthesis with Newton, design with FEA, and test results from fabricated devices is presented. Michael S. McCorquodale, James McCann, Richard B. Brown |
ASP-DAC | 3 |
| 2006 | A 16-bit, low-power microsystem with monolithic MEMS-LC clockingabstractSingle-chip systems save the power dissipation that would be required for chip-to-chip communication, resulting in compact, low-power solutions for battery-powered applications. This paper describes the design and measured performance of a fully-functional digital core with a low-jitter, on-chip, MEMS-LC clock reference. This chip has been fabricated in TSMC's 0.18mum MM/RF bulk CMOS process. Maximum power consumption of the complete microsystem is 48.78mW operating at 90MHz on a 1.8V power supply Robert M. Senger, Eric D. Marsman, Michael S. McCorquodale, Richard B. Brown |
ASP-DAC | 4 |
| 2006 | Clock buffer and wire sizing using sequential programmingabstractThis paper investigates methods for clock skew minimization using buffer and wire sizing. First, a technique that significantly improves solution quality and stability of sequential programming-based buffer/wire sizing is used. Then, a new formulation of clock skew minimization that uses quadratic programming and considers sub-critical skews in addition to the most critical skews is presented. The quality of results are verified to be more robust using Monte Carlo simulations to account for process sensitivity. For the same power budget, the sequential quadratic programming (SQP) method has better expected skew, standard deviation, and overall CPU time on average. Matthew R. Guthaus, Dennis Sylvester, Richard B. Brown |
DAC | 3 |
| 2006 | DSP architecture for cochlear implantsabstractThis paper describes low-power DSP architecture for use in cochlear implants. The microsystem, fabricated in TSMC 0.18mum CMOS, consumes 1.79mW from a 1.2V supply and occupies an area of 9.18mm2while providing the necessary programmability for high speech comprehension by patients. Standby power consumption is 330muW Eric D. Marsman, Robert M. Senger, Gordy A. Carichner, Sundus Kubba, Michael S. McCorquodale, Richard B. Brown |
ISCAS | 6 |
| 2006 | Low-latency, HDL-synthesizable dynamic clock frequency controller with self-referenced hybrid clockingabstractA low-latency, HDL-synthesizable dynamic clock frequency controller is presented as a time-efficient alternative to full-custom implementations. Frequency division of a fully integrated hybrid temperature-compensated LC oscillator (TC-LCO) and ring oscillator clock reference avoids PLL locking delays to enable low-latency, hazard-free frequency selection on an actively running CPU. Fabricated in 0.18mum CMOS as part of a low-power SoC microsystem, the circuit dissipates 480muW at 1.8V Robert M. Senger, Eric D. Marsman, Gordy A. Carichner, Sundus Kubba, Michael S. McCorquodale, Richard B. Brown |
ISCAS | 6 |
| 2006 | Integrated electrochemical neurosensorsabstractArrays of silicon neurosensors have been fabricated in our group for detection of both electrical signals and neurotransmitter levels in human neuron cultures. Neurochemical sensing of dopamine and its metabolites is provided by voltammetry. Passive neurochemical arrays have been tested in living human neuron cultures throughout a study period of seventy-five days. Calibration curves for dopamine taken in culture media with equipment optimized for the sensors suggests detection limits for dopamine below 100nM. To improve response, prototype devices incorporating active circuitry were developed. These active devices were formed by post-processing standard CMOS die fabricated through the MOSIS service, to form the sensor-specific features. We also describe a fully-differential potentiostat which has been developed for implantable biological applications and briefly describe future directions for our work. Timothy D. Strong, Steven M. Martin, Robert F. Franklin, Richard B. Brown |
ISCAS | 4 |
| 2006 | A dual-VDD boosted pulsed bus technique for low power and low leakage operationabstractIn this paper, we propose a new dual-VDD bus technique that is well suited for low power operation. This technique adapts a static pulsed bus architecture to use dual-VDD power supplies. During quiescent periods, the bus system idles at the lower of the two VDD supplies, thereby lowering static power dissipation. When actively transitioning, the inverters in the bus system are temporarily boosted to the higher VDD supply to provide the needed drive strength for performance. Since the VDD boosting is done in a pulsed manner, the bus system is in a high VDD state only when required, ensuring lower power operation without sacrificing performance. This technique yields up to a 50% reduction in total power over traditional static buses and up to a 35% reduction in total power over standard static pulsed buses, with a 12-15% delay improvement. Harmander Singh, Robert M. Senger, Dennis Sylvester, Richard B. Brown, Kevin J. Nowka |
ISLPED | 4 |
| 2005 | Compiler Managed Dynamic Instruction Placement in a Low-Power Code CacheabstractModern embedded microprocessors use low power on-chip memories called scratch-pad memories to store frequently executed instructions and data. Unlike traditional caches, scratch-pad memories lack the complex tag checking and comparison logic, thereby proving to be efficient in area and power. In this work, we focus on exploiting scratch-pad memories for storing hot code segments within an application. Static placement techniques focus on placing the most frequently executed portions of programs into the scratch-pad. However, static schemes are inherently limited by not allowing the contents of the scratch-pad memory to change at run time. In a large fraction of applications, the instruction memory footprints exceed the scratch-pad memory size, thereby limiting the usefulness of the scratch-pad. We propose a compiler managed dynamic placement algorithm, wherein multiple hot code sequences, or traces, are overlapped with each other in the scratch-pad memory at different points in time during execution. Special copy instructions are provided to copy the traces into the scratch-pad memory at run-time. Using a power estimate, the compiler initially selects the most frequent traces in an application for relocation into the scratch-pad memory. Through iterative code motion and redundancy elimination, copy instructions are inserted in infrequently executed regions of the code. For a 64-byte code cache, the compiler managed dynamic placement achieves an average of 64% energy improvement over the static solution in a low-power embedded microcontroller. Rajiv A. Ravindran, Pracheeti D. Nagarkar, Ganesh S. Dasika, Eric D. Marsman, Robert M. Senger, Scott A. Mahlke, Richard B. Brown |
CGO | 7 |
| 2005 | Optimization objectives and models of variation for statistical gate sizingabstractThis paper approaches statistical optimization by examining gate delay variation models and optimization objectives. Most previous work on statistical optimization has focused exclusively on the optimization algorithms without considering the effects of the variation models and objective functions. This work empirically derives a simple variation model that is then used to optimize for robustness. Optimal results from example circuits used to study the effect of the statistical objective function on parametric yield. Matthew R. Guthaus, Natesan Venkateswaran, Vladimir Zolotov, Dennis Sylvester, Richard B. Brown |
ACM Great Lakes Symposium on VLSI | 5 |
| 2005 | Partitioning Variables across Register Windows to Reduce Spill Code in a Low-Power ProcessorabstractLow-power embedded processors utilize compact instruction encodings to achieve small code size. Such encodings place tight restrictions on the number of bits available to encode operand specifiers and, thus, on the number of architected registers. As a result, performance and power are often sacrificed as the burden of operand supply is shifted from the register file to the memory due to the limited number of registers. In this paper, we investigate the use of a windowed register file to address this problem by providing more registers than allowed in the encoding. The registers are organized as a set of identical register windows where, at each point in the execution, there is a single active window. Special window management instructions are used to change the active window and to transfer values between windows. This design gives the appearance of a large register file without compromising the instruction encoding. To support the windowed register file, we designed and implemented a graph partitioning-based compiler algorithm that partitions program variables and temporaries referenced within a procedure across multiple windows. On a 16-bit embedded processor, an average of 11 percent improvement in application performance and 25 percent reduction in system power was achieved as an 8-register design was scaled from one to two windows. Rajiv A. Ravindran, Robert M. Senger, Eric D. Marsman, Ganesh S. Dasika, Matthew R. Guthaus, Scott A. Mahlke, Richard B. Brown |
IEEE Trans. Computers | 7 |
| 2004 | Approaches to run-time and standby mode leakage reduction in global busesabstractIn this paper, we present various design approaches to leakage minimization in global repeaters. We demonstrate the applicability of the MTCMOS scheme to global repeaters for leakage reduction. We then analyze two design approaches called Duplicated Skewed Buses and Skewed Pulsed Buses. We show that significant reduction in standby leakage power can be obtained using these approaches while providing significant improvements in performance. We also illustrate the use of these proposed techniques with the MTCMOS approach to obtain further savings in leakage power. Simulations results in a 90nm process show that skewed pulsed buses with MTCMOS can provide 20% improvement in performance with over 25% reduction in active mode leakage and nearly 100X reduction in standby mode leakage. Rahul M. Rao, Kanak Agarwal 0001, Dennis Sylvester, Richard B. Brown, Kevin J. Nowka, Sani R. Nassif |
ISLPED | 4 |
| 2003 | Increasing the number of effective registers in a low-power processor using a windowed register fileabstractLow-power embedded processors utilize compact instruction encodings to achieve small code size. Instruction sizes of 8 to 16 bits are common. Such encodings place tight restrictions on the number of bits available to encode operand specifiers, and thus on the number of architected registers. The central problem with this approach is that performance and power are often sacrificed as the burden of operand supply is shifted from the register file to the memory due to the limited number of registers. In this paper, we investigate the use of a windowed register file to address this problem by providing more registers than allowed in the encoding. The registers are organized as a set of identical register windows where at each point in the execution there is a single active window. Special window management instructions are used to change the active window and to transfer values between windows. The goal of this design is to give the appearance of a large register file without compromising the instruction encoding. To support the windowed register file, we designed and implemented a novel graph partitioning based compiler algorithm that partitions virtual registers within a given procedure across multiple windows. On a 16-bit embedded processor with a parameterized register window, an average of 10% improvement in application performance and 7% reduction in system power was achieved as an eight-register design was scaled from one to four windows. Rajiv A. Ravindran, Robert M. Senger, Eric D. Marsman, Ganesh S. Dasika, Matthew R. Guthaus, Scott A. Mahlke, Richard B. Brown |
CASES | 7 |
| 2003 | A 16-bit mixed-signal microsystem with integrated CMOS-MEMS clock referenceabstractIn this work, we report on an unprecedented design where digital, analog, and MEMS technologies are combined to realize a general-purpose single-chip CMOS microsystem. The convergence of these technologies has enabled the development of a low power, portable microinstrument ideally suited for controlling environmental and bio-implantable sensors. Robert M. Senger, Eric D. Marsman, Michael S. McCorquodale, Fadi H. Gebara, Keith L. Kraver, Matthew R. Guthaus, Richard B. Brown |
DAC | 7 |
| 2003 | A Top-Down Microsystems Design Methodology and Associated Challenges abstractAn overview of microsystems technology is presented along with a discussion of the recent trends and challenges associated with its development. A typical bottom-up design methodology is described and we propose, in contrast, an efficient and effective top-down methodology. We illustrate its implementation with the development of a microsystem design that has been completed and fabricated in CMOS technology. Gaps in the tool capabilities are identified and suggestions for future directions in CAD tool support for microsystems technology are presented. Michael S. McCorquodale, Fadi H. Gebara, Keith L. Kraver, Eric D. Marsman, Robert M. Senger, Richard B. Brown |
DATE | 6 |
| 2003 | A Heuristic to Determine Low Leakage Sleep State Vectors for CMOS Combinational Circuits
Rahul M. Rao, Frank Liu 0001, Jeffrey L. Burns, Richard B. Brown |
ICCAD | 4 |
| 2003 | New optimal design strategies and analysis of ultra-low leakage circuits for nano-scale SOI technologyabstractThis paper proposes new SOI circuit strategies for simultaneous reduction of standby gate and sub-threshold leakages. Various enhanced MTCMOS design alternatives are analyzed. A new method for assigning the V/sub TH/ and sizes of header and footer transistors is proposed, and stacking of headers/footers is analyzed. The optimum stacking height and tapering/sizing ratio under various design constraints are determined. Our strategies reduce MTCMOS standby leakage further by as much as 20/spl times/ and reduce virtual supply noise by 15%. Koushik K. Das, Rajiv V. Joshi, Ching-Te Chuang, Peter W. Cook, Richard B. Brown |
ISLPED | 5 |
| 2003 | Efficient techniques for gate leakage estimationabstractGate leakage current is expected to be the dominant leakage component in future technology generations. In this paper, we propose methods for steady-state gate leakage estimation based on state characterization. An efficient technique for pattern-dependent gate leakage estimation is presented. Further, we propose the use of this technique for estimating the average gate leakage of a circuit using pattern-independent probabilistic analysis. Results on a large set of benchmark ISCAS circuits show an accuracy within 5 % of SPICE results with 500X to 50000X speed improvement. Rahul M. Rao, Jeffrey L. Burns, Anirudh Devgan, Richard B. Brown |
ISLPED | 4 |
| 2003 | Evaluation of Dynamic-Threshold Logic for Low-Power VLSI Design in 0.13um PD-SOI
Alan J. Drake, Kevin J. Nowka, Richard B. Brown |
VLSI-SOC | 3 |
| 2003 | Microsystem and SoC Design with UMIPS
Michael S. McCorquodale, Eric D. Marsman, Robert M. Senger, Fadi H. Gebara, Richard B. Brown |
VLSI-SOC | 5 |
| 2003 | Micropotentiometric sensorsabstractThis paper covers recent developments in microfabricated potentiometric liquid chemical sensors. It includes a discussion of the various types of solid-state potentiometric sensors, including the inherent and practical challenges involved in implementing and commercializing the devices. The paper also presents recent advances that overcome some of these difficulties, including progress toward compatible microreference electrodes. Hakhyun Nam, Geun S. Cha, Timothy D. Strong, Jeonghan Ha, Jun Ho Sim, Robert W. Hower, Steven M. Martin, Richard B. Brown |
Proc. IEEE | 8 |
| 2000 | CGaAs PowerPC FXUabstractThe development of a PowerPC™ fixed-point execution unit (FXU) in a resource limited, radiation-hard technology is described. Detailed architectural studies led to a design which maximizes performance in a small transistor count implementation. Manufactured in Motorola's 0.5—µm Complementary Gallium Arsenide process, the device operates from 0.9 to 1.9 V with a nominal frequency of 25 MHz at 1.3 V, dissipating 274 mW. Alan J. Drake, Todd D. Basso, Spencer M. Gold, Keith L. Kraver, Phiroze N. Parakh, Claude R. Gauthier, P. Sean Stetson, Richard B. Brown |
DAC | 8 |
| 1999 | Crosstalk constrained global route embeddingabstractArticle Free Access Share on Crosstalk constrained global route embedding Authors: Phiroze N. Parakh Advanced Computer Architecture Laboratory, Department of Electrical Engineering and Computer Science, University of Michigan Advanced Computer Architecture Laboratory, Department of Electrical Engineering and Computer Science, University of MichiganView Profile , Richard B. Brown Advanced Computer Architecture Laboratory, Department of Electrical Engineering and Computer Science, University of Michigan Advanced Computer Architecture Laboratory, Department of Electrical Engineering and Computer Science, University of MichiganView Profile Authors Info & Claims ISPD '99: Proceedings of the 1999 international symposium on Physical designApril 1999 Pages 201–206https://doi.org/10.1145/299996.300077Online:12 April 1999Publication History 9citation236DownloadsMetricsTotal Citations9Total Downloads236Last 12 Months3Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Phiroze N. Parakh, Richard B. Brown |
ISPD | 2 |
| 1998 | Congestion Driven Quadratic PlacementabstractThis paper introduces and demonstrates an extension to quadratic placement that accounts for wiring congestion. The algorithm uses an A* router and line-probe heuristics on region-based routing graphs to compute routing cost. The interplay between routing analysis and quadratic placement using a growth matrix permits global treatment of congestion. Further reduction in congestion is obtained by the relaxation of pin constraints. Experiments show improvements in wireability. Phiroze N. Parakh, Richard B. Brown, Karem A. Sakallah |
DAC | 2 |
| 1998 | High-level design verification of microprocessors via error modelingabstractA design verification methodology for microprocessor hardware based on modeling design errors and generating simulation vectors for the modeled errors via physical fault testing techniques is presented. We have systematically collected design error data from a number of microprocessor design projects. The error data is used to derive error models suitable for design verification testing. A class of basic error models is identified and shown to yield tests that provide good coverage of common error types. To improve coverage for more complex errors, a new class of conditional error models is introduced. An experiment to evaluate the effectiveness of our methodology is presented. Single actual design errors are injected into a correct design, and it is determined if the methodology will generate a test that detects the actual errors. The experiment has been conducted for two microprocessor designs and the results indicate that very high coverage of actual design errors can be obtained with test sets that are complete for a small number of synthetic error models. David Van Campenhout, Hussain Al-Asaad, John P. Hayes, Trevor N. Mudge, Richard B. Brown |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 1998 | Overview of complementary GaAs technology for high-speed VLSI circuitsabstractA self-aligned complementary GaAs (CGaAs) technology (developed at Motorola) for low-power, portable, digital and mixed-mode circuits is being extended to address high-speed VLSI circuit applications. The process supports full complementary, unipolar (pseudo-DCFL), source-coupled, and dynamic (domino) logic families. Though this technology is not yet mature, it is years ahead of CMOS in terms of fast gate delays at low power supply voltages. Complementary circuits operating at 0.9 V have demonstrated power-delay products of 0.01 /spl mu/W/MHz/gate. Propagation delays of unipolar circuits are as low as 25 ps. Logic families can be mixed on a chip to trade power for delay. CGaAs is being evaluated for VLSI applications through the design of a PowerPC-architecture microprocessor. Richard B. Brown, Bruce Bernhardt, M. LaMacchia, J. Abrokwah, Phiroze N. Parakh, Todd D. Basso, Spencer M. Gold, S. Stetson, Claude R. Gauthier, D. Foster, B. Crawforth, T. McQuire, Karem A. Sakallah, Ronald J. Lomax, Trevor N. Mudge |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1997 | Multilevel Optimization of Pipelined CachesabstractThis paper formulates and shows how to solve the problem of selecting the cache size and depth of cache pipelining that maximizes the performance of a given instruction-set architecture. The solution combines trace-driven architectural simulations and the timing analysis of the physical implementation of the cache. Increasing cache size tends to improve performance but this improvement is limited because cache access time increases with its size. This trade-off results in an optimization problem we referred to as multilevel optimization, because it requires the simultaneous consideration of two levels of machine abstraction: the architectural level and the physical implementation level. The introduction of pipelining permits the use of larger caches without increasing their apparent access time, however, the bubbles caused by load and branch delays limit this technique. In this paper we also show how multilevel optimization can be applied to pipelined systems if software- and hardware-based strategies are considered for hiding the branch and load delays. The multilevel optimization technique is illustrated with the design of a pipelined cache for a high clock rate MIPS-based architecture. The results of this design exercise show that, because processors with pipelined caches can have shorter CPU cycle times and larger caches, a significant performance advantage is gained by using two or three pipeline stages to fetch data from the cache. Of course, the results are only optimal for the implementation technologies chosen for the design exercise; other choices could result in quite different optimal designs. The exercise is primarily to illustrate the steps in the design of pipelined caches using multilevel optimization; however, it does exemplify the importance of pipelined caches if high clock rate processors are to achieve high performance. Kunle Olukotun, Trevor N. Mudge, Richard B. Brown |
IEEE Trans. Computers | 3 |
| 1996 | Ravel-XL: a hardware accelerator for assigned-delay compiled-code logic gate simulationabstractRavel-XL is a single-board hardware accelerator for gate-level digital logic simulation. It uses a standard levelized-code approach to statically schedule gate evaluations. However, unlike previous approaches based on levelized-code scheduling, it is not limited to zero- or unit-delay gate models and can provide timing accuracy comparable to that obtained from event-driven methods. We review the synchronous waveform algebra that forms the basis of the Ravel-XL simulation algorithm, present an architecture for its hardware realization, and describe an implementation of this architecture as a single VLSI chip. The chip has about 900000 transistors on a die that is approximately 1.4 cm/sup 2/, requires a 256 pin package and is designed to run at 33 MHz. A Ravel-XL board consisting of the processor chip and local instruction and data memory can simulate up to one billion gates at a rate of approximately 6.6 million gate evaluations per second. To better appreciate the tradeoffs made in designing Ravel-XL, we compare its capabilities to those of other commercial and research software simulators and hardware accelerators. Michael A. Riepe, João Marques-Silva 0001, Karem A. Sakallah, Richard B. Brown |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 1995 | The Aurora RAM CompilerabstractArticle The Aurora RAM compiler Share on Authors: Ajay Chandna University of Michigan, Department of Electrical Engineering & Computer Science, Ann Arbor, MI University of Michigan, Department of Electrical Engineering & Computer Science, Ann Arbor, MIView Profile , C. David Kibler Hewlett Packard Company, 3404 East Harmony Rd., Ft. Collins, CO Hewlett Packard Company, 3404 East Harmony Rd., Ft. Collins, COView Profile , Richard B. Brown University of Michigan, Department of Electrical Engineering & Computer Science, Ann Arbor, MI University of Michigan, Department of Electrical Engineering & Computer Science, Ann Arbor, MIView Profile , Mark Roberts University of Michigan, Department of Electrical Engineering & Computer Science, Ann Arbor, MI University of Michigan, Department of Electrical Engineering & Computer Science, Ann Arbor, MIView Profile , Karem A. Sakallah University of Michigan, Department of Electrical Engineering & Computer Science, Ann Arbor, MI University of Michigan, Department of Electrical Engineering & Computer Science, Ann Arbor, MIView Profile Authors Info & Claims DAC '95: Proceedings of the 32nd annual ACM/IEEE Design Automation ConferenceJanuary 1995 Pages 261–266https://doi.org/10.1145/217474.217539Online:01 January 1995Publication History 0citation326DownloadsMetricsTotal Citations0Total Downloads326Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Ajay Chandna, C. David Kibler, Richard B. Brown, Mark Roberts, Karem A. Sakallah |
DAC | 3 |
| 1994 | Resource Allocation in a High Clock Rate MicroprocessorabstractThis paper discusses the design of a high clock rate (300MHz) processor. The architecture is described, and the goals for the design are explained. The performance of three processor models is evaluated using trace-driven simulation. A cost model is used to estimate the resources required to build processors with varying sizes of on-chip memories, in both single and dual issue models. Recommendations are then made to increase the effectiveness of each of the models. Michael Upton, Thomas R. Huff, Trevor N. Mudge, Richard B. Brown |
ASPLOS | 4 |
| 1994 | Design Tradeoffs for Software-Managed TLBsabstractAn increasing number of architectures provide virtual memory support through software-managed TLBs. However, software management can impose considerable penalties that are highly dependent on the operating system's structure and its use of virtual memory. This work explores software-managed TLB design tradeoffs and their interaction with a range of monolithic and microkernel operating systems. Through hardware monitoring and simulation, we explore TLB performance for benchmarks running on a MIPS R2000-based workstation running Ultrix, OSF/1, and three versions of Mach 3.0. Richard Uhlig, David Nagle, Tim J. Stanley, Trevor N. Mudge, Stuart Sechrest, Richard B. Brown |
ACM Trans. Comput. Syst. | 6 |
| 1993 | Ravel-XL: A Hardware Accelerator for Assigned-Delay Compiled-Code Logic Gate SimulationabstractWe describe the design of Ravel-XL, a hardware accelerator for assigned-delay compiled-code logic gate simulation. After a brief review of the underlying Ravel simulation algorithm, we describe the major factors that influenced the hardware design, particularly the interaction between the instruction execution and operand bandwidth requirements. The initial CMOS VLSI implementation of the accelerator contains a 2K word data cache, occupies approximately 1.9 cm/sup 2/ of die area with 256 pins and approximately 900,000 transistors. Simulation results predicts operation at a clock rate of 33 MHz. This provides a speedup of about 50 over the software implementation of Ravel, about 50 over a compiled event-driven simulator, and about 500 over an interpreted event-driven simulator. We conclude with some planned design improvements that will allow an approximate doubling of the clock rate.> Michael A. Riepe, João Marques-Silva 0001, Karem A. Sakallah, Richard B. Brown |
ICCD | 4 |
| 1993 | Design Tradeoffs for Software-Managed TLBsabstractAn increasing number of architectures provide virtual memory support through software-managed TLBs. However, software management can impose considerable penalties, which are highly dependent on the operating system's structure and its use of virtual memory. This work explores software-managed TLB design tradeoffs and their interaction with a range of operating systems including monolithic and microkernel designs. Through hardware monitoring and simulations, we explore TLB performance for benchmarks running on a MIPS R2000-based workstation running Ultrix, OSF/1, and three versions of mach 3.0. David Nagle, Richard Uhlig, Tim J. Stanley, Stuart Sechrest, Trevor N. Mudge, Richard B. Brown |
ISCA | 6 |
| 1993 | A microarchitectural performance evaluation of a 3.2 Gbyte/s microprocessor busabstractThe conventional classification of inter-instruction dependencies (data, anti and output dependencies) provides a basic scheme for the analysis of pipeline hazards in pipelined instruction set processors. However, it does not consider the relative spatial positions of micro-operations in the pipeline, thus providing limited hints to hardware designers and compiler writers about the hazard resolution in generalized pipeline structures. The authors propose an extension to the conventional classification of dependencies, which is capable of encapsulating the spatial/temporal relationship and providing precise hardware/software resolution strategies. With the extended classification and its associated hardware/software resolution strategies, they are able to systematically analyze the potential register-related pipeline hazards for a given pipeline structure, determine appropriate resolution strategies, and explore the tradeoff between hardware and software complexities. The methodology enables the systematic synthesis of high performance pipelined micro-architectures, and is useful to derive the back-end of the supporting compilers.> Tim J. Stanley, Michael Upton, Patrick Sherhart, Trevor N. Mudge, Richard B. Brown |
MICRO | 5 |
| 1992 | Performance Optimization of Pipelined Primary CachesabstractThe CPU cycle time of a high-performance processor is usually determined by the access time of the primary cache. As processors speeds increase, designers will have to increase the number of pipeline stages used to fetch data from the cache in order to reduce the dependence of CPU cycle time on cache access time. This paper studies the performance advantages of a pipelined cache for a GaAs implementation of the MIPS based architecture using a design methodology that includes long traces of multiprogrammed applications and detailed timing analysis. The study evaluates instruction and data caches with various pipeline depths, cache sizes, block sizes, and refill penalties. The impact on CPU cycle time of these alternatives is also factored into the evaluation. Hardware-based and software-based strategies are considered for hiding the branch and load delays which may be required to avoid pipeline hazards. The results show that software-based methods for mitigating the penalty of branch delays can be as successful as the hardware-based branch-target buffer approach, despite the code-expansion inherent in the software methods. The situation is similar for load delays; while hardware-based dynamic methods hide more delay cycles than do static approaches, they may give up the advantage by extending the cycle time. Because these methods are quite successful at hiding small numbers of branch and load delays, and because processors with pipelined caches also have shorter CPU cycle times and larger caches, a significant performance advantage is gained by using two to three pipeline stages to fetch data from the cache. Kunle Olukotun, Trevor N. Mudge, Richard B. Brown |
ISCA | 3 |
| 1991 | A high resolution current stimulating probe for use in neural prosthesesabstractThe authors describe a high resolution monolithic current stimulating probe which is a multichannel microprobe capable of delivering precisely controlled charge to highly-localized regions of tissue (for example, the auditory nervous system) on a chronic basis. This chip has 8 stimulating sites and provides 7 bits of current level control, that is -126 mu A to +126 mu A with 2 mu A resolution. In order to improve controllability and reduce the cost of the circuitry, the integrated stimulating microprobe is addressable and self-testing. Implemented features include self-test capabilities, a bipolar current forcing scheme for the effective stimulation of the tissue and a user-specified time-out scheme which makes the microstimulator safe. The total number of pads is 14, eight of which are stimulating sites, with the others being used for overall circuit control and power supplies. The chip is designed in an n-epitaxial, p-well, 1.2 mu m, double metal CMOS process with an area of 0.24 mm/sup 2/ excluding bonding pads.> D. Kang, Richard B. Brown, Kensall D. Wise |
Great Lakes Symposium on VLSI | 3 |
| 1991 | Implementing a Cache for a High-Performance GaAs MicroprocessorabstractIn the near future, microprocessor systems with very high clock rates will use multichip module (MCM) pack- aging technology to reduce chip-crossing delays. In this paper we present the results of a study for the design of a 250 MHz Gallium Arsenide (GaAs) microprocessor t,lrat employs h4CM technology to improve performance. The design study for the resulting two-level split cache st.arts with a baseline cache architecture and then ex- amines the following aspects: 1) primary cache size and degree of associativity; 2) primary data-cache write pol- icy; 3) secondary cache size and organization; 4) pri- mary cache fetch size; 5) concurrency between instruc- tion and data accesses. A trace-driven simulator is used to analyze each design's performance. The results show that memory access time and page-size constraints ef- Cectively limit the size of the primary data and instruc- tion caches to 4I<W (16KB). For such cache sizes, a write-through policy is better than a write-back policy. Three cache mechanisms that contribute to improved performance are introduced. The first is a variant of the write-through policy called write-only. This write policy provides most of the performance benefits of sub- Ilod placernenl without extra valid bits. The second, is the use of a split secondary cache. Finally, the third mechanism allows loads to pass stores without associa- tive matching. Keywords-two-level caches, high performance pro- cessors, gallium arsenide, multichip modules, trace- driven cache simulation. Kunle Olukotun, Trevor N. Mudge, Richard B. Brown |
ISCA | 3 |