VLDB 2026 Research / reviewers in the wild / expert
David A. Patterson 0001
dblp:p/DAPatterson
· DBLP profile ↗
100ranked-venue papers
16as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 63 · 13 first-author · 2 since 2021Software engineering, systems software and programming languages · 33 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 15 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-authorArtificial intelligence and machine learning · 3Theory of computation · 3Computer networks · 2Security and privacy · 2Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How to Give AI a Bad Carbon Footprint
David A. Patterson 0001 |
CIDR | 1 |
| 2025 | Databases in the Era of Memory-Centric Computing
Yannis Chronis, Anastasia Ailamaki, Lawrence Benson, Helena Caminal, Jana Giceva, David A. Patterson 0001, Eric Sedlar, Lisa Wu Wills |
CIDR | 6 |
| 2023 | TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for EmbeddingsabstractIn response to innovations in machine learning (ML) models, production workloads changed radically and rapidly. TPU v4 is the fifth Google domain specific architecture (DSA) and its third supercomputer for such ML models. Optical circuit switches (OCSes) dynamically reconfigure its interconnect topology to improve scale, availability, utilization, modularity, deployment, security, power, and performance; users can pick a twisted 3D torus topology if desired. Much cheaper, lower power, and faster than Infiniband, OCSes and underlying optical components are <5% of system cost and <3% of system power. Each TPU v4 includes SparseCores, dataflow processors that accelerate models that rely on embeddings by 5x--7x yet use only 5% of die area and power. Deployed since 2020, TPU v4 outperforms TPU v3 by 2.1x and improves performance/Watt by 2.7x. The TPU v4 supercomputer is 4x larger at 4096 chips and thus nearly 10x faster overall, which along with OCS flexibility and availability allows a large language model to train at an average of ~60% of peak FLOPS/second. For similar sized systems, it is ~4.3x--4.5x faster than the Graphcore IPU Bow and is 1.2x--1.7x faster and uses 1.3x--1.9x less power than the Nvidia A100. TPU v4s inside the energy-optimized warehouse scale computers of Google Cloud use ~2--6x less energy and produce ~20x less CO2e than contemporary DSAs in typical on-premise data centers. Norman P. Jouppi, George Kurian, Sheng Li 0007, Peter C. Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, Cliff Young, Zongwei Zhou, David A. Patterson 0001 |
ISCA | 14 |
| 2021 | Ten Lessons From Three Generations Shaped Google's TPUv4i : Industrial ProductabstractGoogle deployed several TPU generations since 2015, teaching us lessons that changed our views: semi-conductor technology advances unequally; compiler compatibility trumps binary compatibility, especially for VLIW domain-specific architectures (DSA); target total cost of ownership vs initial cost; support multi-tenancy; deep neural networks (DNN) grow 1.5X annually; DNN advances evolve workloads; some inference tasks require floating point; inference DSAs need air-cooling; apps limit latency, not batch size; and backwards ML compatibility helps deploy DNNs quickly. These lessons molded TPUv4i, an inference DSA deployed since 2020. Norman P. Jouppi, Doe Hyun Yoon, Matthew Ashcraft, Mark Gottscho, Thomas B. Jablin, George Kurian, James Laudon, Sheng Li 0007, Peter C. Ma, Thomas Norrie, Nishant Patil, Sushma Prasad, Cliff Young, Zongwei Zhou, David A. Patterson 0001 |
ISCA | 16 |
| 2020 | Google's Training Chips Revealed: TPUv2 and TPUv3abstractThis article consists only of a collection of slides from the author's conference presentation. Thomas Norrie, Nishant Patil, Doe Hyun Yoon, George Kurian, Sheng Li 0007, James Laudon, Cliff Young, Norman P. Jouppi, David A. Patterson 0001 |
Hot Chips Symposium | 9 |
| 2020 | Kira: Processing Astronomy Imagery Using Big Data TechnologyabstractScientific analyses commonly compose multiple single-process programs into a dataflow. An end-to-end dataflow of single-process programs is known as a many-task application. Typically, HPC tools are used to parallelize these analyses. In this work, we investigate an alternate approach that uses Apache Spark-a modern platform for data intensive computing-to parallelize many-task applications. We implement Kira, a flexible and distributed astronomy image processing toolkit, and its Source Extractor (Kira SE) application. Using Kira SE as a case study, we examine the programming flexibility, dataflow richness, scheduling capacity and performance of Apache Spark running on the Amazon EC2 cloud. By exploiting data locality, Kira SE achieves a 4.1× speedup over an equivalent C program when analyzing a 1TB dataset using 512 cores on the Amazon EC2 cloud. Furthermore, Kira SE on the Amazon EC2 cloud achieves a 1.8× speedup over the C program on the NERSC Edison supercomputer. A 128-core Amazon EC2 cloud deployment of Kira SE using Spark Streaming can achieve a second-scale latency with a sustained throughput of 800 MB/s. Our experience with Kira demonstrates that data intensive computing platforms like Apache Spark are a performant alternative for many-task scientific applications. Zhao Zhang 0007, Kyle Barbary, Frank A. Nothaft, Evan Randall Sparks, Oliver Zahn, Michael J. Franklin, David A. Patterson 0001, Saul Perlmutter |
IEEE Trans. Big Data | 7 |
| 2019 | FPGA Accelerated INDEL Realignment in the CloudabstractThe amount of data being generated in genomics is predicted to be between 2 and 40 exabytes per year for the next decade, making genomic analysis the new frontier and the new challenge for precision medicine. This paper explores targeted deployment of hardware accelerators in the cloud to improve the runtime and throughput of immensescale genomic data analyses. In particular, INDEL (INsertion/DELetion) realignment is a critical operation that enables diagnostic testings of cancer through error correction prior to variant calling. It is the slowest part of the somatic (cancer) genomic analysis pipeline, the alignment refinement pipeline, and represents roughly one-third of the execution time of timesensitive diagnostics for acute cancer patients. To accelerate genomic analysis, this paper describes a hardware accelerator for INDEL realignment (IR), and a hardware-software framework leveraging FPGAs-as-a-service in the cloud. We chose to implement genomics analytics on FPGAs because genomic algorithms are still rapidly evolving (e.g. the de facto standard “GATK Best Practices” has had five releases since January of this year). We chose to deploy genomics accelerators in the cloud to reduce capital expenditure and to provide a more quantitative performance and cost analysis. We built and deployed a sea of IR accelerators using our hardware-software accelerator development framework on AWS EC2 F1 instances. We show that our IR accelerator system performed 81× better than multi-threaded genomic analysis software while being 32× more cost efficient. Lisa Wu Wills, David Bruns-Smith, Frank A. Nothaft, Qijing Huang 0001, Sagar Karandikar, Johnny Le, Andrew Lin, Howard Mao, Brendan Sweeney, Krste Asanovic, David A. Patterson 0001, Anthony D. Joseph |
HPCA | 11 |
| 2017 | Reducing Pagerank Communication via Propagation BlockingabstractReducing communication is an important objective, as it can save energy or improve the performance of a communication-bound application. The graph algorithm PageRank computes the importance of vertices in a graph, and it serves as an important benchmark for graph algorithm performance. If the input graph to PageRank has poor locality, the execution will need to read many cache lines from memory, some of which may not be fully utilized. We present propagation blocking, an optimization to improve spatial locality, and we demonstrate its application to PageRank. In contrast to cache blocking which partitions the graph, we partition the data transfers between vertices (propagations). If the input graph has poor locality, our approach will reduce communication. Our approach reduces communication more than conventional cache blocking if the input graph is sufficiently sparse or if number of vertices is sufficiently large relative to the cache size. To evaluate our approach, we use both simple analytic models to gain insights and precise hardware performance counter measurements to compare implementations on a suite of 8 real-world and synthetic graphs. We demonstrate our parallel implementations substantially outperform prior work in execution time and communication volume. Although we present results for PageRank, propagation blocking could be generalized to SpMV (sparse matrix multiplying dense vector) or other graph programming models. Scott Beamer, Krste Asanovic, David A. Patterson 0001 |
IPDPS | 3 |
| 2017 | In-Datacenter Performance Analysis of a Tensor Processing UnitabstractMany architects believe that major improvements in cost-energy-performance must now come from domain-specific hardware. This paper evaluates a custom ASIC---called a Tensor Processing Unit (TPU) --- deployed in datacenters since 2015 that accelerates the inference phase of neural networks (NN). The heart of the TPU is a 65,536 8-bit MAC matrix multiply unit that offers a peak throughput of 92 TeraOps/second (TOPS) and a large (28 MiB) software-managed on-chip memory. The TPU's deterministic execution model is a better match to the 99th-percentile response-time requirement of our NN applications than are the time-varying optimizations of CPUs and GPUs that help average throughput more than guaranteed latency. The lack of such features helps explain why, despite having myriad MACs and a big memory, the TPU is relatively small and low power. We compare the TPU to a server-class Intel Haswell CPU and an Nvidia K80 GPU, which are contemporaries deployed in the same datacenters. Our workload, written in the high-level TensorFlow framework, uses production NN applications (MLPs, CNNs, and LSTMs) that represent 95% of our datacenters' NN inference demand. Despite low utilization for some applications, the TPU is on average about 15X -- 30X faster than its contemporary GPU or CPU, with TOPS/Watt about 30X -- 80X higher. Moreover, using the CPU's GDDR5 memory in the TPU would triple achieved TOPS and raise TOPS/Watt to nearly 70X the GPU and 200X the CPU. Norman P. Jouppi, Cliff Young, Nishant Patil, David A. Patterson 0001, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, Rick Boyle, Pierre-luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tara Vazir Ghaemmaghami, Rajendra Gottipati, William Gulland, Robert Hagmann, Richard Ho 0001, Doug Hogberg, John Hu, Robert Hundt, Dan Hurt, Julian Ibarz, Aaron Jaffey, Alek Jaworski, Alexander Kaplan, Harshit Khaitan, Daniel Killebrew, Andy Koch, Steve Lacy, James Laudon, James Law, Diemthu Le, Chris Leary, Zhuyuan Liu, Kyle Lucke, Alan Lundin, Gordon MacKean, Adriana Maggiore, Maire Mahony, Kieran Miller, Rahul Nagarajan, Ravi Narayanaswami, Ray Ni, Kathy Nix, Thomas Norrie, Mark Omernick, Narayana Penukonda, Andy Phelps, Jonathan Ross, Amir Salek, Emad Samadiani, Chris Severn, Gregory Sizikov, Matthew Snelham, Jed Souter, Dan Steinberg, Andy Swing, Mercedes Tan, Gregory Thorson, Horia Toma, Erick Tuttle, Vijay Vasudevan, Richard Walter, Walter Wang, Eric Wilcox, Doe Hyun Yoon |
ISCA | 4 |
| 2016 | GenAp: a distributed SQL interface for genomic dataabstractBACKGROUND: The impressively low cost and improved quality of genome sequencing provides to researchers of genetic diseases, such as cancer, a powerful tool to better understand the underlying genetic mechanisms of those diseases and treat them with effective targeted therapies. Thus, a number of projects today sequence the DNA of large patient populations each of which produces at least hundreds of terra-bytes of data. Now the challenge is to provide the produced data on demand to interested parties. RESULTS: In this paper, we show that the response to this challenge is a modified version of Spark SQL, a distributed SQL execution engine, that handles efficiently joins that use genomic intervals as keys. With this modification, Spark SQL serves such joins more than 50× faster than its existing brute force approach and 8× faster than similar distributed implementations. Thus, Spark SQL can replace existing practices to retrieve genomic data and, as we show, allow users to reduce the number of lines of software code that needs to be developed to query such data by an order of magnitude. Christos Kozanitis, David A. Patterson 0001 |
BMC Bioinform. | 2 |
| 2015 | DIABLO: A Warehouse-Scale Computer Network Simulator using FPGAsabstractMotivated by rapid software and hardware innovation in warehouse-scale computing (WSC), we visit the problem of warehouse-scale network design evaluation. A WSC is composed of about 30 arrays or clusters, each of which contains about 3000 servers, leading to a total of about 100,000 servers per WSC. We found many prior experiments have been conducted on relatively small physical testbeds, and they often assume the workload is static and that computations are only loosely coupled with the adaptive networking stack. We present a novel and cost-efficient FPGAbased evaluation methodology, called Datacenter-In-A-Box at LOw cost (DIABLO), which treats arrays as whole computers with tightly integrated hardware and software. We have built a 3,000-node prototype running the full WSC software stack. Using our prototype, we have successfully reproduced a few WSC phenomena, such as TCP Incast and memcached request latency long tail, and found that results do indeed change with both scale and with version of the full software stack. Zhangxi Tan, Zhenghao Qian, Krste Asanovic, David A. Patterson 0001 |
ASPLOS | 5 |
| 2015 | Scientific computing meets big data technology: An astronomy use caseabstractScientific analyses commonly compose multiple single-process programs into a dataflow. An end-to-end dataflow of single-process programs is known as a many-task application. Typically, tools from the HPC software stack are used to parallelize these analyses. In this work, we investigate an alternate approach that uses Apache Spark - a modern big data platform - to parallelize many-task applications. We present Kira, a flexible and distributed astronomy image processing toolkit using Apache Spark. We then use the Kira toolkit to implement a Source Extractor application for astronomy images, called Kira SE. With Kira SE as the use case, we study the programming flexibility, dataflow richness, scheduling capacity and performance of Apache Spark running on the EC2 cloud. By exploiting data locality, Kira SE achieves a 3.7 χ speedup over an equivalent C program when analyzing a 1TB dataset using 512 cores on the Amazon EC2 cloud. Furthermore, we show that by leveraging software originally designed for big data infrastructure, Kira SE achieves competitive performance to the C implementation running on the NERSC Edison supercomputer. Our experience with Kira indicates that emerging Big Data platforms such as Apache Spark are a performant alternative for many-task scientific applications. Zhao Zhang 0007, Kyle Barbary, Frank A. Nothaft, Evan Randall Sparks, Oliver Zahn, Michael J. Franklin, David A. Patterson 0001, Saul Perlmutter |
IEEE BigData | 7 |
| 2015 | Rethinking Data-Intensive Science Using Scalable Analytics Systemsabstract"Next generation" data acquisition technologies are allowing scientists to collect exponentially more data at a lower cost. These trends are broadly impacting many scientific fields, including genomics, astronomy, and neuroscience. We can attack the problem caused by exponential data growth by applying horizontally scalable techniques from current analytics systems to accelerate scientific processing pipelines. Frank A. Nothaft, Matt Massie, Timothy Danford, Zhao Zhang 0007, Uri Laserson, Carl Yeksigian, Jey Kottalam, Arun Ahuja, Jeff Hammerbacher, Michael D. Linderman, Michael J. Franklin, Anthony D. Joseph, David A. Patterson 0001 |
SIGMOD Conference | 13 |
| 2015 | The NIH BD2K center for big data in translational genomicsabstractThe world's genomics data will never be stored in a single repository - rather, it will be distributed among many sites in many countries. No one site will have enough data to explain genotype to phenotype relationships in rare diseases; therefore, sites must share data. To accomplish this, the genetics community must forge common standards and protocols to make sharing and computing data among many sites a seamless activity. Through the Global Alliance for Genomics and Health, we are pioneering the development of shared application programming interfaces (APIs) to connect the world's genome repositories. In parallel, we are developing an open source software stack (ADAM) that uses these APIs. This combination will create a cohesive genome informatics ecosystem. Using containers, we are facilitating the deployment of this software in a diverse array of environments. Through benchmarking efforts and big data driver projects, we are ensuring ADAM's performance and utility. Benedict Paten, Mark Diekhans, Brian J. Druker, Stephen H. Friend, Justin Guinney, Nadine Gassner, Mitchell Guttman, W. James Kent, Patrick Mantey, Adam A. Margolin, Matt Massie, Adam M. Novak, Frank A. Nothaft, Lior Pachter, David A. Patterson 0001, Maciej Smuga-Otto, Joshua M. Stuart, Laura J. van't Veer, Barbara J. Wold, David Haussler |
J. Am. Medical Informatics Assoc. | 15 |
| 2014 | Changepoint Analysis for Efficient Variant Calling
Adam E. Bloniarz, Ameet Talwalkar, Jonathan Terhorst, Michael I. Jordan, David A. Patterson 0001, Bin Yu 0001, Yun S. Song |
RECOMB | 5 |
| 2014 | SMaSH: a benchmarking toolkit for human genome variant callingabstractMOTIVATION: Computational methods are essential to extract actionable information from raw sequencing data, and to thus fulfill the promise of next-generation sequencing technology. Unfortunately, computational tools developed to call variants from human sequencing data disagree on many of their predictions, and current methods to evaluate accuracy and computational performance are ad hoc and incomplete. Agreement on benchmarking variant calling methods would stimulate development of genomic processing tools and facilitate communication among researchers. RESULTS: We propose SMaSH, a benchmarking methodology for evaluating germline variant calling algorithms. We generate synthetic datasets, organize and interpret a wide range of existing benchmarking data for real genomes and propose a set of accuracy and computational performance metrics for evaluating variant calling methods on these benchmarking data. Moreover, we illustrate the utility of SMaSH to evaluate the performance of some leading single-nucleotide polymorphism, indel and structural variant calling algorithms. AVAILABILITY AND IMPLEMENTATION: We provide free and open access online to the SMaSH tool kit, along with detailed documentation, at smash.cs.berkeley.edu Ameet Talwalkar, Jesse Liptrap, Julie Newcomb, Christopher Hartl, Jonathan Terhorst, Kristal Curtis, Ma'ayan Bresler, Yun S. Song, Michael I. Jordan, David A. Patterson 0001 |
Bioinform. | 10 |
| 2013 | The RISC-V instruction set
Andrew Waterman, Yunsup Lee, Rimas Avizienis, Henry Cook, David A. Patterson 0001, Krste Asanovic |
Hot Chips Symposium | 5 |
| 2013 | A hardware evaluation of cache partitioning to improve utilization and energy-efficiency while preserving responsivenessabstractComputing workloads often contain a mix of interactive, latency-sensitive foreground applications and recurring background computations. To guarantee responsiveness, interactive and batch applications are often run on disjoint sets of resources, but this incurs additional energy, power, and capital costs. In this paper, we evaluate the potential of hardware cache partitioning mechanisms and policies to improve efficiency by allowing background applications to run simultaneously with interactive foreground applications, while avoiding degradation in interactive responsiveness. We evaluate these tradeoffs using commercial x86 multicore hardware that supports cache partitioning, and find that real hardware measurements with full applications provide different observations than past simulation-based evaluations. Co-scheduling applications without LLC partitioning leads to a 10% energy improvement and average throughput improvement of 54% compared to running tasks separately, but can result in foreground performance degradation of up to 34% with an average of 6%. With optimal static LLC partitioning, the average energy improvement increases to 12% and the average throughput improvement to 60%, while the worst case slowdown is reduced noticeably to 7% with an average slowdown of only 2%. We also evaluate a practical low-overhead dynamic algorithm to control partition sizes, and are able to realize the potential performance guarantees of the optimal static approach, while increasing background throughput by an additional 19%. Henry Cook, Miquel Moretó, Sarah Bird, Khanh Dao, David A. Patterson 0001, Krste Asanovic |
ISCA | 5 |
| 2013 | Generalized scale independence through incremental precomputationabstractDevelopers of rapidly growing applications must be able to anticipate potential scalability problems before they cause performance issues in production environments. A new type of data independence, called scale independence, seeks to address this challenge by guaranteeing a bounded amount of work is required to execute all queries in an application, independent of the size of the underlying data. While optimization strategies have been developed to provide these guarantees for the class of queries that are scale-independent when executed using simple indexes, there are important queries for which such techniques are insufficient. Michael Armbrust, Eric Liang, Tim Kraska, Armando Fox, Michael J. Franklin, David A. Patterson 0001 |
SIGMOD Conference | 6 |
| 2013 | Using clouds for MapReduce measurement assignmentsabstractWe describe our experiences teaching MapReduce in a large undergraduate lecture course using public cloud services and the standard Hadoop API. Using the standard API, students directly experienced the quality of industrial big-data tools. Using the cloud, every student could carry out scalability benchmarking assignments on realistic hardware, which would have been impossible otherwise. Over two semesters, over 500 students took our course. We believe this is the first large-scale demonstration that it is feasible to use pay-as-you-go billing in the cloud for a large undergraduate course. Modest instructor effort was sufficient to prevent students from overspending. Average per-pupil expenses in the Cloud were under $45. Students were excited by the assignment: 90% said they thought it should be retained in future course offerings. Ariel Rabkin, Charles Reiss, Randy H. Katz, David A. Patterson 0001 |
ACM Trans. Comput. Educ. | 4 |
| 2012 | Direction-optimizing breadth-first searchabstractBreadth-First Search is an important kernel used by many graph-processing applications. In many of these emerging applications of BFS, such as analyzing social networks, the input graphs are low-diameter and scale-free. We propose a hybrid approach that is advantageous for low-diameter graphs, which combines a conventional top-down algorithm along with a novel bottom-up algorithm. The bottom-up algorithm can dramatically reduce the number of edges examined, which in turn accelerates the search as a whole. On a multi-socket server, our hybrid approach demonstrates speedups of 3.3 -- 7.8 on a range of standard synthetic graphs and speedups of 2.4 -- 4.6 on graphs from real social networks when compared to a strong baseline. We also typically double the performance of prior leading shared memory (multicore and GPU) implementations. Scott Beamer, Krste Asanovic, David A. Patterson 0001 |
SC | 3 |
| 2012 | Experiences teaching MapReduce in the cloudabstractWe describe our experiences teaching MapReduce in a large undergraduate lecture course using public cloud services. Using the cloud, every student could carry out scalability benchmarking assignments on realistic hardware, which would have been impossible otherwise. Over two semesters, over 500 students took our course. We believe this is the first large-scale demonstration that it is feasible to use pay-as-you-go billing in the Cloud for a large undergraduate course. Modest instructor effort was sufficient to prevent students from overspending. Average per-pupil expenses in the Cloud were under $45, less than half our available grant funding. Students were excited by the assignment: 90% said they thought it should be retained in future course offerings. Ariel Rabkin, Charles Reiss, Randy H. Katz, David A. Patterson 0001 |
SIGCSE | 4 |
| 2011 | The SCADS Director: Scaling a Distributed Storage System Under Stringent Performance Requirements
Beth Trushkowsky, Peter Bodík, Armando Fox, Michael J. Franklin, Michael I. Jordan, David A. Patterson 0001 |
FAST | 6 |
| 2011 | PIQL: Success-Tolerant Query Processing in the CloudabstractNewly-released web applications often succumb to a "Success Disaster," where overloaded database machines and resulting high response times destroy a previously good user experience. Unfortunately, the data independence provided by a traditional relational database system, while useful for agile development, only exacerbates the problem by hiding potentially expensive queries under simple declarative expressions. As a result, developers of these applications are increasingly abandoning relational databases in favor of imperative code written against distributed key/value stores, losing the many benefits of data independence in the process. Instead, we propose PIQL, a declarative language that also provides scale independence by calculating an upper bound on the number of key/value store operations that will be performed for any query. Coupled with a service level objective (SLO) compliance prediction model and PIQL's scalable database architecture, these bounds make it easy for developers to write success-tolerant applications that support an arbitrarily large number of users while still providing acceptable performance. In this paper, we present the PIQL query processing system and evaluate its scale independence on hundreds of machines using two benchmarks, TPC-W and SCADr. Michael Armbrust, Kristal Curtis, Tim Kraska, Armando Fox, Michael J. Franklin, David A. Patterson 0001 |
Proc. VLDB Endow. | 6 |
| 2010 | The case for PIQL: a performance insightful query languageabstractLarge-scale, user-facing applications are increasingly moving from relational databases to distributed key/value stores for high-request-rate, low-latency workloads. Often, this move is motivated not only by key/value stores' ability to scale simply by adding more hardware, but also by the easy to understand predictable performance they provide for all operations. For complex queries, this approach often requires onerous explicit index management and imperative data lookup by the developer. We propose PIQL, a Performance Insightful Query Language that allows developers to express many queries found on these websites while still providing strict bounds on the number of I/O operations that will be performed. Michael Armbrust, Nick Lanham, Stephen Tu, Armando Fox, Michael J. Franklin, David A. Patterson 0001 |
SoCC | 6 |
| 2010 | Characterizing, modeling, and generating workload spikes for stateful servicesabstractEvaluating the resiliency of stateful Internet services to significant workload spikes and data hotspots requires realistic workload traces that are usually very difficult to obtain. A popular approach is to create a workload model and generate synthetic workload, however, there exists no characterization and model of stateful spikes. In this paper we analyze five workload and data spikes and find that they vary significantly in many important aspects such as steepness, magnitude, duration, and spatial locality. We propose and validate a model of stateful spikes that allows us to synthesize volume and data spikes and could thus be used by both cloud computing users and providers to stress-test their infrastructure. Peter Bodík, Armando Fox, Michael J. Franklin, Michael I. Jordan, David A. Patterson 0001 |
SoCC | 5 |
| 2010 | RAMP gold: an FPGA-based architecture simulator for multiprocessorsabstractWe present RAMP Gold, an economical FPGA-based architecture simulator that allows rapid early design-space exploration of manycore systems. The RAMP Gold prototype is a high-throughput, cycle-accurate full-system simulator that runs on a single Xilinx Virtex-5 FPGA board, and which simulates a 64-core shared-memory target machine capable of booting real operating systems. To improve FPGA implementation efficiency, functionality and timing are modeled separately and host multithreading is used in both models. We evaluate the prototype's performance using a modern parallel benchmark suite running on our manycore research operating system, achieving two orders of magnitude speedup compared to a widely-used software-based architecture simulator. Zhangxi Tan, Andrew Waterman, Rimas Avizienis, Yunsup Lee, Henry Cook, David A. Patterson 0001, Krste Asanovic |
DAC | 6 |
| 2010 | Detecting Large-Scale System Problems by Mining Console Logs
Wei Xu 0012, Ling Huang 0001, Armando Fox, David A. Patterson 0001, Michael I. Jordan |
ICML | 4 |
| 2010 | A case for FAME: FPGA architecture model executionabstractGiven the multicore microprocessor revolution, we argue that the architecture research community needs a dramatic increase in simulation capacity. We believe FPGA Architecture Model Execution (FAME) simulators can increase the number of useful architecture research experiments per day by two orders of magnitude over Software Architecture Model Execution (SAME) simulators. To clear up misconceptions about FPGA-based simulation methodologies, we propose a FAME taxonomy to distinguish the costperformance of variations on these ideas. We demonstrate our simulation speedup claim with a case study wherein we employ a prototype FAME simulator, RAMP Gold, to research the interaction between hardware partitioning mechanisms and operating system scheduling policy. The study demonstrates FAME's capabilities: we run a modern parallel benchmark suite on a research operating system, simulate 64-core target architectures with multi-level memory hierarchy timing models, and add experimental hardware mechanisms to the target machine. The simulation speedup achieved by our adoption of FAME-250×-enables experiments with more realistic time scales and data set sizes thanare possible with SAME. Zhangxi Tan, Andrew Waterman, Henry Cook, Sarah Bird, Krste Asanovic, David A. Patterson 0001 |
ISCA | 6 |
| 2010 | PIQL: a performance insightful query languageabstractLarge-scale websites are increasingly moving from relational databases to distributed key-value stores for high request rate, low latency workloads. Often this move is motivated not only by key-value stores' ability to scale simply by adding more hardware, but also by the easy to understand predictable performance they provide for all operations. While this data model works well, lookups are only done by primary key. More complex queries require onerous, explicit index management and imperative data lookups by the developer. We demonstrate PIQL, a Performance Insightful Query Language that allows developers to express many of the queries found on these websites, while still providing strict bounds on the number of I/O operations for any query. Michael Armbrust, Stephen Tu, Armando Fox, Michael J. Franklin, David A. Patterson 0001, Nick Lanham, Beth Trushkowsky, Jesse Trutna |
SIGMOD Conference | 5 |
| 2009 | SCADS: Scale-Independent Storage for Social Computing Applications
Michael Armbrust, Armando Fox, David A. Patterson 0001, Nick Lanham, Beth Trushkowsky, Jesse Trutna, Haruki Oh |
CIDR | 3 |
| 2009 | Predicting Multiple Metrics for Queries: Better Decisions Enabled by Machine LearningabstractOne of the most challenging aspects of managing a very large data warehouse is identifying how queries will behave before they start executing. Yet knowing their performance characteristics - their runtimes and resource usage - can solve two important problems. First, every database vendor struggles with managing unexpectedly long-running queries. When these long-running queries can be identified before they start, they can be rejected or scheduled when they will not cause extreme resource contention for the other queries in the system. Second, deciding whether a system can complete a given workload in a given time period (or a bigger system is necessary) depends on knowing the resource requirements of the queries in that workload. We have developed a system that uses machine learning to accurately predict the performance metrics of database queries whose execution times range from milliseconds to hours. For training and testing our system, we used both real customer queries and queries generated from an extended set of TPC-DS templates. The extensions mimic queries that caused customer problems. We used these queries to compare how accurately different techniques predict metrics such as elapsed time, records used, disk I/Os, and message bytes. The most promising technique was not only the most accurate, but also predicted these metrics simultaneously and using only information available prior to query execution. We validated the accuracy of this machine learning technique on a number of HP Neoview configurations. We were able to predict individual query elapsed time within 20% of its actual time for 85% of the test queries. Most importantly, we were able to correctly identify both the short and long-running (up to two hour) queries to inform workload management and capacity planning. Archana Ganapathi, Harumi A. Kuno, Umeshwar Dayal, Janet L. Wiener, Armando Fox, Michael I. Jordan, David A. Patterson 0001 |
ICDE | 7 |
| 2009 | Online System Problem Detection by Mining Patterns of Console LogsabstractWe describe a novel application of using data mining and statistical learning methods to automatically monitor and detect abnormal execution traces from console logs in an online setting. Different from existing solutions, we use a two stage detection system. The first stage uses frequent pattern mining and distribution estimation techniques to capture the dominant patterns (both frequent sequences and time duration). The second stage use principal component analysis based anomaly detection technique to identify actual problems. Using real system data from a 203-node Hadoop cluster, we show that we can not only achieve highly accurate and fast problem detection, but also help operators better understand execution patterns in their system. Wei Xu 0012, Ling Huang 0001, Armando Fox, David A. Patterson 0001, Michael I. Jordan |
ICDM | 4 |
| 2009 | Detecting large-scale system problems by mining console logsabstractSurprisingly, console logs rarely help operators detect problems in large-scale datacenter services, for they often consist of the voluminous intermixing of messages from many software components written by independent developers. We propose a general methodology to mine this rich source of information to automatically detect system runtime problems. We first parse console logs by combining source code analysis with information retrieval to create composite features. We then analyze these features using machine learning to detect operational problems. We show that our method enables analyses that are impossible with previous methods because of its superior ability to create sophisticated features. We also show how to distill the results of our analysis to an operator-friendly one-page decision tree showing the critical messages associated with the detected problems. We validate our approach using the Darkstar online game server and the Hadoop File System, where we detect numerous real problems with high accuracy and few false positives. In the Hadoop case, we are able to analyze 24 million lines of console logs in 3 minutes. Our methodology works on textual console logs of any size and requires no changes to the service software, no human input, and no knowledge of the software's internals. Wei Xu 0012, Ling Huang 0001, Armando Fox, David A. Patterson 0001, Michael I. Jordan |
SOSP | 4 |
| 2008 | Stencil computation optimization and auto-tuning on state-of-the-art multicore architecturesabstractUnderstanding the most efficient design and utilization of emerging multicore systems is one of the most challenging questions faced by the mainstream and scientific computing industries in several decades. Our work explores multicore stencil (nearest-neighbor) computations - a class of algorithms at the heart of many structured grid codes, including PDE solvers. We develop a number of effective optimization strategies, and build an auto-tuning environment that searches over our optimizations and their parameters to minimize runtime, while maximizing performance portability. To evaluate the effectiveness of these strategies we explore the broadest set of multicore architectures in the current HPC literature, including the Intel Clovertown, AMD Barcelona, Sun Victoria Falls, IBM QS22 PowerXCell 8i, and NVIDIA GTX280. Overall, our auto-tuning optimization methodology results in the fastest multicore stencil performance to date. Finally, we present several key insights into the architectural tradeoffs of emerging multicore designs and their implications on scientific algorithm development. Kaushik Datta, Mark Murphy, Vasily Volkov, Samuel Williams 0001, Jonathan Carter 0002, Leonid Oliker, David A. Patterson 0001, John Shalf, Katherine A. Yelick |
SC | 7 |
| 2008 | Design and implementation trade-offs for wide-area resource discoveryabstractWe describe the design and implementation of SWORD, a scalable resource discovery service for wide-area distributed systems. In contrast to previous systems, SWORD allows users to describe desired resources as a topology of interconnected groups with required intragroup, intergroup, and per-node characteristics, along with the utility that the application derives from specified ranges of metric values. This design gives users the flexibility to find geographically distributed resources for applications that are sensitive to both node and network characteristics, and allows the system to rank acceptable configurations based on their quality for that application. Rather than evaluating a single implementation of SWORD, we explore a variety of architectural designs that deliver the required functionality in a scalable and highly available manner. We discuss the trade-offs of using a centralized architecture as compared to a fully decentralized design to perform wide-area resource discovery. To summarize our results, we found that a centralized architecture based on 4-node server cluster sites at network-peering facilities outperforms a decentralized DHT-based resource discovery infrastructure with respect to query latency for all but the smallest number of sites. However, although a centralized architecture shows significant promise in stable environments, we find that our decentralized implementation has acceptable performance and also benefits from the DHT's self-healing properties in more volatile environments. We evaluate the advantages and disadvantages of centralized and distributed resource discovery architectures on 1000 hosts in emulation and on approximately 200 PlanetLab nodes spread across the Internet. Jeannie R. Albrecht, David Oppenheimer, Amin Vahdat, David A. Patterson 0001 |
ACM Trans. Internet Techn. | 4 |
| 2007 | The parallel computing landscape: a Berkeley viewabstractcontributed articles doi:10.1145/1562764.1562783 Writing programs that scale with increasing numbers of cores should be as easy as writing programs for sequential computers. David A. Patterson 0001 |
ISLPED | 1 |
| 2006 | Research accelerator for multiple processors
David A. Patterson 0001, Arvind 0001, Krste Asanovic, Derek Chiou, James C. Hoe, Christoforos E. Kozyrakis, Shih-Lien Lu, Mark Oskin, Jan M. Rabaey, John Wawrzynek |
Hot Chips Symposium | 1 |
| 2006 | RAMP: research accelerator for multiple processors - a community vision for a shared experimental parallel HW/SW platformabstractSummary form only given. The vast majority of computer architects believe the future of the microprocessor is hundreds to thousands of processors ("cores") on a chip. Given such widespread agreement, it's surprising how much research remains to be done in algorithms, computer architecture, networks, operating systems, file systems, compilers, programming languages, applications, and so on to realize this vision. Fortunately, Moore's law has not only enabled dense multi-core chips, it has also enabled extremely dense FPGAs. Today, one to two dozen soft cores can be programmed into a single FPGA. With multiple FPGAs on a board and multiple boards in a system, 1000-processor designs can be economically and rapidly explored. To make this happen, however, requires a significant amount of infrastructure in hardware, software, and what we call "gateware", the register-transfer level models that fill the FGPAs. By using the Berkeley Emulation Engine boards that were created for other purposes, the hardware is already done. A group of architects plan to design the gateware, create this infrastructure, and share the results in an open-source fashion so that every institution could have their own. Such a system would not just invigorate multiprocessors research in the architecture community. Since processors cores can run at 100 to 200 MHz, a large scale multiprocessor would be fast enough to run operating systems and large programs at speeds sufficient to support software research. Moreover, there is a new generation of FPGAs every 18 months with capacity for twice as many cores and run them faster, so future multiboard FPGA systems are even more attractive. Hence, we believe such a system would accelerate research across all the fields that touch multiple processors. Thus the acronynm RAMP, for Research Accelerator for Multiple Processors. RAMP has the potential to transform the parallel computing community in computer science from a simulation-driven to a prototype-driven discipline, leading to rapid iteration across interfaces of the many fields of multiple processors, and thereby moving much more quickly to a parallel foundation for large-scale computer systems research in the 21st century. David A. Patterson 0001 |
ISPASS | 1 |
| 2006 | Windows XP Kernel Crash Analysis
Archana Ganapathi, Viji Ganapathi, David A. Patterson 0001 |
LISA | 3 |
| 2006 | Service Placement in a Shared Wide-Area Platform
David Oppenheimer, Brent N. Chun, David A. Patterson 0001, Alex C. Snoeren, Amin Vahdat |
USENIX ATC, General Track | 3 |
| 2005 | Crash Data Collection: A Windows Case StudyabstractReliability is a rapidly growing concern in contemporary personal computer (PC) industry, both for computer users as well as product developers. To improve dependability, systems designers and programmers must consider failure and usage data for operating systems as well as applications. In this paper, we discuss our experience with crash and usage data collection for Windows machines. We analyze results based on crashes in the UC Berkeley EECS department. Archana Ganapathi, David A. Patterson 0001 |
DSN | 2 |
| 2005 | Design and implementation tradeoffs for wide-area resource discoveryabstractThis paper describes the design and implementation of SWORD, a scalable resource discovery service for wide-area distributed systems. In contrast to previous systems, SWORD allows users to describe desired resources as a topology of interconnected groups with required intragroup, intergroup, and per-node characteristics, along with the utility that the application derives from various ranges of values of those characteristics. This design gives users the flexibility to find geographically distributed resources for applications that are sensitive to both node and network characteristics, and allows the system to rank acceptable configurations based on their quality for that application. We explore a variety of architectures to deliver SWORD's functionality in a scalable and highly-available manner. A 1000-node ModelNet evaluation using a workload of measurements collected from PlanetLab shows that an architecture based on 4-node server cluster sites at network peering facilities outperforms a decentralized DHT-based resource discovery infrastructure for all but the smallest number of sites. While such a centralized architecture shows significant promise, we find that our decentralized implementation, both in emulation and running continuously on over 200 PlanetLab nodes, performs well while benefiting from the DHT's self-healing properties. David Oppenheimer, Jeannie R. Albrecht, David A. Patterson 0001, Amin Vahdat |
HPDC | 3 |
| 2005 | Latency Lags BandwidthabstractSummary form only given. As I review performance trends, I am struck by a consistent theme across many technologies over many years: bandwidth improves much more quickly than latency for four different technologies: disks, networks, memories and processors. A rule of thumb to quantify the imbalance is: bandwidth improves by more than the square of the improvement in latency. This paper lists a half-dozen performance milestones to document this observation, many reasons why it happens, a few ways to cope with it, and two small examples of how you might design systems differently if you kept this simple rule of thumb in mind. David A. Patterson 0001 |
ICCD | 1 |
| 2005 | Service placement in shared wide-area platformsabstractFederated geographically-distributed computing platforms such as PlanetLab [1] and the Grid [2, 3] have recently become popular for evaluating and deploying network services and scientific computations. As the size, reach, and user population of such infrastructures grow, resource discovery and resource selection become increasingly important. Although a number of resource discovery and allocation services have been built, there is little data on the utilization of the distributed computing platforms they target. Yet the design and efficacy of such services depends on the characteristics of the target platform. David Oppenheimer, Brent N. Chun, David A. Patterson 0001, Alex C. Snoeren, Amin Vahdat |
SOSP | 3 |
| 2004 | Experience with Evaluating Human-Assisted Recovery ProcessesabstractWe describe an approach to quantitatively evaluating human-assisted failure-recovery tools and processes in the environment of modern Internetand enterprise-class server systems. Our approach can quantify the dependability impact of a single recovery system, and also enables comparisons between different recovery approaches. The approach combines aspects of dependability benchmarking with human user studies, incorporating human participants in the system evaluations yet still producing typical dependability-related metrics as results. We illustrate our methodology via a case study of a system-wide undo/redo recovery tool for e-mail services; our approach is able to expose the dependability benefits of the tool as well as point out areas where its behavior could use improvement. Aaron B. Brown, Leonard Chung, William Kakes, Calvin Ling, David A. Patterson 0001 |
DSN | 5 |
| 2004 | Path-Based Failure and Evolution Management
Mike Y. Chen, Anthony J. Accardi, Emre Kiciman, David A. Patterson 0001, Armando Fox, Eric A. Brewer |
NSDI | 4 |
| 2003 | Overcoming the Limitations of Conventional Vector ProcessorsabstractDespite their superior performance for multimedia applications, vector processors have three limitations that hinder their widespread acceptance. First, the complexity and size of the centralized vector register file limits the number of functional units. Second, precise exceptions for vector instructions are difficult to implement. Third, vector processors require an expensive on-chip memory system that supports high bandwidth at low access latency.This paper introduces CODE, a scalable vector microarchitecture that addresses these three shortcomings. It is designed around a clustered vector register file and uses a separate network for operand transfers across functional units. With extensive use of decoupling, it can hide the latency of communication across functional units and provides 26% performance improvement over a centralized organization. CODE scales efficiently to 8 functional units without requiring wide instruction issue capabilities. A renaming table makes the clustered register file transparent at the instruction set level. Renaming also enables precise exceptions for vector instructions at a performance loss of less than 5%. Finally, decoupling allows CODE to tolerate large increases in memory latency at sub-linear performance degradation without using on-chip caches. Thus, CODE can use economical, off-chip, memory systems. Christoforos E. Kozyrakis, David A. Patterson 0001 |
ISCA | 2 |
| 2003 | Undo for Operators: Building an Undoable E-mail Store
Aaron B. Brown, David A. Patterson 0001 |
USENIX ATC, General Track | 2 |
| 2002 | Recovery Oriented Computing: A New Research Agenda for a New CenturyabstractSummary form only given, as follows. After 15 years of successfully improving cost-performance, it's time for new challenges for the systems research community. As a result of the focus on cost-performance, the fabled five 9s of availability (99.999% uptime) looks to be much easier to achieve in advertising than in computers, and the cost of managing systems can be five times the cost of the hardware. In a Post-PC Era of wireless gadgets using services on the Internet, one new challenge is building services that really are dependable and much less expensive to maintain. Traditional Fault-Tolerant Computing concentrates on tolerating hardware and operating system faults, ignoring faults by human operators and even applications. Recovery Oriented Computing (ROC) aims at improving Mean Time To Recover to both lower the cost of management and improve at the availability of whole system, including the people who operate it. We look to civil engineering and diplomacy to inspire principles for ROC design. This talk outlines motivation for and proposed principles of ROC design, plus some concrete results in the area of benchmarking of availability. David A. Patterson 0001 |
HPCA | 1 |
| 2002 | A Simple Way to Estimate the Cost of Downtime
David A. Patterson 0001 |
LISA | 1 |
| 2002 | Vector vs. superscalar and VLIW architectures for embedded multimedia benchmarksabstractMultimedia processing on embedded devices requires an architecture that leads to high performance, low power consumption, reduced design complexity, and small code size. In this paper, we use EEMBC, an industrial benchmark suite, to compare the VIRAM vector architecture to superscalar and VLIW processors for embedded multimedia applications. The comparison covers the VIRAM instruction set, vectorizing compiler and the prototype chip that integrates a vector processor with DRAM main memory. We demonstrate that executable code for VIRAM is up to 10 times smaller than VLIW code and comparable to /spl times/86 CISC code. The simple, cache-less VIRAM chip is 2 times faster than a 4-way superscalar RISC processor that uses a 5 times faster clock frequency and consumes 10 times more power VIRAM is also 10 times faster than cache-based VLIW processors. Even after manual optimization of the VLIW code and insertion of SIMD and DSP instructions, the single-issue VIRAM processor is 60% faster than 5-way to 8-way VLIW designs. Christoforos E. Kozyrakis, David A. Patterson 0001 |
MICRO | 2 |
| 2002 | ROC-1: Hardware Support for Recovery-Oriented ComputingabstractWe introduce the ROC-1 hardware platform, a large-scale cluster system designed to provide high availability for Internet service applications. The ROC-1 prototype embodies our philosophy of recovery-oriented computing (ROC) by emphasizing detection and recovery from the failures that inevitably occur in Internet service environments, rather than simple avoidance of such failures. ROC-1 promises greater availability than existing server systems by incorporating four techniques applied from the ground up to both hardware and software: redundancy and isolation, online self-testing and verification, support for problem diagnosis and concern for human interaction with the system. David Oppenheimer, Aaron B. Brown, James Beck, Daniel Hettena, Jon Kuroda, Noah Treuhaft, David A. Patterson 0001, Katherine A. Yelick |
IEEE Trans. Computers | 7 |
| 2001 | Hardware/compiler codevelopment for an embedded media processorabstractEmbedded and portable systems running multimedia applications create a new challenge for hardware architects. A microprocessor for such applications needs to be easy to program like a general-purpose processor and have the performance and power efficiency of a digital signal processor. This paper presents the codevelopment of the instruction set, the hardware, and the compiler for the Vector IRAM media processor. A vector architecture is used to exploit the data parallelism of multimedia programs, which allows the use of highly modular hardware and enables implementations that combine high performance, low power consumption, and reduced design complexity. It also leads to a compiler model that is efficient both in terms of performance and executable code size. The memory system for the vector processor is implemented using embedded DRAM technology, which provides high bandwidth in an integrated, cost-effective manner. The hardware and the compiler for this architecture make complementary contributions to the efficiency of the overall system. This paper explores the interactions and tradeoffs between them, as well as the enhancements to a vector architecture necessary for multimedia processing. We also describe how the architecture, design, and compiler features come together in a prototype system-on-a-chip, able to execute 3.2 billion operations per second per watt. Christoforos E. Kozyrakis, David Judd, Joseph Gebis, Samuel Williams 0001, David A. Patterson 0001, Katherine A. Yelick |
Proc. IEEE | 5 |
| 2000 | Towards Availability Benchmarks: A Case Study of Software RAID Systems
Aaron B. Brown, David A. Patterson 0001 |
USENIX ATC, General Track | 2 |
| 1999 | A Retrospective on Twelve Years of LISA Proceedings
David A. Patterson 0001 |
LISA | 2 |
| 1999 | Virtual Log Based File Systems for a Programmable Disk
Randolph Y. Wang, Thomas E. Anderson, David A. Patterson 0001 |
OSDI | 3 |
| 1998 | The Architectural Costs of Streaming I/O: A Comparison of Workstations, Clusters, and SMPsabstractWe investigate resource usage while performing streaming I/O by contrasting three architectures, a single workstation, a cluster, and an SMP, under various I/O benchmarks. We derive analytical and empirically-based models of resource usage during data transfer, examining the I/O bus, memory bus, network, and processor of each system. By investigating each resource in detail, we assess what comprises a well-balanced system for these workloads. We find that the architectures we study are not well balanced for streaming I/O applications. Across the platforms, the main limitation to attaining peak performance is the CPU, due to lack of data locality. Increasing processor performance (especially with improved block operation performance) will be of great aid for these workloads in the future. For a cluster workstation, the I/O bus is a major system bottleneck, because of the increased load placed on it from network communication. A well-balanced cluster workstation should have copious I/O bus bandwidth, perhaps via multiple I/O busses. The SMP suffers from poor memory-system performance; even when there is true parallelism in the benchmark, contention in the shared-memory system leads to reduced performance. As a result, the clustered workstations provide higher absolute performance for streaming I/O workloads. Remzi H. Arpaci-Dusseau, Andrea C. Arpaci-Dusseau, David E. Culler, Joseph M. Hellerstein, David A. Patterson 0001 |
HPCA | 5 |
| 1998 | Performance Characterization of a Quad Pentium Pro SMP using OLTP WorkloadsabstractCommercial applications are an important, yet often overlooked, workload with significantly different characteristics from technical workloads. The potential impact of these differences is that computers optimized for technical workloads may not provide good performance for commercial applications, and these applications may not fully exploit advances in processor design. To evaluate these issues, we use hardware counters to measure architectural features of a four-processor Pentium Pro-based server running a TPC-C-like workload on an Informix database. We examine the effectiveness of out-of-order execution, branch prediction, speculative execution, superscalar issue and retire, caching and multiprocessor scaling. We find that out-of-order execution, superscalar issue and retire, and branch prediction are not as effective for database workloads as they are for technical workloads, such as SPEC. We find that caches are effective at reducing processor traffic to memory; even larger caches would be helpful to satisfy more data requests. Multiprocessor scaling of this workload is good, but even modest bus utilization degrades application memory latency, limiting database throughput. Kimberly Keeton, David A. Patterson 0001, Yong Qiang He, Roger C. Raphael, Walter E. Baker |
ISCA | 2 |
| 1997 | A new voting based hardware data prefetch schemeabstractThe dramatic increase in the processor memory gap in recent years has led to the development of techniques like data prefetching that hide the latency of cache misses. Two such hardware techniques are the stream buffer and the stride predictor. They have dissimilar architectures, are effective for different kinds of memory access patterns and require different amounts of extra memory bandwidth. We compare the performance of these two techniques and propose a scheme that unifies them. Simulation studies on six benchmark programs confirm that the combined scheme is more effective in reducing the average memory access time (AMAT) than either of the two individually. Gurmeet Singh Manku, Mukul R. Prasad, David A. Patterson 0001 |
HiPC | 3 |
| 1997 | Intelligent RAM (IRAM): The Industrial Setting, Applications and ArchitecturesabstractThe goal of intelligent RAM (IRAM) is to design a cost-effective computer by designing a processor in a memory fabrication process, instead of in a conventional logic fabrication process, and include memory on-chip. To design a processor in a DRAM process one must learn about the business and culture of the DRAMs, which is quite different from microprocessors. The authors describe some of those differences and their current vision of IRAM applications, architectures, and implementations. David A. Patterson 0001, Krste Asanovic, Aaron B. Brown, Richard Fromm, Jason Golbus, Benjamin Gribstad, Kimberly Keeton, Christoforos E. Kozyrakis, David R. Martin 0001, Stylianos Perissakis, Randi Thomas, Noah Treuhaft, Katherine A. Yelick |
ICCD | 1 |
| 1997 | The Energy Efficiency of IRAM ArchitecturesabstractPortable systems demand energy efficiency in order to maximize battery life. IRAM architectures, which combine DRAM and a processor on the same chip in a DRAM process, are more energy efficient than conventional systems. The high density of DRAM permits a much larger amount of memory on-chip than a traditional SRAM cache design in a logic process. This allows most or all IRAM memory accesses to be satisfied on-chip. Thus there is much less need to drive high-capacitance off-chip buses, which contribute significantly to the energy consumption of a system. To quantify this advantage we apply models of energy consumption in DRAM and SRAM memories to results from cache simulations of applications reflective of personal productivity tasks on low power systems. We find that IRAM memory hierarchies consume as little as 22% of the energy consumed by a conventional memory hierarchy for memory-intensive applications, while delivering comparable performance. Furthermore, the energy consumed by a system consisting of an IRAM memory hierarchy combined with an energy efficient CPU core is as little as 40% of that of the same CPU core with a traditional memory hierarchy. Richard Fromm, Stylianos Perissakis, Neal Cardwell, Christoforos E. Kozyrakis, Bruce McGaughy, David A. Patterson 0001, Thomas E. Anderson, Katherine A. Yelick |
ISCA | 6 |
| 1997 | Extensible, Scalable Monitoring for Clusters of Computers
David A. Patterson 0001 |
LISA | 2 |
| 1997 | High-Performance Sorting on Networks of WorkstationsabstractWe report the performance of NOW-Sort, a collection of sorting implementations on a Network of Workstations (NOW). We find that parallel sorting on a NOW is competitive to sorting on the large-scale SMPs that have traditionally held the performance records. On a 64-node cluster, we sort 6.0 GB in just under one minute, while a 32-node cluster finishes the Datamation benchmark in 2.41 seconds. Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau, David E. Culler, Joseph M. Hellerstein, David A. Patterson 0001 |
SIGMOD Conference | 5 |
| 1996 | Serverless Network File SystemsabstractWe propose a new paradigm for network file system design:serverless network file systems. While traditional network file systems rely on a central server machine, a serverless system utilizes workstations cooperating as peers to provide all file system services. Any machine in the system can store, cache, or control any block of data. Our approach uses this location independence, in combination with fast local area networks, to provide better performance and scalability than traditional file systems. Furthermore, because any machine in the system can assume the responsibilities of a failed component, our serverless design also provides high availability via redundatn data storage. To demonstrate our approach, we have implemented a prototype serverless network file system called xFS. Preliminary performance measurements suggest that our architecture achieves its goal of scalability. For instance, in a 32-node xFS system with 32 active clients, each client receives nearly as much read or write throughput as it would see if it were the only active client. Thomas E. Anderson, Michael Dahlin, Jeanna Matthews, David A. Patterson 0001, Drew S. Roselli, Randolph Y. Wang |
ACM Trans. Comput. Syst. | 4 |
| 1995 | Choosing the Best Storage System for Video ServiceabstractNo abstract available. Ann L. Chervenak, David A. Patterson 0001, Randy H. Katz |
ACM Multimedia | 2 |
| 1995 | A Case for NOW (Networks of Workstations) - AbstractabstractNo abstract available. David A. Patterson 0001, David E. Culler, Thomas E. Anderson |
PODC | 1 |
| 1995 | The Interaction of Parallel and Sequential Workloads on a Network of WorkstationsabstractThis paper examines the plausibility of using a network of workstations (NOW) for a mixture of parallel and sequential jobs. Through simulations, our study examines issues that arise when combining these two workloads on a single platform. Starting from a dedicated NOW just for parallel programs, we incrementally relax uniprogramming restrictions until we have a multi-programmed, multi-user NOW for both interactive sequential users and parallel programs. We show that a number of issues associated with the distributed NOW environment (e.g., daemon activity, coscheduling skew) can have a small but noticeable effect on parallel program performance. We also find that efficient migration to idle workstations is necessary to maintain acceptable parallel application performance. Furthermore, we present a methodology for deriving an optimal delay time for recruiting idle machines for use by parallel programs; this recruitment threshold was just 3 minutes for the research cluster we measured. Finally, we quantify the effects of the additional parallel load upon interactive users by keeping track of the potential number of user delays in our simulations. When we limit the maximum number of delays per user, we can still maintain acceptable parallel program performance. In summary, we find that for our workloads a 2:1 rule applies: a NOW cluster of approximately 60 machines can sustain a 32-node parallel workload in addition to the sequential load placed upon it by interactive users. Remzi H. Arpaci-Dusseau, Andrea C. Arpaci-Dusseau, Amin Vahdat, Lok T. Liu, Thomas E. Anderson, David A. Patterson 0001 |
SIGMETRICS | 6 |
| 1995 | Serverless Network File SystemsabstractArticle Serverless network file systems Share on Authors: T. E. Anderson Computer Science Division, University of California at Berkeley Computer Science Division, University of California at BerkeleyView Profile , M. D. Dahlin Computer Science Division, University of California at Berkeley Computer Science Division, University of California at BerkeleyView Profile , J. M. Neefe Computer Science Division, University of California at Berkeley Computer Science Division, University of California at BerkeleyView Profile , D. A. Patterson Computer Science Division, University of California at Berkeley Computer Science Division, University of California at BerkeleyView Profile , D. S. Roselli Computer Science Division, University of California at Berkeley Computer Science Division, University of California at BerkeleyView Profile , R. Y. Wang Computer Science Division, University of California at Berkeley Computer Science Division, University of California at BerkeleyView Profile Authors Info & Claims SOSP '95: Proceedings of the fifteenth ACM symposium on Operating systems principlesDecember 1995 Pages 109–126https://doi.org/10.1145/224056.224066Online:03 December 1995Publication History 255citation3,156DownloadsMetricsTotal Citations255Total Downloads3,156Last 12 Months71Last 6 weeks6 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Thomas E. Anderson, Michael Dahlin, Jeanna Matthews, David A. Patterson 0001, Drew S. Roselli, Randolph Y. Wang |
SOSP | 4 |
| 1994 | RAID-II: A High-Bandwidth Network File ServerabstractIn 1989, the RAID (Redundant Arrays of Inexpensive Disks) group at U.C. Berkeley built a prototype disk array called RAID-I. The bandwidth delivered to clients by RAID-I was severely limited by the memory system bandwidth of the disk array's host workstation. They designed their second prototype, RAID-II, to deliver more of the disk array bandwidth to file server clients. A custom-built crossbar memory system called the XBUS board connects the disks directly to the high-speed network, allowing data for large requests to bypass the server workstation. RAID-II runs Log-Structured File System (LFS) software to optimize performance for bandwidth-intensive applications. The RAID-II hardware with a single XBUS controller board delivers 20 megabytes/second for large, random read operations and up to 31 megabytes/second for sequential read operations. A preliminary implementation of LFS on RAID-II delivers 21 megabytes/second on large read requests and 15 megabytes/second on large write operations.> Ann L. Drapeau, Ken Shirriff, John H. Hartman, Ethan L. Miller, Srinivasan Seshan, Randy H. Katz, Ken Lutz, David A. Patterson 0001, Edward K. Lee 0001, Peter M. Chen, Garth A. Gibson |
ISCA | 8 |
| 1994 | Cooperative Caching: Using Remote Client Memory to Improve File System Performance
Michael Dahlin, Randolph Y. Wang, Thomas E. Anderson, David A. Patterson 0001 |
OSDI | 4 |
| 1994 | A Quantitative Analysis of Cache Policies for Scalable Network File SystemsabstractCurrent network file system protocols rely heavily on a central server to coordinate file activity among client workstations. This central server can become a bottleneck that limits scalability for environments with large numbers of clients. In central server systems such as NFS and AFS, all client writes, cache misses, and coherence messages are handled by the server. To keep up with this workload, expensive server machines are needed, configured with high-performance CPUs, memory systems, and I/O channels. Since the server stores all data, it must be physically capable of connecting to many disks. This reliance on a central server also makes current systems inappropriate for wide area network use where the network bandwidth to the server may be limited. Michael Dahlin, Clifford Mather, Randolph Y. Wang, Thomas E. Anderson, David A. Patterson 0001 |
SIGMETRICS | 5 |
| 1994 | Toward Workload Characterization of Video Server and Digital Library ApplicationsabstractNo abstract available. Ann L. Drapeau, David A. Patterson 0001, Randy H. Katz |
SIGMETRICS | 2 |
| 1994 | Coding Techniques for Handling Failures in Large Disk Arrays
Lisa Hellerstein, Garth A. Gibson, Richard M. Karp, Randy H. Katz, David A. Patterson 0001 |
Algorithmica | 5 |
| 1994 | Performance and Design Evaluation of the RAID-II Storage Server
Peter M. Chen, Edward K. Lee 0001, Ann L. Drapeau, Ken Lutz, Ethan L. Miller, Srinivasan Seshan, Ken Shirriff, David A. Patterson 0001, Randy H. Katz |
Distributed Parallel Databases | 8 |
| 1994 | A New Approach to I/O Performance Evaluation - Self-Scaling I/O Benchmarks, Predicted I/O PerformanceabstractCurrent I/O benchmarks suffer from several chronic problems: they quickly become obsolete; they do not stress the I/O system; and they do not help much in understanding I/O system performance. We propose a new approach to I/O performance analysis. First, we propose a self-scaling benchmark that dynamically adjusts aspects of its workload according to the performance characteristic of the system being measured. By doing so, the benchmark automatically scales across current and future systems. The evaluation aids in understanding system performance by reporting how performance varies according to each of five workload parameters. Second, we propose predicted performance, a technique for using the results from the self-scaling evaluation to estimate quickly the performance for workloads that have not been measured. We show that this technique yields reasonably accurate performance estimates and argue that this method gives a far more accurate comparative performance evaluation than traditional single-point benchmarks. We apply our new evaluation technique by measuring a SPARCstation 1+ with one SCSI disk, an HP 730 with one SCSI-II disk, a DECstation 5000/200 running the Sprite LFS operating system with a three-disk disk array, a Convex C240 minisupercomputer with a four-disk disk array, and a Solbourne 5E/905 fileserver with a two-disk disk array. Peter M. Chen, David A. Patterson 0001 |
ACM Trans. Comput. Syst. | 2 |
| 1993 | LogP: Towards a Realistic Model of Parallel ComputationabstractA vast body of theoretical research has focused either on overly simplistic models of parallel computation, notably the PRAM, or overly specific models that have few representatives in the real world. Both kinds of models encourage exploitation of formal loopholes, rather than rewarding development of techniques that yield performance across a range of current and future parallel machines. This paper offers a new parallel machine model, called LogP, that reflects the critical technology trends underlying parallel computers. it is intended to serve as a basis for developing fast, portable parallel algorithms and to offer guidelines to machine designers. Such a model must strike a balance between detail and simplicity in order to reveal important bottlenecks without making analysis of interesting problems intractable. The model is based on four parameters that specify abstractly the computing bandwidth, the communication bandwidth, the communication delay, and the efficiency of coupling communication and computation. Portable parallel algorithms typically adapt to the machine configuration, in terms of these parameters. The utility of the model is demonstrated through examples that are implemented on the CM-5. David E. Culler, Richard M. Karp, David A. Patterson 0001, Abhijit Sahay, Klaus E. Schauser, Eunice E. Santos, Ramesh Subramonian, Thorsten von Eicken |
PPoPP | 3 |
| 1993 | A New Approach to I/O Performance Evaluation - Self-Scaling I/O Benchmarks, Predicted I/O PerformanceabstractCurrent I/O benchmarks suffer from several chronic problems: they quickly become obsolete, they do not stress the I/O system, and they do not help in understanding I/O system performance. We propose a new approach to I/O performance analysis. First, we propose a self-scaling benchmark that dynamically adjusts aspects of its workload according to the performance characteristic of the system being measured. By doing so, the benchmark automatically scales across current and future systems. The evaluation aids in understanding system performance by reporting how performance varies according to each of fie workload parameters. Second, we propose predicted performance, a technique for using the results from the self-scaling evaluation to quickly estimate the performance for workloads that have not been measured. We show that this technique yields reasonably accurate performance estimates and argue that this method gives a far more accurate comparative performance evaluation than traditional single point benchmarks. We apply our new evaluation technique by measuring a SPARCstation 1+ with one SCSI disk, an HP 730 with one SCSI-II disk, a Sprite LFS DECstation 5000/200 with a three-disk disk array, a Convex C240 minisupercomputer with a four-disk disk array, and a Solbourne 5E/905 fileserver with a two-disk disk array. Peter M. Chen, David A. Patterson 0001 |
SIGMETRICS | 2 |
| 1993 | Designing Disk Arrays for High Data Reliability
Garth A. Gibson, David A. Patterson 0001 |
J. Parallel Distributed Comput. | 2 |
| 1993 | Storage performance-metrics and benchmarksabstractThe metrics and benchmarks used in storage performance evaluation are discussed. The technology trends taking place in storage systems, such as disk and tape evolution, disk arrays, and solid-state disks, are highlighted. The current popular I/O benchmarks are then described, reviewed, and run on three systems: a DECstation 5000/200 running the Sprite Operating System, a SPARCstation 1+ running SunOS, and an HP Series 700 (Model 730) running HP-UX. Two approaches to storage benchmarks-LADDIS and a self-scaling benchmark with predicted performance-are also described.> Peter M. Chen, David A. Patterson 0001 |
Proc. IEEE | 2 |
| 1992 | Tradeoffs in Supporting Two Page SizesabstractAs computer system main memories get larger and processor cycles-per-instruction (CPIs) get smaller, the time spent in handling translation lookaside buffer (TLB) misses could become a performance bottleneck. We explore relieving this bottleneck by (a) increasing the page size and (b) supporting two page sizes. Madhusudhan Talluri, Shing I. Kong, Mark D. Hill, David A. Patterson 0001 |
ISCA | 4 |
| 1990 | Maximizing Performance in a Striped Disk ArrayabstractImprovements in disk speeds have not kept up with improvements in processor and memory speeds. One way to correct the resulting speed mismatch is to stripe data across many disks. In this paper, we address how to stripe data to get maximum performance from the disks. Specifically, we examine how to choose the striping unit, i.e. the amount of logically contiguous data on each disk. We synthesize rules for determining the best striping unit for a given range of workloads. Peter M. Chen, David A. Patterson 0001 |
ISCA | 2 |
| 1990 | An Evaluation of Redundant Arrays of Disks Using an Amdahl 5890abstractRecently we presented several disk array architectures designed to increase the data rate and I/O rate of supercomputing applications, transaction processing, and file systems [Patterson 88]. In this paper we present a hardware performance measurement of two of these architectures, mirroring and rotated parity. We see how throughput for these two architectures is affected by response time requirements, request sizes, and read to write ratios. We find that for applications with large accesses, such as many supercomputing applications, a rotated parity disk array far outperforms traditional mirroring architecture. For applications dominated by small accesses, such as transaction processing, mirroring architectures have higher performance per disk than rotated parity architectures. Peter M. Chen, Garth A. Gibson, Randy H. Katz, David A. Patterson 0001 |
SIGMETRICS | 4 |
| 1989 | Failure Correction Techniques for Large Disk Arrays
Garth A. Gibson, Lisa Hellerstein, Richard M. Karp, Randy H. Katz, David A. Patterson 0001 |
ASPLOS | 5 |
| 1989 | Disk system architectures for high performance computingabstractFollowing a brief review of the fundamentals of disk system architecture, the characteristics of the applications that demand high I/O system performance are described. Conventional ways to improve disk performance are discussed. New developments in disk array systems are introduced, and controller architectures are described.> Randy H. Katz, Garth A. Gibson, David A. Patterson 0001 |
Proc. IEEE | 3 |
| 1988 | A Case for Redundant Arrays of Inexpensive Disks (RAID)abstractIncreasing performance of CPUs and memories will be squandered if not matched by a similar performance increase in I/O. While the capacity of Single Large Expensive Disks (SLED) has grown rapidly, the performance improvement of SLED has been modest. Redundant Arrays of Inexpensive Disks (RAID), based on the magnetic disk technology developed for personal computers, offers an attractive alternative to SLED, promising improvements of an order of magnitude in performance, reliability, power consumption, and scalability. This paper introduces five levels of RAIDs, giving their relative cost/performance, and compares RAID to an IBM 3380 and a Fujitsu Super Eagle. David A. Patterson 0001, Garth A. Gibson, Randy H. Katz |
SIGMOD Conference | 1 |
| 1988 | The Design of XPRS
Michael Stonebraker, Randy H. Katz, David A. Patterson 0001, John K. Ousterhout |
VLDB | 3 |
| 1987 | Fast multiply and divide for a VLSI floating-point unitabstractThis paper presents the design of a fast and area-efficient multiply-divide unit used in building a VLSI floating-point processor (FPU), conforming to the IEEE standard 754. Details of the algorithms, implementation techniques and design tradeoffs are presented, The multiplier and divider are implemented in 2 micron CMOS technology with two layers of metal, and occupy 23 square mm (23% of the entire FPU). We expect to perform extended-precision multiplication and division in 1.1 and 2.8 microseconds, respectively. Bidyut Kumar Bose, Li-fan Pei, George S. Taylor, David A. Patterson 0001 |
IEEE Symposium on Computer Arithmetic | 4 |
| 1986 | Evaluation of the SPUR Lisp ArchitectureabstractThe SPUR microprocessor has a 40-bit tagged architecture designed to improve its performance for Lisp programs. Although SPUR includes just a small set of enhancements to the Berkeley RISC-II architecture, simulation results show that with a 150-ns cycle time SPUR will run Common Lisp programs at least as fast as a Symbolies 3600 or a DEC VAX 8600. This paper explains SPUR's instruction set architecture and provides measurements of how certain components of the architecture perform. George S. Taylor, Paul N. Hilfinger, James R. Larus, David A. Patterson 0001, Benjamin G. Zorn |
ISCA | 4 |
| 1986 | An In-Cache Address Translation MechanismabstractIn the design of SPUR, a high-performance multiprocessor workstation, the use of large caches and hardware-supported cache consistency suggests a new approach to virtual address translation. By performing translation in each processor's virtually-tagged cache, the need for separate translation lookaside buffers (TLBs) is eliminated. Eliminating the TLB substantially reduces the hardware cost and complexity of the translation mechanism and eliminates the translation consistency problem. Trace-driven simulations show that normal cache behavior is only minimally affected by caching page table entries, and that in many cases, using a separate device would actually reduce system performance. David A. Wood 0001, Susan J. Eggers, Garth A. Gibson, Mark D. Hill, Joan M. Pendleton, Scott A. Ritchie, George S. Taylor, Randy H. Katz, David A. Patterson 0001 |
ISCA | 9 |
| 1984 | Architecture of SOAR: Smalltalk on a RISCabstractSmalltalk on a RISC (SOAR) is a simple, Von Neumann computer that is designed to execute the Smalltalk-80 system much faster than existing VLSI microcomputers. The Smalltalk-80 system is a highly productive programming environment but poses tough challenges for implementors: dynamic data typing, a high level instruction set, frequent and expensive procedure calls, and object-oriented storage management. SOAR compiles programs to a low level, efficient instruction set. Parallel tag checks permit high performance for the simple common cases and cause traps to software routines for the complex cases. Parallel register initialization and multiple on-chip register windows speed procedure calls. Sophisticated software techniques relieve the hardware of the burden of managing objects. We have initial evaluations of the effectiveness of the SOAR architecture by compiling and simulating benchmarks, and will prove SOAR's feasibility by fabricating a 35,000-transistor SOAR chip. These early results suggest that a Reduced Instruction Set Computer can provide high performance in an exploratory programming environment. David M. Ungar, Ricki Blau, Peter Foley, A. Dain Samples, David A. Patterson 0001 |
ISCA | 5 |
| 1983 | Architecture of a VLSI Instruction Cache for a RISCabstractA cache was first used in a commercial computer in 1968,1 and researchers have spent the last 15 years analyzing caches and suggesting improvements. In designing a VLSI instruction cache for a RISC microprocessor we have uncovered four ideas potentially applicable to other VLSI machines. These ideas provide expansible cache memory, increased cache speed, reduced program code size, and decreased manufacturing costs. These improvements blur the habitual distinction between an instruction cache and an instruction fetch unit. David A. Patterson 0001, Phil Garrison, Mark D. Hill, Dimitris Lioupis, Chris Nyberg, Tim Sippel, Korbin S. Van Dyke |
ISCA | 1 |
| 1982 | RISC assessment: A high-level language experimentabstractWe present the result of an informal experiment comparing the performance of one Reduced Instruction Set Computer, RISC I, to five traditional computers, VAX-11/780, PDP-11/70, BBN C/70, MC68000, and Z8000, in a high-level language environment. Measuring either absolute performance or the penalty for using high-level languages, the best computer is RISC I. David A. Patterson 0001, Richard S. Piepho |
ISCA | 1 |
| 1981 | VAX hardware for the proposed IEEE floating-point standardabstractThe proposed IEEE floating-point standard has been implemented in a substitute floating-point accelerator for the VAX∗∗11/780. We explain how features of the proposed standard influenced the design of the new processor. By comparing it with the original VAX accelerator, we illustrate the differences between hardware for the proposed standard and hardware for a more traditional floating-point architecture. George S. Taylor, David A. Patterson 0001 |
IEEE Symposium on Computer Arithmetic | 2 |
| 1981 | RISC I: A Reduced Instruction Set VLSI Computer
David A. Patterson 0001, Carlo H. Séquin |
ISCA | 1 |
| 1980 | Retrospective on High-Level Language Computer ArchitectureabstractHigh-level language computers (HLLC) have attracted interest in the architectural and programming community during the last 15 years; proposals have been made for machines directed towards the execution of various languages such as ALGOL,1,2 APL,3,4,5 BASIC,6,7 COBOL,8,9 FORTRAN,10,ll LISP,12,13 PASCAL,14 PL/I,15,16,17 SNOBOL,18,19 and a host of specialized languages. Though numerous designs have been proposed, only a handful of high-level language computers have actually been implemented.4,7,9,20,21 In examining the goals and successes of high-level language computers, the authors have found that most designs suffer from fundamental problems stemming from a misunderstanding of the issues involved in the design, use, and implementation of cost-effective computer systems. It is the intent of this paper to identify and discuss several issues applicable to high-level language computer architecture, to provide a more concrete definition of high-level language computers, and to suggest a direction for high-level language computer architectures of the future. David R. Ditzel, David A. Patterson 0001 |
ISCA | 2 |
| 1980 | Design Considerations for Single-Chip Computers of the FutureabstractIn the mid 1980's it will be possible to put a million devices (transistors or active MOS gate electrodes) onto a single silicon chip. General trends in the evolution of silicon integrated circuits are reviewed and design constraints for emerging VLSI circuits are analyzed. Desirable architectural features in modern computers are then discussed and consequences for an implementation with large-scale integrated circuits are investigated. The resulting recommended processor design includes features such as an on-chip memory hierarchy, multiple homogeneous caches for enhanced execution parallelism, support for complex data structures and high-level languages, a flexible instruction set, and communication hardware. It is concluded that a viable modular building block for the next generation of computing systems will be a self-contained computer on a single chip. A tentative allocation of the one milion transistors to the various functional blocks is given, and the result is a memory intensive design. David A. Patterson 0001, Carlo H. Séquin |
IEEE Trans. Computers | 1 |
| 1979 | Design Considerations for the VLSI Processor of X-treeabstractX-NODE is a single-chip VLSI processor to be realized in the mid 1980's and to be used as a building block for a tree-structured multiprocessor system (X-TREE). Three major trends influence the design of this processor: the continuing evolution of VLSI technology, the requirements for parallelism and communication in a multiprocessor system, and the need for better support of software and high level language constructs. The influence of these trends on the processor architecture are discussed and the current state of the design of X-NODE is outlined. X-NODE will introduce several new features exploiting the full potential of VLSI technology. The processor and hierarchical memory of multiple device types will be combined on a single chip to provide a powerful processor. With basically a memory-to-memory architecture, an on-chip caching scheme provides the performance of a register based architecture. This on-chip memory hierarchy contains program and data, as well as microcode. The instruction set of any processor can thus be dynamically changed and tailored to the specific problem being executed. It is planned to support high level language constructs directly in hardware through mechanisms such as bounds checking. David A. Patterson 0001, E. Scott Fehr, Carlo H. Séquin |
ISCA | 1 |
| 1978 | X-Tree: A Tree Structured Multi-Processor Computer ArchitectureabstractThe problem of organizing multiple, monolithic microprocessors into an effective general purpose computer structure is examined. A tree structure with extra interconnections was found to be especially attractive. It provides a structured hierarchy for control, addressing and message routing. More important, it appears to provide a mechanism to automatically migrate data abstractions and processes over the network of processors. The network can be expanded to any desired size and no global control or routine mechanisms are needed. Alvin M. Despain, David A. Patterson 0001 |
ISCA | 2 |
| 1976 | Strum: Structured Microprogram Development System for Correct FirmwareabstractAn approach to the development of correct microprograms is to use the methodologies that have been beneficial in the generation of correct user programs, i. e., structured programming, high-level languages (HLL's), and formal program verification using Floyd's inductive assertion method. This paper presents a system that combines these techniques to simplify the design and implementation of correct microprograms for a real microprogrammable computer. It gives some statistics which support our emphasis on generation as well as correctness and some preliminary results on the use of our system. David A. Patterson 0001 |
IEEE Trans. Computers | 1 |