EDBT 2026 Demo / reviewers in the wild / expert
James Laudon
dblp:65/336
· DBLP profile ↗
18ranked-venue papers
3as first author
4since 2021 · last 2023
0000-0003-1061-1780ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 10 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Deep learning architectures and training · 42% Efficient and distributed learning · 19% Learning paradigms · 17% | |
| Computer architecture, parallel and distributed computing, and storage systems
12 papers |
Hardware accelerators and domain-specific architectures · 55% Processor architecture and microarchitecture · 20% Memory systems · 10% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 30 heaviest of 55, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
mixture of experts |
1.9 | 3 | 2023 | Brainformers: Trading Simplicity for Efficiency · ICML 2023 Lifelong Language Pretraining with Distribution-Specialized Experts · ICML 2023 Mixture-of-Experts with Expert Choice Routing · NeurIPS 2022 |
Machine learning › Representation and self-supervised learning
pre-training |
0.8 | 2 | 2023 | Lifelong Language Pretraining with Distribution-Specialized Experts · ICML 2023 Mixture-of-Experts with Expert Choice Routing · NeurIPS 2022 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.8 | 2 | 2021 | Ten Lessons From Three Generations Shaped Google's TPUv4i : Industrial Product · ISCA 2021 In-Datacenter Performance Analysis of a Tensor Processing Unit · ISCA 2017 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.7 | 1 | 2023 | Lifelong Language Pretraining with Distribution-Specialized Experts · ICML 2023 |
Machine learning › Deep learning architectures and training › transformer
efficient transformer |
0.7 | 1 | 2023 | Brainformers: Trading Simplicity for Efficiency · ICML 2023 |
Machine learning › Learning paradigms
lifelong learning |
0.7 | 1 | 2023 | Lifelong Language Pretraining with Distribution-Specialized Experts · ICML 2023 |
Machine learning › Optimization for machine learning
sparse model |
0.7 | 1 | 2023 | Brainformers: Trading Simplicity for Efficiency · ICML 2023 |
Machine learning › Deep learning architectures and training
transformer |
0.7 | 1 | 2023 | Brainformers: Trading Simplicity for Efficiency · ICML 2023 |
Machine learning › Efficient and distributed learning › model compression › sparsity
activation sparsity |
0.6 | 1 | 2022 | Mixture-of-Experts with Expert Choice Routing · NeurIPS 2022 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
inference accelerator |
0.5 | 1 | 2021 | Ten Lessons From Three Generations Shaped Google's TPUv4i : Industrial Product · ISCA 2021 |
Processor architecture and microarchitecture › instruction-level parallelism
VLIW |
0.5 | 1 | 2021 | Ten Lessons From Three Generations Shaped Google's TPUv4i : Industrial Product · ISCA 2021 |
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
device placement |
0.4 | 1 | 2020 | Transferable Graph Optimizers for ML Compilers · NeurIPS 2020 |
Machine learning › Efficient and distributed learning
distributed training |
0.4 | 1 | 2020 | Transferable Graph Optimizers for ML Compilers · NeurIPS 2020 |
Compilers and program optimization › compiler optimization
computation graph optimization |
0.4 | 1 | 2020 | Transferable Graph Optimizers for ML Compilers · NeurIPS 2020 |
Compilers and program optimization › compiler optimization
machine learning for compiler optimization |
0.4 | 1 | 2020 | Transferable Graph Optimizers for ML Compilers · NeurIPS 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › inference accelerator
neural network inference accelerator |
0.3 | 1 | 2017 | In-Datacenter Performance Analysis of a Tensor Processing Unit · ISCA 2017 |
Hardware accelerators and domain-specific architectures › tensor accelerator
tensor processing unit |
0.3 | 1 | 2017 | In-Datacenter Performance Analysis of a Tensor Processing Unit · ISCA 2017 |
Natural language and speech › Language models and text generation
large language model |
0.2 | 1 | 2023 | Brainformers: Trading Simplicity for Efficiency · ICML 2023 |
Performance modeling and evaluation
benchmarking |
0.1 | 2 | 2017 | In-Datacenter Performance Analysis of a Tensor Processing Unit · ISCA 2017 The SGI Origin: A ccNUMA Highly Scalable Server · ISCA 1997 |
Performance modeling and evaluation › benchmarking › computer architecture benchmarking
accelerator benchmarking |
0.1 | 1 | 2017 | In-Datacenter Performance Analysis of a Tensor Processing Unit · ISCA 2017 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 2 | 2007 | Fair Queuing Memory Systems · MICRO 2006 Virtual private caches · ISCA 2007 |
Cloud and datacenter computing
quality of service |
0.1 | 2 | 2007 | Fair Queuing Memory Systems · MICRO 2006 Virtual private caches · ISCA 2007 |
Memory systems › cache management
cache capacity management |
0.1 | 1 | 2007 | Virtual private caches · ISCA 2007 |
Memory systems
cache management |
0.1 | 1 | 2007 | Virtual private caches · ISCA 2007 |
Memory systems
cache coherence |
0.0 | 4 | 1997 | The SGI Origin: A ccNUMA Highly Scalable Server · ISCA 1997 The DASH Prototype: Logic Overhead and Performance · IEEE Trans. Parallel Distributed Syst. 1993 The DASH Prototype: Implementation and Performance · ISCA 1992 |
Memory systems › cache coherence
directory-based coherence |
0.0 | 4 | 1997 | The SGI Origin: A ccNUMA Highly Scalable Server · ISCA 1997 The DASH Prototype: Logic Overhead and Performance · IEEE Trans. Parallel Distributed Syst. 1993 The DASH Prototype: Implementation and Performance · ISCA 1992 |
Parallel and multicore computing
multiprocessor system |
0.0 | 2 | 1997 | The SGI Origin: A ccNUMA Highly Scalable Server · ISCA 1997 The DASH Prototype: Logic Overhead and Performance · IEEE Trans. Parallel Distributed Syst. 1993 |
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor |
0.0 | 3 | 1993 | The DASH Prototype: Logic Overhead and Performance · IEEE Trans. Parallel Distributed Syst. 1993 The DASH Prototype: Implementation and Performance · ISCA 1992 Memory Consistency and Event Ordering in Scalable Shared-Memory Multiprocessors · ISCA 1990 |
Processor architecture and microarchitecture › multithreading
multiple-context processor |
0.0 | 2 | 1994 | Interleaving: A Multithreading Technique Targeting Multiprocessors and Workstations · ASPLOS 1994 Architectural and implementation tradeoffs in the design of multiple-context processors · ISCA 1992 |
Processor architecture and microarchitecture
multithreading |
0.0 | 2 | 1994 | Interleaving: A Multithreading Technique Targeting Multiprocessors and Workstations · ASPLOS 1994 Architectural and implementation tradeoffs in the design of multiple-context processors · ISCA 1992 |
Methods — techniques the papers use, named apart from their topics
sequential attention · 0.9graph neural network · 0.9deep reinforcement learning · 0.9regularization · 0.7neural architecture search · 0.7layer normalization · 0.7expert addition · 0.7top-k gating · 0.6sparse activation · 0.6simulation · 0.1fair queuing · 0.1hardware performance monitor · 0.0workload characterization · 0.0performance measurement · 0.0protocol verification · 0.0static instruction scheduling · 0.0dynamic scheduling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Lifelong Language Pretraining with Distribution-Specialized ExpertsabstractPretraining on a large-scale corpus has become a standard method to build general language models (LMs). Adapting a model to new data distributions targeting different downstream tasks poses significant challenges. Naive fine-tuning may incur catastrophic forgetting when the over-parameterized LMs overfit the new data but fail to preserve the pretrained features. Lifelong learning (LLL) aims to enable information systems to learn from a continuous data stream across time. However, most prior work modifies the training recipe assuming a static fixed network architecture. We find that additional model capacity and proper regularization are key elements to achieving strong LLL performance. Thus, we propose Lifelong-MoE, an extensible MoE (Mixture-of-Experts) architecture that dynamically adds model capacity via adding experts with regularized pretaining. Our results show that by only introducing a limited number of extra experts while keeping the computation cost constant, our model can steadily adapt to data distribution shifts while preserving the previous knowledge. Compared to existing lifelong learning approaches, Lifelong-MoE achieves better few-shot performance on NLP tasks. More impressively, Lifelong-MoE surpasses multi-task learning on 19 downstream NLU tasks. Wuyang Chen 0001, Yanqi Zhou, Nan Du 0002, Yanping Huang, James Laudon, Claire Cui |
ICML | 5 |
| 2023 | Brainformers: Trading Simplicity for EfficiencyabstractTransformers are central to recent successes in natural language processing and computer vision. Transformers have a mostly uniform backbone where layers alternate between feed-forward and self-attention in order to build a deep network. Here we investigate this design choice and find that more complex blocks that have different permutations of layer primitives can be more efficient. Using this insight, we develop a complex block, named Brainformer, that consists of a diverse sets of layers such as sparsely gated feed-forward layers, dense feed-forward layers, attention layers, and various forms of layer normalization and activation functions. Brainformer consistently outperforms the state-of-the-art dense and sparse Transformers, in terms of both quality and efficiency. A Brainformer model with 8 billion activated parameters per token demonstrates 2x faster training convergence and 5x faster step time compared to its GLaM counterpart. In downstream task evaluation, Brainformer also demonstrates a 3% higher SuperGLUE score with fine-tuning compared to GLaM with a similar number of activated parameters. Finally, Brainformer largely outperforms a Primer dense model derived with NAS with similar computation per token on fewshot evaluations. Yanqi Zhou, Nan Du 0002, Yanping Huang, Daiyi Peng, Chang Lan, Siamak Shakeri, David R. So, Andrew M. Dai, Yifeng Lu, Quoc V. Le, Claire Cui, James Laudon, Jeffrey Dean |
ICML | 14 |
| 2022 | Mixture-of-Experts with Expert Choice RoutingabstractSparsely-activated Mixture-of-experts (MoE) models allow the number of parameters to greatly increase while keeping the amount of computation for a given token or a given sample unchanged. However, a poor expert routing strategy (e.g. one resulting in load imbalance) can cause certain experts to be under-trained, leading to an expert being under or over-specialized. Prior work allocates a fixed number of experts to each token using a top-k function regardless of the relative importance of different tokens. To address this, we propose a heterogeneous mixture-of-experts employing an expert choice method. Instead of letting tokens select the top-k experts, we have experts selecting the top-k tokens. As a result, each token can be routed to a variable number of experts and each expert can have a fixed bucket size. We systematically study pre-training speedups using the same computational resources of the Switch Transformer top-1 and GShard top-2 gating of prior work and find that our method improves training convergence time by more than 2×. For the same computational cost, our method demonstrates higher performance in fine-tuning 11 selected tasks in the GLUE and SuperGLUE benchmarks. For a smaller activation cost, our method outperforms the T5 dense model in 7 out of the 11 tasks. Yanqi Zhou, Tao Lei 0001, Hanxiao Liu, Nan Du 0002, Yanping Huang, Vincent Y. Zhao, Andrew M. Dai, Quoc V. Le, James Laudon |
NeurIPS | 10 |
| 2021 | Ten Lessons From Three Generations Shaped Google's TPUv4i : Industrial ProductabstractGoogle deployed several TPU generations since 2015, teaching us lessons that changed our views: semi-conductor technology advances unequally; compiler compatibility trumps binary compatibility, especially for VLIW domain-specific architectures (DSA); target total cost of ownership vs initial cost; support multi-tenancy; deep neural networks (DNN) grow 1.5X annually; DNN advances evolve workloads; some inference tasks require floating point; inference DSAs need air-cooling; apps limit latency, not batch size; and backwards ML compatibility helps deploy DNNs quickly. These lessons molded TPUv4i, an inference DSA deployed since 2020. Norman P. Jouppi, Doe Hyun Yoon, Matthew Ashcraft, Mark Gottscho, Thomas B. Jablin, George Kurian, James Laudon, Sheng Li 0007, Peter C. Ma, Thomas Norrie, Nishant Patil, Sushma Prasad, Cliff Young, Zongwei Zhou, David A. Patterson 0001 |
ISCA | 7 |
| 2020 | Google's Training Chips Revealed: TPUv2 and TPUv3abstractThis article consists only of a collection of slides from the author's conference presentation. Thomas Norrie, Nishant Patil, Doe Hyun Yoon, George Kurian, Sheng Li 0007, James Laudon, Cliff Young, Norman P. Jouppi, David A. Patterson 0001 |
Hot Chips Symposium | 6 |
| 2020 | Transferable Graph Optimizers for ML CompilersabstractMost compilers for machine learning (ML) frameworks need to solve many correlated optimization problems to generate efficient machine code. Current ML compilers rely on heuristics based algorithms to solve these optimization problems one at a time. However, this approach is not only hard to maintain but often leads to sub-optimal solutions especially for newer model architectures. Existing learning based approaches in the literature are sample inefficient, tackle a single optimization problem, and do not generalize to unseen graphs making them infeasible to be deployed in practice. To address these limitations, we propose an end-to-end, transferable deep reinforcement learning method for computational graph optimization (GO), based on a scalable sequential attention mechanism over an inductive graph neural network. GO generates decisions on the entire graph rather than on each individual node autoregressively, drastically speeding up the search compared to prior methods. Moreover, we propose recurrent attention layers to jointly optimize dependent graph optimization tasks and demonstrate 33%-60% speedup on three graph optimization tasks compared to TensorFlow default optimization. On a diverse set of representative graphs consisting of up to 80,000 nodes, including Inception-v3, Transformer-XL, and WaveNet, GO achieves on average 21% improvement over human experts and 18% improvement over the prior state of the art with 15x faster convergence, on a device placement task evaluated in real systems. Yanqi Zhou, Sudip Roy 0002, AmirAli Abdolrashidi, Daniel Wong 0001, Peter C. Ma, Qiumin Xu, Hanxiao Liu, Mangpo Phitchaya Phothilimtha, Anna Goldie, Azalia Mirhoseini, James Laudon |
NeurIPS | 12 |
| 2017 | In-Datacenter Performance Analysis of a Tensor Processing UnitabstractMany architects believe that major improvements in cost-energy-performance must now come from domain-specific hardware. This paper evaluates a custom ASIC---called a Tensor Processing Unit (TPU) --- deployed in datacenters since 2015 that accelerates the inference phase of neural networks (NN). The heart of the TPU is a 65,536 8-bit MAC matrix multiply unit that offers a peak throughput of 92 TeraOps/second (TOPS) and a large (28 MiB) software-managed on-chip memory. The TPU's deterministic execution model is a better match to the 99th-percentile response-time requirement of our NN applications than are the time-varying optimizations of CPUs and GPUs that help average throughput more than guaranteed latency. The lack of such features helps explain why, despite having myriad MACs and a big memory, the TPU is relatively small and low power. We compare the TPU to a server-class Intel Haswell CPU and an Nvidia K80 GPU, which are contemporaries deployed in the same datacenters. Our workload, written in the high-level TensorFlow framework, uses production NN applications (MLPs, CNNs, and LSTMs) that represent 95% of our datacenters' NN inference demand. Despite low utilization for some applications, the TPU is on average about 15X -- 30X faster than its contemporary GPU or CPU, with TOPS/Watt about 30X -- 80X higher. Moreover, using the CPU's GDDR5 memory in the TPU would triple achieved TOPS and raise TOPS/Watt to nearly 70X the GPU and 200X the CPU. Norman P. Jouppi, Cliff Young, Nishant Patil, David A. Patterson 0001, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, Rick Boyle, Pierre-luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tara Vazir Ghaemmaghami, Rajendra Gottipati, William Gulland, Robert Hagmann, Richard Ho 0001, Doug Hogberg, John Hu, Robert Hundt, Dan Hurt, Julian Ibarz, Aaron Jaffey, Alek Jaworski, Alexander Kaplan, Harshit Khaitan, Daniel Killebrew, Andy Koch, Steve Lacy, James Laudon, James Law, Diemthu Le, Chris Leary, Zhuyuan Liu, Kyle Lucke, Alan Lundin, Gordon MacKean, Adriana Maggiore, Maire Mahony, Kieran Miller, Rahul Nagarajan, Ravi Narayanaswami, Ray Ni, Kathy Nix, Thomas Norrie, Mark Omernick, Narayana Penukonda, Andy Phelps, Jonathan Ross, Amir Salek, Emad Samadiani, Chris Severn, Gregory Sizikov, Matthew Snelham, Jed Souter, Dan Steinberg, Andy Swing, Mercedes Tan, Gregory Thorson, Horia Toma, Erick Tuttle, Vijay Vasudevan, Richard Walter, Walter Wang, Eric Wilcox, Doe Hyun Yoon |
ISCA | 38 |
| 2007 | Virtual private cachesabstractVirtual Private Machines (VPM) provide a framework for Quality of Service (QoS) in CMP-based computer systems. VPMs incorporate microarchitecture mechanisms that allow shares of hardware resources to be allocated to executing threads, thus providing applications with an upper bound on execution time regardless of other thread activity. Virtual Private Caches (VPCs) are an important element of VPMs. VPC hardware consists of two major components: the VPC Arbiter, which manages shared cache bandwidth, and the VPC Capacity Manager, which manages the cache storage. Both the VPC Arbiter and VPC Capacity Manager provide minimum service guarantees that, when combined, achieve QoS for the cache subsystem. Simulation-based evaluation shows that conventional cache bandwidth management policies allow concurrently executing threads to affect each other significantly in an uncontrollable manner. The evaluation targets cache bandwidth because the effects of cache capacity sharing have been studied elsewhere. In contrast with the conventional policies, the VPC Arbiter meets its QoS performance objectives on all workloads studied and over a range of allocated bandwidth levels. The VPC Arbiter’s fairness policy, which distributes leftover bandwidth, mitigates the effects of cache preemption latencies, thus ensuring threads a high-degree of performance isolation. Furthermore, the VPC Arbiter eliminates negative bandwidth interference which can improve aggregate throughput and resource utilization. Kyle J. Nesbit, James Laudon, James E. Smith 0001 |
ISCA | 2 |
| 2006 | Fair Queuing Memory SystemsabstractWe propose and evaluate a multi-thread memory scheduler that targets high performance CMPs. The proposed memory scheduler is based on concepts originally developed for network fair queuing scheduling algorithms. The memory scheduler is fair and provides quality of service (QoS) while improving system performance. On a four processor CMP running workloads containing a mix of applications with a range of memory bandwidth demands, the proposed memory scheduler provides QoS to all of the threads in all of the workloads, improves system performance by an average of 14% (41% in the best case), and reduces the variance in the threads' target memory bandwidth utilization from .2 to .0058 Kyle J. Nesbit, Nidhi Aggarwal, James Laudon, James E. Smith 0001 |
MICRO | 3 |
| 1997 | The SGI Origin: A ccNUMA Highly Scalable ServerabstractThe SGI Origin 2000 is a cache-coherent non-uniform memory access (ccNUMA) multiprocessor designed and manufactured by Silicon Graphics, Inc. The Origin system was designed from the ground up as a multiprocessor capable of scaling to both small and large processor counts without any bandwidth, latency, or cost cliffs. The Origin system consists of up to 512 nodes interconnected by a scalable Craylink network. Each node consists of one or two R10000 processors, up to 4 GB of coherent memory, and a connection to a portion of the XIO IO subsystem. This paper discusses the motivation for building the Origin 2000 and then describes its architecture and implementation. In addition, performance results are presented for the NAS Parallel Benchmarks V2.2 and the SPLASH2 applications. Finally, the Origin system is compared to other contemporary commercial ccNUMA systems. James Laudon, Daniel Lenoski |
ISCA | 1 |
| 1994 | Interleaving: A Multithreading Technique Targeting Multiprocessors and WorkstationsabstractThere is an increasing trend to use commodity microprocessors as the compute engines in large-scale multiprocessors. However, given that the majority of the microprocessors are sold in the workstation market, not in the multiprocessor market, it is only natural that architectural features that benefit only multiprocessors are less likely to be adopted in commodity microprocessors. In this paper, we explore multiple-context processors, an architectural technique proposed to hide the large memory latency in multiprocessors. We show that while current multiple-context designs work reasonably well for multiprocessors, they are ineffective in hiding the much shorter uniprocessor latencies using the limited parallelism found in workstation environments. We propose an alternative design that combines the best features of two existing approaches, and present simulation results that show it yields better performance for both multiprogrammed workloads on a workstation and parallel applications on a multiprocessor. By addressing the needs of the workstation environment, our proposal makes multiple contexts more attractive for commodity microprocessors. James Laudon, Anoop Gupta, Mark Horowitz |
ASPLOS | 1 |
| 1993 | The DASH Prototype: Logic Overhead and PerformanceabstractThe fundamental premise behind the DASH project is that it is feasible to build large-scale shared-memory multiprocessors with hardware cache coherence. The hardware overhead of directory-based cache coherence in a 48-processor is examined. The data show that the overhead is only about 10-15%, which appears to be a small cost for the ease of programming offered by coherent caches and the potential for higher performance. The performance of the system is discussed, and the speedups obtained by a variety of parallel applications running on the prototype are shown. Using a sophisticated hardware performance monitor, the effectiveness of coherent caches and the relationship between an application's reference behavior and its speedup are characterized. The optimizations incorporated in the DASH protocol are evaluated in terms of their effectiveness on parallel applications and on atomic tests that stress the memory system.> Daniel Lenoski, James Laudon, Truman Joe, David Nakahira, Luis Stevens, Anoop Gupta, John L. Hennessy |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1992 | Architectural and implementation tradeoffs in the design of multiple-context processorsabstractWe examine two multiple-context schemes in the context of scalable shared-memory multiprocessors. The blocked scheme switches between contexts at cache misses. The proposed interleaved scheme switches between available contexts on a cycle-by-cycle basis, while providing full pipeline interlocks for good single-context performance. We show the interleaved scheme to have a performance advantage over the blocked scheme due to its ability to hide pipeline dependencies and reduce the context switch cost. We also show that, while the implementation of the interleaved scheme is more complex, this complexity is not overwhelming. James Laudon, Anoop Gupta, Mark Horowitz |
ISCA | 1 |
| 1992 | The DASH Prototype: Implementation and PerformanceabstractThe fundamental premise behind the DASH project is that it is feasible to build large-scale shared-memory multiprocessors with hardware cache coherence. While paper studies and software simulators are useful for understanding many high-level design trade-offs, prototypes are essential to ensure that no critical details are overlooked. A prototype provides convincing evidence of the feasibility of the design allows one to accurately estimate both the hardware and the complexity cost of various features, and provides a platform for studying real workloads. A 16-processor prototype of the DASH multiprocessor has been operational for the last six months. In this paper, the hardware overhead of directory-based cache coherence in the prototype is examined. We also discuss the performance of the system, and the speedups obtained by parallel applications running on the prototype. Using a sophisticated hardware performance monitor, we characterize the effectiveness of coherent caches and the relationship between an application's reference behavior and its speedup. Daniel Lenoski, James Laudon, Truman Joe, David Nakahira, Luis Stevens, Anoop Gupta, John L. Hennessy |
ISCA | 2 |
| 1990 | Memory Consistency and Event Ordering in Scalable Shared-Memory MultiprocessorsabstractScalable shared-memory multiprocessors distribute memory among the processors and use scalable interconnection networks to provide high bandwidth and low latency communication. In addition, memory accesses are cached, buffered, and pipelined to bridge the gap between the slow shared memory and the fast processors. Unless carefully controlled, such architectural optimizations can cause memory accesses to be executed in an order different from what the programmer expects. The set of allowable memory access orderings forms the memory consistency model or event ordering model for an architecture. Kourosh Gharachorloo, Daniel Lenoski, James Laudon, Phillip B. Gibbons, Anoop Gupta, John L. Hennessy |
ISCA | 3 |
| 1990 | The Directory-Based Cache Coherence Protocol for the DASH MultiprocessorabstractDASH is a scalable shared-memory multiprocessor currently being developed at Stanford's Computer Systems Laboratory. The architecture consists of powerful processing nodes, each with a portion of the shared-memory, connected to a scalable interconnection network. A key feature of DASH is its distributed directory-based cache coherence protocol. Unlike traditional snoopy coherence protocols, the DASH protocol does not rely on broadcast; instead it uses point-to-point messages sent between the processors and memories to keep caches consistent. Furthermore, the DASH system does not contain any single serialization or control point. While these features provide the basis for scalability, they also force a reevaluation of many fundamental issues involved in the design of a protocol. These include the issues of correctness, performance and protocol complexity. In this paper, we present the design of the DASH coherence protocol and discuss how it addresses the above issues. We also discuss our strategy for verifying the correctness of the protocol and briefly compare our protocol to the IEEE Scalable Coherent Interface protocol. Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, John L. Hennessy |
ISCA | 2 |
| 1988 | The Astronautics ZS-1 processorabstractThe Astronautics ZS-1 is a high speed minisupercomputer system designed for scientific and engineering applications. The ZS-1 central processor uses a decoupled architecture which splits instruction words into two streams, one for fixed-point/memory-address computation and the other for floating-point operations. The two instruction streams are then processed in parallel, with architectural queues providing communication between the streams. Pipelining is used extensively throughout the ZS-1. The combination of decoupling and pipelining provides overall, sustained performance of about one third a CRAY-XMP-1 on compiled double-precision Fortran. The authors give a description of the decoupled architecture and discuss other forms of parallelism used in the ZS-1. A discussion of static and dynamic instruction scheduling is also included.> James E. Smith 0001, Greg Dermer, B. D. Vanderwarn, S. D. Klinger, C. M. Rozewski, D. L. Fowler, K. R. Scidmore, James Laudon |
ICCD | 8 |
| 1987 | The ZS-1 Central ProcessorabstractThe Astronautics ZS-1 is a high speed, 64-bit computer system designed for scientific and engineering applications. The ZS-1 central processor uses a decoupled architecture, which splits instructions into two streams---one for fixed point/memory address computation and the other for floating point operations. The two instruction streams are then processed in parallel. Pipelining is also used extensively throughout the ZS-1.This paper describes the architecture and implementation of the ZS-1 central processor, beginning with some of the basic design objectives. Descriptions of the instruction set, pipeline structure, and virtual memory implementation demonstrate the methods used to satisfy the objectives. High performance is achieved through a combination of static (compile-time) instruction scheduling and dynamic (run-time) scheduling. Both types of scheduling are illustrated with examples. James E. Smith 0001, Greg Dermer, B. D. Vanderwarn, S. D. Klinger, C. M. Rozewski, D. L. Fowler, K. R. Scidmore, James Laudon |
ASPLOS | 8 |