Mark Gottscho

dblp:115/9136 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
2since 2021 · last 2024
0000-0001-8370-4158ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Theory of computation · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Memory systems · 34% Hardware accelerators and domain-specific architectures · 26% Energy-efficient computing · 15%
Theoretical computer science
1 paper
Coding theory · 100%

Topics — the 22 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.322024
SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts · MICRO 2024
Ten Lessons From Three Generations Shaped Google's TPUv4i : Industrial Product · ISCA 2021
Memory systems
memory hierarchy
0.812024
SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts · MICRO 2024
Hardware accelerators and domain-specific architectures › machine learning accelerator
inference accelerator
0.512021
Ten Lessons From Three Generations Shaped Google's TPUv4i : Industrial Product · ISCA 2021
Processor architecture and microarchitecture › instruction-level parallelism
VLIW
0.512021
Ten Lessons From Three Generations Shaped Google's TPUv4i : Industrial Product · ISCA 2021
Memory systems
DRAM
0.412019
Context-Aware Resiliency: Unequal Message Protection for Random-Access Memories · IEEE Trans. Inf. Theory 2019
Hardware reliability and fault tolerance
error-correcting codes for memory
0.412019
Context-Aware Resiliency: Unequal Message Protection for Random-Access Memories · IEEE Trans. Inf. Theory 2019
Coding theory
error-correcting codes
0.412019
Context-Aware Resiliency: Unequal Message Protection for Random-Access Memories · IEEE Trans. Inf. Theory 2019
Memory systems › cache › cache technology
fault-tolerant cache
0.322015
Power / Capacity Scaling: Energy Savings With Simple Fault-Tolerant Caches · DAC 2014
DPCS: Dynamic Power/Capacity Scaling for SRAM Caches in the Nanoscale Era · ACM Trans. Archit. Code Optim. 2015
Reconfigurable computing and FPGAs › reconfigurable computing
reconfigurable dataflow
0.212024
SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts · MICRO 2024
Memory systems
cache
0.212015
DPCS: Dynamic Power/Capacity Scaling for SRAM Caches in the Nanoscale Era · ACM Trans. Archit. Code Optim. 2015
Energy-efficient computing › power management › memory power management
cache energy reduction
0.212015
DPCS: Dynamic Power/Capacity Scaling for SRAM Caches in the Nanoscale Era · ACM Trans. Archit. Code Optim. 2015
Energy-efficient computing › power management › memory power management
DRAM power reduction
0.212015
ViPZonE: Hardware Power Variability-Aware Virtual Memory Management for Energy Savings · IEEE Trans. Computers 2015
Energy-efficient computing › voltage scaling
dynamic voltage scaling
0.212015
DPCS: Dynamic Power/Capacity Scaling for SRAM Caches in the Nanoscale Era · ACM Trans. Archit. Code Optim. 2015
Memory systems › cache › cache technology
SRAM cache
0.212015
DPCS: Dynamic Power/Capacity Scaling for SRAM Caches in the Nanoscale Era · ACM Trans. Archit. Code Optim. 2015
Memory systems
virtual memory management
0.212015
ViPZonE: Hardware Power Variability-Aware Virtual Memory Management for Energy Savings · IEEE Trans. Computers 2015
Memory systems
cache design
0.212014
Power / Capacity Scaling: Energy Savings With Simple Fault-Tolerant Caches · DAC 2014
Hardware reliability and fault tolerance
memory fault tolerance
0.212014
Multi-Layer Memory Resiliency · DAC 2014
Energy-efficient computing
power gating
0.212014
Power / Capacity Scaling: Energy Savings With Simple Fault-Tolerant Caches · DAC 2014
Memory systems
non-volatile memory
0.112015
ViPZonE: Hardware Power Variability-Aware Virtual Memory Management for Energy Savings · IEEE Trans. Computers 2015
Hardware reliability and fault tolerance
process variation
0.112015
DPCS: Dynamic Power/Capacity Scaling for SRAM Caches in the Nanoscale Era · ACM Trans. Archit. Code Optim. 2015
Energy-efficient computing
power management
0.112014
Power / Capacity Scaling: Energy Savings With Simple Fault-Tolerant Caches · DAC 2014
Energy-efficient computing
voltage scaling
0.112014
Power / Capacity Scaling: Energy Savings With Simple Fault-Tolerant Caches · DAC 2014

Methods — techniques the papers use, named apart from their topics

streaming dataflow · 0.8expert composition · 0.8combinatorial bounds · 0.8extended-hamming codes · 0.4extended hamming code · 0.4simulation · 0.2opportunistic computing · 0.2application annotations · 0.2voltage scaling · 0.2fault disabling · 0.2aging mitigation · 0.2
YearPublicationVenuePosition
2024 SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts
abstract
Monolithic large language models (LLMs) like GPT-4 have paved the way for modern generative AI applications. Training, serving, and maintaining monolithic LLMs at scale, however, remains prohibitively expensive and challenging. The disproportionate increase in compute-to-memory ratio of modern AI accelerators have created a memory wall, necessitating new methods to deploy AI. Recent research has shown that a composition of many smaller expert models, each with several orders of magnitude fewer parameters, can match or exceed the capabilities of monolithic LLMs. Composition of Experts (CoE) is a modular approach that lowers the cost and complexity of training and serving. However, this approach presents two key challenges when using conventional hardware: (1) without fused operations, smaller models have lower operational intensity, which makes high utilization more challenging to achieve; and (2) hosting a large number of models can be either prohibitively expensive or slow when dynamically switching between them. In this paper, we describe how combining CoE, streaming dataflow, and a three-tier memory system scales the AI memory wall. We describe Samba-CoE, a CoE system with 150 experts and a trillion total parameters. We deploy Samba-CoE on the SambaNova SN40L Reconfigurable Dataflow Unit (RDU) -a commercial dataflow accelerator architecture that has been codesigned for enterprise inference and training applications. The chip introduces a new three-tier memory system with on-chip distributed SRAM, on-package HBM, and off-package DDR DRAM. A dedicated inter-RDU network enables scaling up and out over multiple sockets. We demonstrate speedups ranging from 2× to 13× on various benchmarks running on eight RDU sockets compared with an unfused baseline. We show that for CoE inference deployments, the 8-socket RDU Node reduces machine footprint by up to 19 ×, speeds up model switching time by 15× to 31×, and achieves an overall speedup of 3.7× over a DGX H100 and 6.6× over a DGX A100.
Raghu Prabhakar, Ram Sivaramakrishnan, Darshan Gandhi, Mingran Wang, Kejie Zhang, Tianren Gao, Angela Wang, Yongning Sheng, Joshua Brot, Denis Sokolov, Apurv Vivek, Calvin Leung, Arjun Sabnis, Jiayu Bai, Tuowen Zhao, Mark Gottscho, Mark Luttrell, Manish K. Shah, Zhengyu Chen 0002, Kaizhao Liang, Swayambhoo Jain, Urmish Thakker, Dawei Huang, Sumti Jairath, Kevin J. Brown, Kunle Olukotun
MICRO19
2021 Ten Lessons From Three Generations Shaped Google's TPUv4i : Industrial Product
abstract
Google deployed several TPU generations since 2015, teaching us lessons that changed our views: semi-conductor technology advances unequally; compiler compatibility trumps binary compatibility, especially for VLIW domain-specific architectures (DSA); target total cost of ownership vs initial cost; support multi-tenancy; deep neural networks (DNN) grow 1.5X annually; DNN advances evolve workloads; some inference tasks require floating point; inference DSAs need air-cooling; apps limit latency, not batch size; and backwards ML compatibility helps deploy DNNs quickly. These lessons molded TPUv4i, an inference DSA deployed since 2020.
Norman P. Jouppi, Doe Hyun Yoon, Matthew Ashcraft, Mark Gottscho, Thomas B. Jablin, George Kurian, James Laudon, Sheng Li 0007, Peter C. Ma, Thomas Norrie, Nishant Patil, Sushma Prasad, Cliff Young, Zongwei Zhou, David A. Patterson 0001
ISCA4
2019 Context-Aware Resiliency: Unequal Message Protection for Random-Access Memories
abstract
A common way to protect data stored in DRAM and related memory systems is through the use of an error-correcting code such as the extended Hamming code. Traditionally, these error-correcting codes provide equal protection guarantees to all messages. In this paper, we focus on unequal message protection (UMP), in which a subset of messages is deemed as special, and is afforded additional error-correction protection while maintaining the same number of redundancy bits as the baseline code. UMP is a powerful approach when the special messages are chosen based on the knowledge of data patterns in context. Our objective is to construct deterministic, algebraic codes with guaranteed UMP properties, derive their cardinality bounds using novel combinatorial techniques, and to demonstrate their efficacy for realistic memory benchmarks. We first introduce a UMP alternative to the single-bit parity-check code, and then we generalize to a broader UMP code family, including a UMP alternative to the extended Hamming code, offering full double-error correction protection to special messages. Our UMP constructions, applied to main memory in high-performance computing applications, could lead to significant system-level benefits such as less frequent checkpoints in supercomputers and decreased risk of catastrophic failure from erroneous special messages.
Clayton Schoeny, Frederic Sala, Mark Gottscho, Irina Alam, Puneet Gupta 0001, Lara Dolecek
IEEE Trans. Inf. Theory3
2018 Error Correction and Detection for Computing Memories Using System Side Information
abstract
Error correction and detection are the core components of all modern memory systems. Current computing memory systems use simple coding schemes to simultaneously meet the resiliency and latency requirements. In this paper, we review our recent results on context-aware coding for computing memories, an approach that explicitly takes into account various intrinsic side information for improved robustness to faults. We discuss both error correction and detection, codes' theoretical properties, and provide examples of how these solutions can be implemented in practice. We explicitly describe the special case of the error localization codes. We also discuss promising future directions and connections with classical information theoretic concepts.
Clayton Schoeny, Irina Alam, Mark Gottscho, Puneet Gupta 0001, Lara Dolecek
ITW3
2017 Context-aware resiliency: Unequal message protection for random-access memories
abstract
A common way to protect data stored in DRAM and related memory systems is through the use of a single-error-correcting/double-error-detecting (SECDED) code. Traditionally, these error-correcting codes provide equal protection guarantees to all messages. In a recent work, we demonstrated enhanced error correction capabilities for SECDED codes by taking into account contextual side-information about the data. This paper is concerned with a closely related scenario: unequal message protection (UMP), where a subset of special messages is afforded additional error-correction ability. UMP is relevant to computing systems where certain messages are critical and failures cannot be tolerated. We study practical UMP constructions where messages are guaranteed either one or two bit-error-correction. We provide upper and lower bounds on the number of special messages. We introduce an explicit and practical code construction based on BCH subcodes and demonstrate the efficacy of our technique on data from the AxBench and SPEC CPU2006 benchmark suites.
Clayton Schoeny, Frederic Sala, Mark Gottscho, Irina Alam, Puneet Gupta 0001, Lara Dolecek
ITW3
2017 Low-Cost Memory Fault Tolerance for IoT Devices
abstract
IoT devices need reliable hardware at low cost. It is challenging to efficiently cope with both hard and soft faults in embedded scratchpad memories. To address this problem, we propose a two-step approach: FaultLink and Software-Defined Error-Localizing Codes (SDELC). FaultLink avoids hard faults found during testing by generating a custom-tailored application binary image for each individual chip. During software deployment-time, FaultLink optimally packs small sections of program code and data into fault-free segments of the memory address space and generates a custom linker script for a lazy-linking procedure. During run-time, SDELC deals with unpredictable soft faults via novel and inexpensive Ultra-Lightweight Error-Localizing Codes (UL-ELCs). These require fewer parity bits than single-error-correcting Hamming codes. Yet our UL-ELCs are more powerful than basic single-error-detecting parity: they localize single-bit errors to a specific chunk of a codeword. SDELC then heuristically recovers from these localized errors using a small embedded C library that exploits observable side information (SI) about the application’s memory contents. SI can be in the form of redundant data (value locality), legal/illegal instructions, etc. Our combined FaultLink+SDELC approach improves min-VDD by up to 440 mV and correctly recovers from up to 90% (70%) of random single-bit soft faults in data (instructions) with just three parity bits per 32-bit word.
Mark Gottscho, Irina Alam, Clayton Schoeny, Lara Dolecek, Puneet Gupta 0001
ACM Trans. Embed. Comput. Syst.1
2016 Multi-story power distribution networks for GPUs
Liangzhen Lai, Mark Gottscho, Puneet Gupta 0001
DATE3
2016 X-Mem: A cross-platform and extensible memory characterization tool for the cloud
abstract
Effective use of the memory hierarchy is crucial to cloud computing. Platform memory subsystems must be carefully provisioned and configured to minimize overall cost and energy for cloud providers. For cloud subscribers, the diversity of available platforms complicates comparisons and the optimization of performance. To address these needs, we present X-Mem, a new open-source software tool that characterizes the memory hierarchy for cloud computing.
Mark Gottscho, Sriram Govindan, Bikash Sharma, Mohammed Shoaib, Puneet Gupta 0001
ISPASS1
2015 DPCS: Dynamic Power/Capacity Scaling for SRAM Caches in the Nanoscale Era
abstract
Fault-Tolerant Voltage-Scalable (FTVS) SRAM cache architectures are a promising approach to improve energy efficiency of memories in the presence of nanoscale process variation. Complex FTVS schemes are commonly proposed to achieve very low minimum supply voltages, but these can suffer from high overheads and thus do not always offer the best power/capacity trade-offs. We observe on our 45nm test chips that the “fault inclusion property” can enable lightweight fault maps that support multiple runtime supply voltages. Based on this observation, we propose a simple and low-overhead FTVS cache architecture for power/capacity scaling. Our mechanism combines multilevel voltage scaling with optional architectural support for power gating of blocks as they become faulty at low voltages. A static (SPCS) policy sets the runtime cache VDD once such that a only a few cache blocks may be faulty in order to minimize the impact on performance. We describe a Static Power/Capacity Scaling (SPCS) policy and two alternate Dynamic Power/Capacity Scaling (DPCS) policies that opportunistically reduce the cache voltage even further for more energy savings. This architecture achieves lower static power for all effective cache capacities than a recent more complex FTVS scheme. This is due to significantly lower overheads, despite the inability of our approach to match the min-VDD of the competing work at a fixed target yield. Over a set of SPEC CPU2006 benchmarks on two system configurations, the average total cache (system) energy saved by SPCS is 62% (22%), while the two DPCS policies achieve roughly similar energy reduction, around 79% (26%). On average, the DPCS approaches incur 2.24% performance and 6% area penalties.
Mark Gottscho, Abbas BanaiyanMofrad, Nikil Dutt, Alexandru Nicolau, Puneet Gupta 0001
ACM Trans. Archit. Code Optim.1
2015 ViPZonE: Hardware Power Variability-Aware Virtual Memory Management for Energy Savings
abstract
Hardware variability is predicted to increase dramatically over the coming years as a consequence of continued technology scaling. In this paper, we apply the Underdesigned and Opportunistic Computing (UnO) paradigm by exposing system-level power variability to software to improve energy efficiency. We present ViPZonE, a memory management solution in conjunction with application annotations that opportunistically performs memory allocations to reduce DRAM energy. ViPZonE's components consist of a physical address space with DIMM-aware zones, a modified page allocation routine, and a new virtual memory system call for dynamic allocations from userspace. We implemented ViPZonE in the Linux kernel with GLIBC API support, running on a real x86-64 testbed with significant access power variation in its DDR3 DIMMs. We demonstrate that on our testbed, ViPZonE can save up to 27.80 percent memory energy, with no more than 4.80 percent performance degradation across a set of PARSEC benchmarks tested with respect to the baseline Linux software. Furthermore, through a hypothetical “what-if” extension, we predict that in future non-volatile memory systems which consume almost no idle power, ViPZonE could yield even greater benefits, demonstrating the ability to exploit memory hardware variability through opportunistic software.
Mark Gottscho, Luis Angel D. Bathen, Nikil Dutt, Alexandru Nicolau, Puneet Gupta 0001
IEEE Trans. Computers1
2014 Multi-Layer Memory Resiliency
abstract
With memories continuing to dominate the area, power, cost and performance of a design, there is a critical need to provision reliable, high-performance memory bandwidth for emerging applications. Memories are susceptible to degradation and failures from a wide range of manufacturing, operational and environmental effects, requiring a multi-layer hardware/software approach that can tolerate, adapt and even opportunistically exploit such effects. The overall memory hierarchy is also highly vulnerable to the adverse effects of variability and operational stress. After reviewing the major memory degradation and failure modes, this paper describes the challenges for dependability across the memory hierarchy, and outlines research efforts to achieve multi-layer memory resilience using a hardware/software approach. Two specific exemplars are used to illustrate multilayer memory resilience: first we describe static and dynamic policies to achieve energy savings in caches using aggressive voltage scaling combined with disabling faulty blocks; and second we show how software characteristics can be exposed to the architecture in order to mitigate the aging of large register files in GPGPUs. These approaches can further benefit from semantic retention of application intent to enhance memory dependability across multiple abstraction levels, including applications, compilers, run-time systems, and hardware platforms.
Nikil Dutt, Puneet Gupta 0001, Alexandru Nicolau, Abbas BanaiyanMofrad, Mark Gottscho, Majid Namaki-Shoushtari
DAC5
2014 Power / Capacity Scaling: Energy Savings With Simple Fault-Tolerant Caches
abstract
Complicated approaches to fault-tolerant voltage-scalable (FTVS) SRAM cache architectures can suffer from high overheads. We propose static (SPCS) and dynamic (DPCS) variants of power/capacity scaling, a simple and low-overhead fault-tolerant cache architecture that utilizes insights gained from our 45nm SOI test chip. Our mechanism combines multi-level voltage scaling with power gating of blocks that become faulty at each voltage level. The SPCS policy sets the runtime cache VDD statically such that almost all of the cache blocks are not faulty. The DPCS policy opportunistically reduces the voltage further to save more power than SPCS while limiting the impact on performance caused by additional faulty blocks. Through an analytical evaluation, we show that our approach can achieve lower static power for all effective cache capacities than a recent complex FTVS work. This is due to significantly lower overheads, despite the failure of our approach to match the min-VDD of the competing work at fixed yield. Through architectural simulations, we find that the average energy saved by SPCS is 55%, while DPCS saves an average of 69% of energy with respect to baseline caches at 1 V. Our approach incurs no more than 4% performance and 5% area penalties in the worst case cache configuration.
Mark Gottscho, Abbas BanaiyanMofrad, Nikil Dutt, Alexandru Nicolau, Puneet Gupta 0001
DAC1
2013 Variability-aware memory management for nanoscale computing
abstract
As the semiconductor industry continues to push the limits of sub-micron technology, the ITRS expects hardware (e.g., die-to-die, wafer-to-wafer, and chip-to-chip) variations to continue increasing over the next few decades. As a result, it is imperative for designers to build variation-aware software stacks that may adapt and opportunistically exploit said variations to increase system performance/responsiveness as well as minimize power consumption. The memory subsystem is one of the largest components in today's computing system, a main contributor to the overall power consumption of the system, and therefore one of the most vulnerable components to the effects of variations (e.g., power). This paper discusses the concept of variability-aware memory management for nanoscale computing systems. We show how to opportunistically exploit the hardware variations in on-chip and off-chip memory at the system level through the deployment of variation-aware software stacks.
Nikil Dutt, Puneet Gupta 0001, Alexandru Nicolau, Luis Angel D. Bathen, Mark Gottscho
ASP-DAC5