Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

David Nagle

dblp:55/1314 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 14 · 2 first-authorSystems, architecture and hardware · 13 · 2 first-authorDatabases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
17 papers
Distributed systems · 44% Storage systems · 33% Memory systems · 8%
Databases, data mining, and information retrieval
1 paper
Data mining · 77% Transaction processing and concurrency control · 23%

Topics — the 30 heaviest of 41, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems
distributed storage
0.312018
Sharding the Shards: Managing Datastore Locality at Scale with Akkio · OSDI 2018
Distributed systems
distributed database
0.322013
Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013
Spanner: Google's Globally-Distributed Database · OSDI 2012
Distributed systems
consensus
0.212013
Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013
Distributed systems
replication
0.212013
Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013
Distributed systems › replication › update propagation
synchronous replication
0.212013
Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013
Cloud and datacenter computing
cluster resource management and scheduling
0.112018
Sharding the Shards: Managing Datastore Locality at Scale with Akkio · OSDI 2018
Storage systems › storage devices
MEMS-based storage
0.132000
Modeling and performance of MEMS-based storage devices · SIGMETRICS 2000
Operating System Management of MEMS-based Storage Devices · OSDI 2000
Designing computer systems with MEMS-based storage · ASPLOS 2000
Storage systems › i/o scheduling
disk scheduling
0.122000
Data Mining on an OLTP System (Nearly) for Free · SIGMOD Conference 2000
Towards Higher Disk Head Utilization: Extracting "Free" Bandwidth from Busy Disk Drives · OSDI 2000
Memory systems
non-volatile memory
0.122000
Modeling and performance of MEMS-based storage devices · SIGMETRICS 2000
Designing computer systems with MEMS-based storage · ASPLOS 2000
Distributed systems › distributed coordination and fault tolerance
consensus and replication
0.012012
Spanner: Google's Globally-Distributed Database · OSDI 2012
Storage systems › file systems
distributed file system
0.021998
A Cost-Effective, High-Bandwidth Storage Architecture · ASPLOS 1998
File Server Scaling with Network-Attached Secure Disks · SIGMETRICS 1997
Storage systems › network-attached storage
network-attached secure disk
0.021998
A Cost-Effective, High-Bandwidth Storage Architecture · ASPLOS 1998
File Server Scaling with Network-Attached Secure Disks · SIGMETRICS 1997
Storage systems
network-attached storage
0.021998
A Cost-Effective, High-Bandwidth Storage Architecture · ASPLOS 1998
File Server Scaling with Network-Attached Secure Disks · SIGMETRICS 1997
Storage systems › magnetic storage
hard disk drive
0.012000
Towards Higher Disk Head Utilization: Extracting "Free" Bandwidth from Busy Disk Drives · OSDI 2000
Memory systems › memory management › virtual memory › address translation
TLB
0.021994
Design Tradeoffs for Software-Managed TLBs · ACM Trans. Comput. Syst. 1994
Design Tradeoffs for Software-Managed TLBs · ISCA 1993
Storage systems › file systems › distributed file system
file server scaling
0.011997
File Server Scaling with Network-Attached Secure Disks · SIGMETRICS 1997
Performance modeling and evaluation
simulation
0.021994
Trap-driven Simulation with Tapeworm II · ASPLOS 1994
Kernel-Based Memory Simulation · SIGMETRICS 1994
Operating systems › resource management › memory management
virtual memory
0.021994
Design Tradeoffs for Software-Managed TLBs · ISCA 1993
Design Tradeoffs for Software-Managed TLBs · ACM Trans. Comput. Syst. 1994
Processor architecture and microarchitecture
instruction fetch
0.011995
Instruction Fetching: Coping with Code Bloat · ISCA 1995
Memory systems › cache
prefetching
0.011995
Instruction Fetching: Coping with Code Bloat · ISCA 1995
Memory systems › memory hierarchy
cache and TLB effects
0.011994
Trap-driven Simulation with Tapeworm II · ASPLOS 1994
Memory systems
memory referencing behavior
0.011994
Trap-driven Simulation with Tapeworm II · ASPLOS 1994
Performance modeling and evaluation › simulation › architectural simulation
memory system simulation
0.011994
Kernel-Based Memory Simulation · SIGMETRICS 1994
Memory systems › memory management › on-chip memory management
on-chip memory allocation
0.011994
Optimal Allocation of On-Chip Memory for Multiple-API Operating Systems · ISCA 1994
Memory systems › virtual memory management
software TLB
0.011994
Design Tradeoffs for Software-Managed TLBs · ACM Trans. Comput. Syst. 1994
Memory systems › memory management
virtual memory
0.011994
Design Tradeoffs for Software-Managed TLBs · ACM Trans. Comput. Syst. 1994
Operating systems › resource management
storage management
0.012000
Operating System Management of MEMS-based Storage Devices · OSDI 2000
Storage systems
file systems
0.012000
Modeling and performance of MEMS-based storage devices · SIGMETRICS 2000
Memory systems
memory hierarchy
0.012000
Designing computer systems with MEMS-based storage · ASPLOS 2000
Performance modeling and evaluation
benchmarking
0.011995
Instruction Fetching: Coping with Code Bloat · ISCA 1995

Methods — techniques the papers use, named apart from their topics

sharding · 0.3truetime · 0.2hardware monitoring · 0.1trace-driven simulation · 0.1disk request scheduling · 0.1active disk environment · 0.1trap-driven simulation · 0.0simulation model · 0.0mechanics equation · 0.0design flow comparison · 0.0simulation · 0.0kernel-based analysis · 0.0
YearPublicationVenuePosition
2018 Sharding the Shards: Managing Datastore Locality at Scale with Akkio
Muthukaruppan Annamalai, Kaushik Ravichandran 0003, Harish Srinivas, Igor Zinkovsky, Luning Pan, Tony Savor, David Nagle, Michael Stumm
OSDI7
2013 Spanner: Google's Globally Distributed Database
abstract
Spanner is Google’s scalable, multiversion, globally distributed, and synchronously replicated database. It is the first system to distribute data at global scale and support externally-consistent distributed transactions. This article describes how Spanner is structured, its feature set, the rationale underlying various design decisions, and a novel time API that exposes clock uncertainty. This API and its implementation are critical to supporting external consistency and a variety of powerful features: nonblocking reads in the past, lock-free snapshot transactions, and atomic schema changes, across all of Spanner.
James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost 0001, J. J. Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Peter Hochschild, Wilson C. Hsieh, Sebastian Kanthak, Eugene Kogan, Alexander Lloyd, Sergey Melnik 0001, David Mwaura, David Nagle, Sean Quinlan, Rajesh Rao, Lindsay Rolig, Yasushi Saito, Michal Szymaniak, Ruth Wang, Dale Woodford
ACM Trans. Comput. Syst.18
2012 Spanner: Google's Globally-Distributed Database
James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost 0001, J. J. Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Peter Hochschild, Wilson C. Hsieh, Sebastian Kanthak, Eugene Kogan, Alexander Lloyd, Sergey Melnik 0001, David Mwaura, David Nagle, Sean Quinlan, Rajesh Rao, Lindsay Rolig, Yasushi Saito, Michal Szymaniak, Ruth Wang, Dale Woodford
OSDI18
2001 Better Security via Smarter Devices
abstract
This white paper promotes a new approach to network security in which each individual device erects its own security perimeter and defends its own critical resources (e.g., network link or storage media). Together with conventional border defenses, such self-securing devices could provide a flexible infrastructure for dynamic prevention, detection, diagnosis, isolation, and repair of successful breaches in borders and device security perimeters. We overview the self-securing devices approach and the siege warfare analogy that inspired it. We also describe several examples of how different devices might be extended with embedded security functionality and outline some challenges of designing and managing self-securing devices.
Gregory R. Ganger, David Nagle
HotOS2
2000 Designing computer systems with MEMS-based storage
abstract
For decades the RAM-to-disk memory hierarchy gap has plagued computer architects. An exciting new storage technology based on microelectromechanical systems (MEMS) is poised to fill a large portion of this performance gap, significantly reduce system power consumption, and enable many new applications. This paper explores the system-level implications of integrating MEMS-based storage into the memory hierarchy. Results show that standalone MEMS-based storage reduces I/O stall times by 4-74X over disks and improves overall application runtimes by 1.9-4.4X. When used as on-board caches for disks, MEMS-based storage improves I/O response time by up to 3.5X. Further, the energy consumption of MEMS-based storage is 10-54X less than that of state-of-the-art low-power disk drives. The combination of the high-level physical characteristics of MEMS-based storage (small footprints, high shock tolerance) and the ability to directly integrate MEMS-based storage with processing leads to such new applications as portable gigabit storage systems and ubiquitous active storage nodes.
Steven W. Schlosser, John Linwood Griffin, David Nagle, Gregory R. Ganger
ASPLOS3
2000 Operating System Management of MEMS-based Storage Devices
John Linwood Griffin, Steven W. Schlosser, Gregory R. Ganger, David Nagle
OSDI4
2000 Towards Higher Disk Head Utilization: Extracting "Free" Bandwidth from Busy Disk Drives
Christopher R. Lumb, Jiri Schindler, Gregory R. Ganger, David Nagle, Erik Riedel
OSDI4
2000 Modeling and performance of MEMS-based storage devices
abstract
MEMS-based storage devices are seen by many as promising alternatives to disk drives. Fabricated using conventional CMOS processes, MEMS-based storage consists of thousands of small, mechanical probe tips that access gigabytes of high-density, nonvolatile magnetic storage. This paper takes a first step towards understanding the performance characteristics of these devices by mapping them onto a disk-like metaphor. Using simulation models based on the mechanics equations governing the devices' operation, this work explores how different physical characteristics (e.g., actuator forces and per-tip data rates) impact the design trade-offs and performance of MEMS-based storage. Overall results indicate that average access times for MEMS-based storage are 6.5 times faster than for a modern disk (1.5 ms vs. 9.7 ms). Results from filesystem and database bench-marks show that this improvement reduces application I/O stall times up to 70%, resulting in overall performance improvements of 3X.
John Linwood Griffin, Steven W. Schlosser, Gregory R. Ganger, David Nagle
SIGMETRICS4
2000 Data Mining on an OLTP System (Nearly) for Free
abstract
This paper proposes a scheme for scheduling disk requests that takes advantage of the ability of high-level functions to operate directly at individual disk drives. We show that such a scheme makes it possible to support a Data Mining workload on an OLTP system almost for free: there is only a small impact on the throughput and response time of the existing workload. Specifically, we show that an OLTP system has the disk resources to consistently provide one third of its sequential bandwidth to a background Data Mining task with close to zero impact on OLTP throughput and response time at high transaction loads. At low transaction loads, we show much lower impact than observed in previous work. This means that a production OLTP system can be used for Data Mining tasks without the expense of a second dedicated system. Our scheme takes advantage of close interaction with the on-disk scheduler by reading blocks for the Data Mining workload as the disk head “passes over” them while satisfying demand blocks from the OLTP request stream. We show that this scheme provides a consistent level of throughput for the background workload even at very high foreground loads. Such a scheme is of most benefit in combination with an Active Disk environment that allows the background Data Mining application to also take advantage of the processing power and memory available directly on the disk drives.
Erik Riedel, Christos Faloutsos, Gregory R. Ganger, David Nagle
SIGMOD Conference4
2000 Reducing power by optimizing the necessary precision/range of floating-point arithmetic
abstract
Low-power systems often find the power cost of floating-point (FP) hardware prohibitively expensive. This paper explores ways of reducing FP power consumption by minimizing the bitwidth representation of FP data. Analysis of several FP programs that manipulate low-resolution human sensory data shows that these programs suffer no loss of accuracy even with a significant reduction in bitwidth. Most FP programs in our benchmark suite maintain the same output even when the mantissa bitwidth is reduced by half. This FP bitwidth reduction can deliver a significant power saving through the use of a variable bitwidth FP unit. Our results show that up to 66% reduction in multiplier energy/operation can be achieved in the FP unit by this bitwidth reduction technique without sacrificing any program accuracy.
J. Y. F. Tong, David Nagle, Rob A. Rutenbar
IEEE Trans. Very Large Scale Integr. Syst.2
1999 Vertical Benchmarks for CAD
abstract
Vertical benchmarks are complex system designs represented at multiple levels of abstraction. More effective than componentbased CAD benchmarks, vertical benchmarks enable quantitative comparison of CAD techniques within or across design flows. This work describes the notion of vertical benchmarks and presents our benchmark, which is based on a commercial DSP, by comparing two alternative design flows. 2.
Christopher Inacio, Herman Schmit, David Nagle, Andrew Ryan, Donald E. Thomas, Yingfai Tong, Ben Klass
DAC3
1998 A Cost-Effective, High-Bandwidth Storage Architecture
abstract
This paper describes the Network-Attached Secure Disk (NASD) storage architecture, prototype implementations oj NASD drives, array management for our architecture, and three, filesystems built on our prototype. NASD provides scalable storage bandwidth without the cost of servers used primarily, for transferring data from peripheral networks (e.g. SCSI) to client networks (e.g. ethernet). Increasing datuset sizes, new attachment technologies, the convergence of peripheral and interprocessor switched networks, and the increased availability of on-drive transistors motivate and enable this new architecture. NASD is based on four main principles: direct transfer to clients, secure interfaces via cryptographic support, asynchronous non-critical-path oversight, and variably-sized data objects. Measurements of our prototype system show that these services can be cost-effectively integrated into a next generation disk drive ASK. End-to-end measurements of our prototype drive andfilesysterns suggest that NASD cun support conventional distributed filesystems without performance degradation. More importantly, we show scaluble bandwidth for NASD-specialized filesystems. Using a parallel data mining application, NASD drives deliver u linear scaling of 6.2 MB/s per clientdrive pair, tested with up to eight pairs in our lab.
Garth A. Gibson, David Nagle, Khalil Amiri, Jeff Butler, Fay W. Chang, Howard Gobioff, Charles Hardin, Erik Riedel, David Rochberg, Jim Zelenka
ASPLOS2
1997 File Server Scaling with Network-Attached Secure Disks
abstract
By providing direct data transfer between storage and client, network-attached storage devices have the potential to improve scalability for existing distributed file systems (by removing the server as a bottleneck) and bandwidth for new parallel and distributed file systems (through network striping and more efficient data paths). Together, these advantages influence a large enough fraction of the storage market to make commodity network-attached storage feasible. Realizing the technology's full potential requires careful consideration across a wide range of file system, networking and security issues. This paper contrasts two network-attached storage architectures---(1) Networked SCSI disks (NetSCSI) are network-attached storage devices with minimal changes from the familiar SCSI interface, while (2) Network-Attached Secure Disks (NASD) are drives that support independent client access to drive object services. To estimate the potential performance benefits of these architectures, we develop an analytic model and perform trace-driven replay experiments based on AFS and NFS traces. Our results suggest that NetSCSI can reduce file server load during a burst of NFS or AFS activity by about 30%. With the NASD architecture, server load (during burst activity) can be reduced by a factor of up to five for AFS and up to ten for NFS.
Garth A. Gibson, David Nagle, Khalil Amiri, Fay W. Chang, Eugene M. Feinberg, Howard Gobioff, Chen Lee, Berend Ozceri, Erik Riedel, David Rochberg, Jim Zelenka
SIGMETRICS2
1995 Instruction Fetching: Coping with Code Bloat
abstract
Previous research has shown that the SPEC benchmarks achieve low miss ratios in relatively small instruction caches. This paper presents evidence that current software-development practices produce applications that exhibit substantially higher instruction-cache miss ratios than do the SPEC benchmarks. To represent these trends, we have assembled a collection of applications, called the Instruction Benchmark Suite (IBS), that provides a better test of instruction-cache performance. We discuss the rationale behind the design of IBS and characterize its behavior relative to the SPEC benchmark suite. Our analysis is based on trace-driven and trap-driven simulations and takes into full account both the application and operating-system components of the workloads.This paper then reexamines a collection of previously-proposed hardware mechanisms for improving instruction-fetch performance in the context of the IBS workloads. We study the impact of cache organization, transfer bandwidth, prefetching, and pipelined memory systems on machines that rely on the use of relatively small primary instruction caches to facilitate increased clock rates. We find that, although of little use for SPEC, the right combination of these techniques substantially benefits IBS. Even so, under IBS, a stubborn lower bound on the instruction-fetch CPI remains as an obstacle to improving overall processor performance.
Richard Uhlig, David Nagle, Trevor N. Mudge, Stuart Sechrest, Joel S. Emer
ISCA2
1994 Trap-driven Simulation with Tapeworm II
abstract
Tapeworm II is a software-based simulation tool that evaluates the cache and TLB performance of multiple-task and operating system intensive workloads. Tapeworm resides in an OS kernel and causes a host machine's hardware to drive simulations with kernel traps instead of with address traces, as is conventionally done. This allows Tapeworm to quickly and accurately capture complete memory referencing behavior with a limited degradation in overall system performance. This paper compares trap-driven simulation, as implemented in Tapeworm, with the more common technique of trace-driven memory simulation with respect to speed, accuracy, portability and flexibility.
Richard Uhlig, David Nagle, Trevor N. Mudge, Stuart Sechrest
ASPLOS2
1994 Optimal Allocation of On-Chip Memory for Multiple-API Operating Systems
abstract
The allocation of die area to different processor components is a central issue in the design of single-chip microprocessors. Chip area is occupied by both core execution logic, such as ALU and FPU datapaths, and memory structures, such as caches, TLBs, and write buffers. The authors focus on the allocation of die area to memory structures through a cost/benefit analysis. The cost of memory structures with different sizes and associativities is estimated by using an established area model for on-chip memory. The performance benefits of selecting a given structure are measured through a collection of methods including on-the-fly hardware monitoring, trace-driven simulation and kernel-based analysis. Special consideration is given to operating systems that support multiple application programming interfaces (APIs), a software trend that substantially affects on-chip memory allocation decisions. Results: Small adjustments in cache and TLB design parameters can significantly impact overall performance. Operating systems that support multiple APIs, such as Mach 3.0, increase the relative importance of on-chip instruction caches and TLBs when compared against single-API systems such as Ultrix.>
David Nagle, Richard Uhlig, Trevor N. Mudge, Stuart Sechrest
ISCA1
1994 Kernel-Based Memory Simulation
abstract
No abstract available.
Richard Uhlig, David Nagle, Trevor N. Mudge, Stuart Sechrest
SIGMETRICS2
1994 Design Tradeoffs for Software-Managed TLBs
abstract
An increasing number of architectures provide virtual memory support through software-managed TLBs. However, software management can impose considerable penalties that are highly dependent on the operating system's structure and its use of virtual memory. This work explores software-managed TLB design tradeoffs and their interaction with a range of monolithic and microkernel operating systems. Through hardware monitoring and simulation, we explore TLB performance for benchmarks running on a MIPS R2000-based workstation running Ultrix, OSF/1, and three versions of Mach 3.0.
Richard Uhlig, David Nagle, Tim J. Stanley, Trevor N. Mudge, Stuart Sechrest, Richard B. Brown
ACM Trans. Comput. Syst.2
1993 Design Tradeoffs for Software-Managed TLBs
abstract
An increasing number of architectures provide virtual memory support through software-managed TLBs. However, software management can impose considerable penalties, which are highly dependent on the operating system's structure and its use of virtual memory. This work explores software-managed TLB design tradeoffs and their interaction with a range of operating systems including monolithic and microkernel designs. Through hardware monitoring and simulations, we explore TLB performance for benchmarks running on a MIPS R2000-based workstation running Ultrix, OSF/1, and three versions of mach 3.0.
David Nagle, Richard Uhlig, Tim J. Stanley, Stuart Sechrest, Trevor N. Mudge, Richard B. Brown
ISCA1