EDBT 2026 Demo / reviewers in the wild / expert
David Nagle
dblp:55/1314
· DBLP profile ↗
19ranked-venue papers
2as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 14 · 2 first-authorSystems, architecture and hardware · 13 · 2 first-authorDatabases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
17 papers |
Distributed systems · 44% Storage systems · 33% Memory systems · 8% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 77% Transaction processing and concurrency control · 23% |
Topics — the 30 heaviest of 41, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems
distributed storage |
0.3 | 1 | 2018 | Sharding the Shards: Managing Datastore Locality at Scale with Akkio · OSDI 2018 |
Distributed systems
distributed database |
0.3 | 2 | 2013 | Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013 Spanner: Google's Globally-Distributed Database · OSDI 2012 |
Distributed systems
consensus |
0.2 | 1 | 2013 | Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013 |
Distributed systems
replication |
0.2 | 1 | 2013 | Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013 |
Distributed systems › replication › update propagation
synchronous replication |
0.2 | 1 | 2013 | Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.1 | 1 | 2018 | Sharding the Shards: Managing Datastore Locality at Scale with Akkio · OSDI 2018 |
Storage systems › storage devices
MEMS-based storage |
0.1 | 3 | 2000 | Modeling and performance of MEMS-based storage devices · SIGMETRICS 2000 Operating System Management of MEMS-based Storage Devices · OSDI 2000 Designing computer systems with MEMS-based storage · ASPLOS 2000 |
Storage systems › i/o scheduling
disk scheduling |
0.1 | 2 | 2000 | Data Mining on an OLTP System (Nearly) for Free · SIGMOD Conference 2000 Towards Higher Disk Head Utilization: Extracting "Free" Bandwidth from Busy Disk Drives · OSDI 2000 |
Memory systems
non-volatile memory |
0.1 | 2 | 2000 | Modeling and performance of MEMS-based storage devices · SIGMETRICS 2000 Designing computer systems with MEMS-based storage · ASPLOS 2000 |
Distributed systems › distributed coordination and fault tolerance
consensus and replication |
0.0 | 1 | 2012 | Spanner: Google's Globally-Distributed Database · OSDI 2012 |
Storage systems › file systems
distributed file system |
0.0 | 2 | 1998 | A Cost-Effective, High-Bandwidth Storage Architecture · ASPLOS 1998 File Server Scaling with Network-Attached Secure Disks · SIGMETRICS 1997 |
Storage systems › network-attached storage
network-attached secure disk |
0.0 | 2 | 1998 | A Cost-Effective, High-Bandwidth Storage Architecture · ASPLOS 1998 File Server Scaling with Network-Attached Secure Disks · SIGMETRICS 1997 |
Storage systems
network-attached storage |
0.0 | 2 | 1998 | A Cost-Effective, High-Bandwidth Storage Architecture · ASPLOS 1998 File Server Scaling with Network-Attached Secure Disks · SIGMETRICS 1997 |
Storage systems › magnetic storage
hard disk drive |
0.0 | 1 | 2000 | Towards Higher Disk Head Utilization: Extracting "Free" Bandwidth from Busy Disk Drives · OSDI 2000 |
Memory systems › memory management › virtual memory › address translation
TLB |
0.0 | 2 | 1994 | Design Tradeoffs for Software-Managed TLBs · ACM Trans. Comput. Syst. 1994 Design Tradeoffs for Software-Managed TLBs · ISCA 1993 |
Storage systems › file systems › distributed file system
file server scaling |
0.0 | 1 | 1997 | File Server Scaling with Network-Attached Secure Disks · SIGMETRICS 1997 |
Performance modeling and evaluation
simulation |
0.0 | 2 | 1994 | Trap-driven Simulation with Tapeworm II · ASPLOS 1994 Kernel-Based Memory Simulation · SIGMETRICS 1994 |
Operating systems › resource management › memory management
virtual memory |
0.0 | 2 | 1994 | Design Tradeoffs for Software-Managed TLBs · ISCA 1993 Design Tradeoffs for Software-Managed TLBs · ACM Trans. Comput. Syst. 1994 |
Processor architecture and microarchitecture
instruction fetch |
0.0 | 1 | 1995 | Instruction Fetching: Coping with Code Bloat · ISCA 1995 |
Memory systems › cache
prefetching |
0.0 | 1 | 1995 | Instruction Fetching: Coping with Code Bloat · ISCA 1995 |
Memory systems › memory hierarchy
cache and TLB effects |
0.0 | 1 | 1994 | Trap-driven Simulation with Tapeworm II · ASPLOS 1994 |
Memory systems
memory referencing behavior |
0.0 | 1 | 1994 | Trap-driven Simulation with Tapeworm II · ASPLOS 1994 |
Performance modeling and evaluation › simulation › architectural simulation
memory system simulation |
0.0 | 1 | 1994 | Kernel-Based Memory Simulation · SIGMETRICS 1994 |
Memory systems › memory management › on-chip memory management
on-chip memory allocation |
0.0 | 1 | 1994 | Optimal Allocation of On-Chip Memory for Multiple-API Operating Systems · ISCA 1994 |
Memory systems › virtual memory management
software TLB |
0.0 | 1 | 1994 | Design Tradeoffs for Software-Managed TLBs · ACM Trans. Comput. Syst. 1994 |
Memory systems › memory management
virtual memory |
0.0 | 1 | 1994 | Design Tradeoffs for Software-Managed TLBs · ACM Trans. Comput. Syst. 1994 |
Operating systems › resource management
storage management |
0.0 | 1 | 2000 | Operating System Management of MEMS-based Storage Devices · OSDI 2000 |
Storage systems
file systems |
0.0 | 1 | 2000 | Modeling and performance of MEMS-based storage devices · SIGMETRICS 2000 |
Memory systems
memory hierarchy |
0.0 | 1 | 2000 | Designing computer systems with MEMS-based storage · ASPLOS 2000 |
Performance modeling and evaluation
benchmarking |
0.0 | 1 | 1995 | Instruction Fetching: Coping with Code Bloat · ISCA 1995 |
Methods — techniques the papers use, named apart from their topics
sharding · 0.3truetime · 0.2hardware monitoring · 0.1trace-driven simulation · 0.1disk request scheduling · 0.1active disk environment · 0.1trap-driven simulation · 0.0simulation model · 0.0mechanics equation · 0.0design flow comparison · 0.0simulation · 0.0kernel-based analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Sharding the Shards: Managing Datastore Locality at Scale with Akkio
Muthukaruppan Annamalai, Kaushik Ravichandran 0003, Harish Srinivas, Igor Zinkovsky, Luning Pan, Tony Savor, David Nagle, Michael Stumm |
OSDI | 7 |
| 2013 | Spanner: Google's Globally Distributed DatabaseabstractSpanner is Google’s scalable, multiversion, globally distributed, and synchronously replicated database. It is the first system to distribute data at global scale and support externally-consistent distributed transactions. This article describes how Spanner is structured, its feature set, the rationale underlying various design decisions, and a novel time API that exposes clock uncertainty. This API and its implementation are critical to supporting external consistency and a variety of powerful features: nonblocking reads in the past, lock-free snapshot transactions, and atomic schema changes, across all of Spanner. James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost 0001, J. J. Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Peter Hochschild, Wilson C. Hsieh, Sebastian Kanthak, Eugene Kogan, Alexander Lloyd, Sergey Melnik 0001, David Mwaura, David Nagle, Sean Quinlan, Rajesh Rao, Lindsay Rolig, Yasushi Saito, Michal Szymaniak, Ruth Wang, Dale Woodford |
ACM Trans. Comput. Syst. | 18 |
| 2012 | Spanner: Google's Globally-Distributed Database
James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost 0001, J. J. Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Peter Hochschild, Wilson C. Hsieh, Sebastian Kanthak, Eugene Kogan, Alexander Lloyd, Sergey Melnik 0001, David Mwaura, David Nagle, Sean Quinlan, Rajesh Rao, Lindsay Rolig, Yasushi Saito, Michal Szymaniak, Ruth Wang, Dale Woodford |
OSDI | 18 |
| 2001 | Better Security via Smarter DevicesabstractThis white paper promotes a new approach to network security in which each individual device erects its own security perimeter and defends its own critical resources (e.g., network link or storage media). Together with conventional border defenses, such self-securing devices could provide a flexible infrastructure for dynamic prevention, detection, diagnosis, isolation, and repair of successful breaches in borders and device security perimeters. We overview the self-securing devices approach and the siege warfare analogy that inspired it. We also describe several examples of how different devices might be extended with embedded security functionality and outline some challenges of designing and managing self-securing devices. Gregory R. Ganger, David Nagle |
HotOS | 2 |
| 2000 | Designing computer systems with MEMS-based storageabstractFor decades the RAM-to-disk memory hierarchy gap has plagued computer architects. An exciting new storage technology based on microelectromechanical systems (MEMS) is poised to fill a large portion of this performance gap, significantly reduce system power consumption, and enable many new applications. This paper explores the system-level implications of integrating MEMS-based storage into the memory hierarchy. Results show that standalone MEMS-based storage reduces I/O stall times by 4-74X over disks and improves overall application runtimes by 1.9-4.4X. When used as on-board caches for disks, MEMS-based storage improves I/O response time by up to 3.5X. Further, the energy consumption of MEMS-based storage is 10-54X less than that of state-of-the-art low-power disk drives. The combination of the high-level physical characteristics of MEMS-based storage (small footprints, high shock tolerance) and the ability to directly integrate MEMS-based storage with processing leads to such new applications as portable gigabit storage systems and ubiquitous active storage nodes. Steven W. Schlosser, John Linwood Griffin, David Nagle, Gregory R. Ganger |
ASPLOS | 3 |
| 2000 | Operating System Management of MEMS-based Storage Devices
John Linwood Griffin, Steven W. Schlosser, Gregory R. Ganger, David Nagle |
OSDI | 4 |
| 2000 | Towards Higher Disk Head Utilization: Extracting "Free" Bandwidth from Busy Disk Drives
Christopher R. Lumb, Jiri Schindler, Gregory R. Ganger, David Nagle, Erik Riedel |
OSDI | 4 |
| 2000 | Modeling and performance of MEMS-based storage devicesabstractMEMS-based storage devices are seen by many as promising alternatives to disk drives. Fabricated using conventional CMOS processes, MEMS-based storage consists of thousands of small, mechanical probe tips that access gigabytes of high-density, nonvolatile magnetic storage. This paper takes a first step towards understanding the performance characteristics of these devices by mapping them onto a disk-like metaphor. Using simulation models based on the mechanics equations governing the devices' operation, this work explores how different physical characteristics (e.g., actuator forces and per-tip data rates) impact the design trade-offs and performance of MEMS-based storage. Overall results indicate that average access times for MEMS-based storage are 6.5 times faster than for a modern disk (1.5 ms vs. 9.7 ms). Results from filesystem and database bench-marks show that this improvement reduces application I/O stall times up to 70%, resulting in overall performance improvements of 3X. John Linwood Griffin, Steven W. Schlosser, Gregory R. Ganger, David Nagle |
SIGMETRICS | 4 |
| 2000 | Data Mining on an OLTP System (Nearly) for FreeabstractThis paper proposes a scheme for scheduling disk requests that takes advantage of the ability of high-level functions to operate directly at individual disk drives. We show that such a scheme makes it possible to support a Data Mining workload on an OLTP system almost for free: there is only a small impact on the throughput and response time of the existing workload. Specifically, we show that an OLTP system has the disk resources to consistently provide one third of its sequential bandwidth to a background Data Mining task with close to zero impact on OLTP throughput and response time at high transaction loads. At low transaction loads, we show much lower impact than observed in previous work. This means that a production OLTP system can be used for Data Mining tasks without the expense of a second dedicated system. Our scheme takes advantage of close interaction with the on-disk scheduler by reading blocks for the Data Mining workload as the disk head “passes over” them while satisfying demand blocks from the OLTP request stream. We show that this scheme provides a consistent level of throughput for the background workload even at very high foreground loads. Such a scheme is of most benefit in combination with an Active Disk environment that allows the background Data Mining application to also take advantage of the processing power and memory available directly on the disk drives. Erik Riedel, Christos Faloutsos, Gregory R. Ganger, David Nagle |
SIGMOD Conference | 4 |
| 2000 | Reducing power by optimizing the necessary precision/range of floating-point arithmeticabstractLow-power systems often find the power cost of floating-point (FP) hardware prohibitively expensive. This paper explores ways of reducing FP power consumption by minimizing the bitwidth representation of FP data. Analysis of several FP programs that manipulate low-resolution human sensory data shows that these programs suffer no loss of accuracy even with a significant reduction in bitwidth. Most FP programs in our benchmark suite maintain the same output even when the mantissa bitwidth is reduced by half. This FP bitwidth reduction can deliver a significant power saving through the use of a variable bitwidth FP unit. Our results show that up to 66% reduction in multiplier energy/operation can be achieved in the FP unit by this bitwidth reduction technique without sacrificing any program accuracy. J. Y. F. Tong, David Nagle, Rob A. Rutenbar |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1999 | Vertical Benchmarks for CADabstractVertical benchmarks are complex system designs represented at multiple levels of abstraction. More effective than componentbased CAD benchmarks, vertical benchmarks enable quantitative comparison of CAD techniques within or across design flows. This work describes the notion of vertical benchmarks and presents our benchmark, which is based on a commercial DSP, by comparing two alternative design flows. 2. Christopher Inacio, Herman Schmit, David Nagle, Andrew Ryan, Donald E. Thomas, Yingfai Tong, Ben Klass |
DAC | 3 |
| 1998 | A Cost-Effective, High-Bandwidth Storage ArchitectureabstractThis paper describes the Network-Attached Secure Disk (NASD) storage architecture, prototype implementations oj NASD drives, array management for our architecture, and three, filesystems built on our prototype. NASD provides scalable storage bandwidth without the cost of servers used primarily, for transferring data from peripheral networks (e.g. SCSI) to client networks (e.g. ethernet). Increasing datuset sizes, new attachment technologies, the convergence of peripheral and interprocessor switched networks, and the increased availability of on-drive transistors motivate and enable this new architecture. NASD is based on four main principles: direct transfer to clients, secure interfaces via cryptographic support, asynchronous non-critical-path oversight, and variably-sized data objects. Measurements of our prototype system show that these services can be cost-effectively integrated into a next generation disk drive ASK. End-to-end measurements of our prototype drive andfilesysterns suggest that NASD cun support conventional distributed filesystems without performance degradation. More importantly, we show scaluble bandwidth for NASD-specialized filesystems. Using a parallel data mining application, NASD drives deliver u linear scaling of 6.2 MB/s per clientdrive pair, tested with up to eight pairs in our lab. Garth A. Gibson, David Nagle, Khalil Amiri, Jeff Butler, Fay W. Chang, Howard Gobioff, Charles Hardin, Erik Riedel, David Rochberg, Jim Zelenka |
ASPLOS | 2 |
| 1997 | File Server Scaling with Network-Attached Secure DisksabstractBy providing direct data transfer between storage and client, network-attached storage devices have the potential to improve scalability for existing distributed file systems (by removing the server as a bottleneck) and bandwidth for new parallel and distributed file systems (through network striping and more efficient data paths). Together, these advantages influence a large enough fraction of the storage market to make commodity network-attached storage feasible. Realizing the technology's full potential requires careful consideration across a wide range of file system, networking and security issues. This paper contrasts two network-attached storage architectures---(1) Networked SCSI disks (NetSCSI) are network-attached storage devices with minimal changes from the familiar SCSI interface, while (2) Network-Attached Secure Disks (NASD) are drives that support independent client access to drive object services. To estimate the potential performance benefits of these architectures, we develop an analytic model and perform trace-driven replay experiments based on AFS and NFS traces. Our results suggest that NetSCSI can reduce file server load during a burst of NFS or AFS activity by about 30%. With the NASD architecture, server load (during burst activity) can be reduced by a factor of up to five for AFS and up to ten for NFS. Garth A. Gibson, David Nagle, Khalil Amiri, Fay W. Chang, Eugene M. Feinberg, Howard Gobioff, Chen Lee, Berend Ozceri, Erik Riedel, David Rochberg, Jim Zelenka |
SIGMETRICS | 2 |
| 1995 | Instruction Fetching: Coping with Code BloatabstractPrevious research has shown that the SPEC benchmarks achieve low miss ratios in relatively small instruction caches. This paper presents evidence that current software-development practices produce applications that exhibit substantially higher instruction-cache miss ratios than do the SPEC benchmarks. To represent these trends, we have assembled a collection of applications, called the Instruction Benchmark Suite (IBS), that provides a better test of instruction-cache performance. We discuss the rationale behind the design of IBS and characterize its behavior relative to the SPEC benchmark suite. Our analysis is based on trace-driven and trap-driven simulations and takes into full account both the application and operating-system components of the workloads.This paper then reexamines a collection of previously-proposed hardware mechanisms for improving instruction-fetch performance in the context of the IBS workloads. We study the impact of cache organization, transfer bandwidth, prefetching, and pipelined memory systems on machines that rely on the use of relatively small primary instruction caches to facilitate increased clock rates. We find that, although of little use for SPEC, the right combination of these techniques substantially benefits IBS. Even so, under IBS, a stubborn lower bound on the instruction-fetch CPI remains as an obstacle to improving overall processor performance. Richard Uhlig, David Nagle, Trevor N. Mudge, Stuart Sechrest, Joel S. Emer |
ISCA | 2 |
| 1994 | Trap-driven Simulation with Tapeworm IIabstractTapeworm II is a software-based simulation tool that evaluates the cache and TLB performance of multiple-task and operating system intensive workloads. Tapeworm resides in an OS kernel and causes a host machine's hardware to drive simulations with kernel traps instead of with address traces, as is conventionally done. This allows Tapeworm to quickly and accurately capture complete memory referencing behavior with a limited degradation in overall system performance. This paper compares trap-driven simulation, as implemented in Tapeworm, with the more common technique of trace-driven memory simulation with respect to speed, accuracy, portability and flexibility. Richard Uhlig, David Nagle, Trevor N. Mudge, Stuart Sechrest |
ASPLOS | 2 |
| 1994 | Optimal Allocation of On-Chip Memory for Multiple-API Operating SystemsabstractThe allocation of die area to different processor components is a central issue in the design of single-chip microprocessors. Chip area is occupied by both core execution logic, such as ALU and FPU datapaths, and memory structures, such as caches, TLBs, and write buffers. The authors focus on the allocation of die area to memory structures through a cost/benefit analysis. The cost of memory structures with different sizes and associativities is estimated by using an established area model for on-chip memory. The performance benefits of selecting a given structure are measured through a collection of methods including on-the-fly hardware monitoring, trace-driven simulation and kernel-based analysis. Special consideration is given to operating systems that support multiple application programming interfaces (APIs), a software trend that substantially affects on-chip memory allocation decisions. Results: Small adjustments in cache and TLB design parameters can significantly impact overall performance. Operating systems that support multiple APIs, such as Mach 3.0, increase the relative importance of on-chip instruction caches and TLBs when compared against single-API systems such as Ultrix.> David Nagle, Richard Uhlig, Trevor N. Mudge, Stuart Sechrest |
ISCA | 1 |
| 1994 | Kernel-Based Memory SimulationabstractNo abstract available. Richard Uhlig, David Nagle, Trevor N. Mudge, Stuart Sechrest |
SIGMETRICS | 2 |
| 1994 | Design Tradeoffs for Software-Managed TLBsabstractAn increasing number of architectures provide virtual memory support through software-managed TLBs. However, software management can impose considerable penalties that are highly dependent on the operating system's structure and its use of virtual memory. This work explores software-managed TLB design tradeoffs and their interaction with a range of monolithic and microkernel operating systems. Through hardware monitoring and simulation, we explore TLB performance for benchmarks running on a MIPS R2000-based workstation running Ultrix, OSF/1, and three versions of Mach 3.0. Richard Uhlig, David Nagle, Tim J. Stanley, Trevor N. Mudge, Stuart Sechrest, Richard B. Brown |
ACM Trans. Comput. Syst. | 2 |
| 1993 | Design Tradeoffs for Software-Managed TLBsabstractAn increasing number of architectures provide virtual memory support through software-managed TLBs. However, software management can impose considerable penalties, which are highly dependent on the operating system's structure and its use of virtual memory. This work explores software-managed TLB design tradeoffs and their interaction with a range of operating systems including monolithic and microkernel designs. Through hardware monitoring and simulations, we explore TLB performance for benchmarks running on a MIPS R2000-based workstation running Ultrix, OSF/1, and three versions of mach 3.0. David Nagle, Richard Uhlig, Tim J. Stanley, Stuart Sechrest, Trevor N. Mudge, Richard B. Brown |
ISCA | 1 |