Noah Watkins

dblp:34/10441 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 1 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Storage systems · 62% Cloud and datacenter computing · 24% Parallel and multicore computing · 7%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
cluster resource management and scheduling
0.322013
SIDR: structure-aware intelligent data routing in Hadoop · SC 2013
SciHadoop: array-based query processing in Hadoop · SC 2011
Storage systems › distributed storage
distributed shared log
0.312017
Malacology: A Programmable Storage System · EuroSys 2017
Storage systems
file systems
0.312017
Malacology: A Programmable Storage System · EuroSys 2017
Storage systems › file systems › distributed file system
metadata load balancing
0.312017
Malacology: A Programmable Storage System · EuroSys 2017
Storage systems › computational storage
programmable storage
0.312017
Malacology: A Programmable Storage System · EuroSys 2017
Storage systems › file systems
distributed file system
0.212015
Mantle: a programmable metadata load balancer for the ceph file system · SC 2015
Parallel and multicore computing
load balancing
0.212015
Mantle: a programmable metadata load balancer for the ceph file system · SC 2015
Storage systems
metadata management
0.212015
Mantle: a programmable metadata load balancer for the ceph file system · SC 2015
Storage systems
flash and SSD
0.212014
Flash on Rails: Consistent Flash Performance through Redundancy · USENIX ATC 2014
Cloud and datacenter computing › cloud data management
intermediate data partitioning
0.212013
SIDR: structure-aware intelligent data routing in Hadoop · SC 2013
Cloud and datacenter computing › cluster resource management and scheduling › cluster scheduling
mapreduce scheduling
0.212013
SIDR: structure-aware intelligent data routing in Hadoop · SC 2013
Query processing and optimization › runtime optimization › data skipping
partition pruning
0.112011
SciHadoop: array-based query processing in Hadoop · SC 2011
High-performance computing
scientific computing systems
0.122013
SIDR: structure-aware intelligent data routing in Hadoop · SC 2013
SciHadoop: array-based query processing in Hadoop · SC 2011
Cloud and datacenter computing
datacenter storage
0.112017
Malacology: A Programmable Storage System · EuroSys 2017
Storage systems › flash and SSD
flash memory
0.112014
Flash on Rails: Consistent Flash Performance through Redundancy · USENIX ATC 2014
Memory systems
non-volatile memory
0.112014
Flash on Rails: Consistent Flash Performance through Redundancy · USENIX ATC 2014

Methods — techniques the papers use, named apart from their topics

mapreduce · 0.2logical query specification · 0.2data dependency analysis · 0.2structure-aware partitioning · 0.2query-aware routing · 0.2
YearPublicationVenuePosition
2018 Tintenfisch: File System Namespace Schemas and Generators
Michael Sevilla, Reza Nasirigerdeh, Carlos Maltzahn, Jeff LeFevre, Noah Watkins, Peter Alvaro, Margaret Lawson, Jay F. Lofstead, James Pivarski
HotStorage5
2018 Cudele: An API and Framework for Programmable Consistency and Durability in a Global Namespace
abstract
HPC and data center scale application developers are abandoning POSIX IO because file system metadata synchronization and serialization overheads of providing strong consistency and durability are too costly - and often unnecessary - for their applications. Unfortunately, designing file systems with weaker consistency or durability semantics excludes applications that rely on stronger guarantees, forcing developers to re-write their applications or deploy them on a different system. We present a framework and API that lets administrators specify their consistency/durability requirements and dynamically assign them to subtrees in the same namespace, allowing administrators to optimize subtrees over time and space for different workloads. We show similar speedups to related work but more importantly, we show performance improvements when we custom fit subtree semantics to applications such as checkpoint-restart (91.7x speedup), user home directories (0.03 standard deviation from optimal), and users checking for partial results (2% overhead).
Michael Sevilla, Ivo Jimenez, Noah Watkins, Jeff LeFevre, Peter Alvaro, Shel Finkelstein, Patrick Donnelly, Carlos Maltzahn
IPDPS3
2018 quiho: Automated Performance Regression Testing Using Inferred Resource Utilization Profiles
abstract
We introduce quiho, a framework for profiling application performance that can be used in automated performance regression tests. quiho profiles an application by applying sensitivity analysis, in particular statistical regression analysis (SRA), using application-independent performance feature vectors that characterize the performance of machines. The result of the SRA, feature importance specifically, is used as a proxy to identify hardware and low-level system software behavior. The relative importance of these features serve as a performance profile of an application (termed inferred resource utilization profile or IRUP), which is used to automatically validate performance behavior across multiple revisions of an application»s code base without having to instrument code or obtain performance counters. We demonstrate that quiho can successfully discover performance regressions by showing its effectiveness in profiling application performance for synthetically introduced regressions as well as those found in real-world applications.
Ivo Jimenez, Noah Watkins, Michael Sevilla, Jay F. Lofstead, Carlos Maltzahn
ICPE2
2017 Malacology: A Programmable Storage System
abstract
Storage systems need to support high-performance for special-purpose data processing applications that run on an evolving storage device technology landscape. This puts tremendous pressure on storage systems to support rapid change both in terms of their interfaces and their performance. But adapting storage systems can be difficult because unprincipled changes might jeopardize years of code-hardening and performance optimization efforts that were necessary for users to entrust their data to the storage system. We introduce the programmable storage approach, which exposes internal services and abstractions of the storage stack as building blocks for higher-level services. We also build a prototype to explore how existing abstractions of common storage system services can be leveraged to adapt to the needs of new data processing systems and the increasing variety of storage devices. We illustrate the advantages and challenges of this approach by composing existing internal abstractions into two new higher-level services: a file system metadata load balancer and a high-performance distributed shared-log. The evaluation demonstrates that our services inherit desirable qualities of the back-end storage system, including the ability to balance load, efficiently propagate service metadata, recover from failure, and navigate trade-offs between latency and throughput using leases.
Michael Sevilla, Noah Watkins, Ivo Jimenez, Peter Alvaro, Shel Finkelstein, Jeff LeFevre, Carlos Maltzahn
EuroSys2
2017 Integrating External Resources with a Task-Based Programming Model
abstract
Accessing external resources (e.g., loading input data, checkpointing snapshots, and out-of-core processing) can have a significant impact on the performance of supercomputer applications. However, no existing programming systems for high-performance computing directly manage and optimize these external accesses. As a result, users must explicitly manage external accesses alongside their computation at the application level, which can result in both correctness and performance issues. We address this limitation by introducing Iris, a task-based programming model with semantics for external resources. Iris allows applications to describe their access requirements to external resources and the relationship of those accesses to the computation. Iris incorporates external I/O into a deferred execution model, reschedules external I/O to overlap I/O with computation, and reduces external I/O when possible. We evaluate Iris on three microbenchmarks representative of important workloads in HPC and a full combustion simulation, S3D. We demonstrate that the Iris implementation of S3D reduces the external I/O overhead by up to 20×, compared to the Legion and the Fortran implementations.
Sean Treichler, Galen M. Shipman, Michael Bauer 0001, Noah Watkins, Carlos Maltzahn, Patrick S. McCormick, Alex Aiken
HiPC5
2017 DeclStore: Layering Is for the Faint of Heart
Noah Watkins, Michael Sevilla, Ivo Jimenez, Kathryn Dahlgren, Peter Alvaro, Shel Finkelstein, Carlos Maltzahn
HotStorage1
2016 ZEA, A Data Management Approach for SMR
Adam Manzanares, Noah Watkins, Cyril Guyot, Damien Le Moal, Carlos Maltzahn, Zvonimir Bandic
HotStorage2
2015 Mantle: a programmable metadata load balancer for the ceph file system
abstract
Migrating resources is a useful tool for balancing load in a distributed system, but it is difficult to determine when to move resources, where to move resources, and how much of them to move. We look at resource migration for file system metadata and show how CephFS's dynamic subtree partitioning approach can exploit varying degrees of locality and balance because it can partition the namespace into variable sized units. Unfortunately, the current metadata balancer is complicated and difficult to control because it struggles to address many of the general resource migration challenges inherent to the metadata management problem. To help decouple policy from mechanism, we introduce a programmable storage system that lets the designer inject custom balancing logic. We show the flexibility and transparency of this approach by replicating the strategy of a state-of-the-art metadata balancer and conclude by comparing this strategy to other custom balancers on the same system.
Michael Sevilla, Noah Watkins, Carlos Maltzahn, Ike Nassi, Scott A. Brandt, Sage A. Weil, Greg Farnum, Samuel A. Fineberg
SC2
2014 Flash on Rails: Consistent Flash Performance through Redundancy
Dimitrios Skourtis, Dimitris Achlioptas, Noah Watkins, Carlos Maltzahn, Scott A. Brandt
USENIX ATC3
2013 SIDR: structure-aware intelligent data routing in Hadoop
abstract
The MapReduce framework is being extended for domains quite different from the web applications for which it was designed, including the processing of big structured data, e.g., scientific and financial data. Previous work using MapReduce to process scientific data ignores existing structure when assigning intermediate data and scheduling tasks. In this paper, we present a method for incorporating knowledge of the structure of scientific data and executing query into the MapReduce communication model. Built in SciHadoop, a version of the Hadoop MapReduce framework for scientific data, SIDR intelligently partitions and routes intermediate data, allowing it to: remove Hadoop's global barrier and execute Reduce tasks prior to all Map tasks completing; minimize intermediate key skew; and produce early, correct results. SIDR executes queries up to 2.5 times faster than Hadoop and 37% faster than SciHadoop; produces initial results with only 6% of the query completed; and produces dense, contiguous output.
Joe B. Buck, Noah Watkins, Greg Levin, Adam Crume, Kleoni Ioannidou, Scott A. Brandt, Carlos Maltzahn, Neoklis Polyzotis, Aaron Torres
SC2
2011 SciHadoop: array-based query processing in Hadoop
abstract
Hadoop has become the de facto platform for large-scale data analysis in commercial applications, and increasingly so in scientific applications. However, Hadoop's byte stream data model causes inefficiencies when used to process scientific data that is commonly stored in highly-structured, array-based binary file formats resulting in limited scalability of Hadoop applications in science. We introduce Sci-Hadoop, a Hadoop plugin allowing scientists to specify logical queries over array-based data models. Sci-Hadoop executes queries as map/reduce programs defined over the logical data model. We describe the implementation of a Sci-Hadoop prototype for NetCDF data sets and quantify the performance of five separate optimizations that address the following goals for several representative aggregate queries: reduce total data transfers, reduce remote reads, and reduce unnecessary reads. Two optimizations allow holistic aggregate queries to be evaluated opportunistically during the map phase; two additional optimizations intelligently partition input data to increase read locality, and one optimization avoids block scans by examining the data dependencies of an executing query to prune input partitions. Experiments involving a holistic function show run-time improvements of up to 8x, with drastic reductions of IO, both locally and over the network.
Joe B. Buck, Noah Watkins, Jeff LeFevre, Kleoni Ioannidou, Carlos Maltzahn, Neoklis Polyzotis, Scott A. Brandt
SC2