Eric A. Perlman

dblp:93/2783 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
0since 2021 · last 2018
0000-0001-5542-1302ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
High-performance computing · 63% Cloud and datacenter computing · 18% Storage systems · 18%
Databases, data mining, and information retrieval
3 papers
Spatial and temporal data management · 65% Distributed and cloud data management · 26% Query processing and optimization · 9%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Computational science and engineering · 100%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
scientific data management
0.222011
I/O streaming evaluation of batch queries for data-intensive computational turbulence · SC 2011
JAWS: Job-Aware Workload Scheduling for the Exploration of Turbulence Simulations · SC 2010
Spatial and temporal data management
spatial indexing
0.222008
Organizing and indexing non-convex regions · Proc. VLDB Endow. 2008
Data exploration of turbulence simulations using a database cluster · SC 2007
Computational science and engineering › computational fluid dynamics
turbulence simulation
0.122007
Data exploration of turbulence simulations using a database cluster · SC 2007
Poster reception - Engineering the 100 terabyte turbulence database (or how to track particles at home) · SC 2006
High-performance computing
streaming i/o
0.112011
I/O streaming evaluation of batch queries for data-intensive computational turbulence · SC 2011
Storage systems
i/o optimization
0.112010
JAWS: Job-Aware Workload Scheduling for the Exploration of Turbulence Simulations · SC 2010
Cloud and datacenter computing › job scheduling
query scheduling
0.112010
JAWS: Job-Aware Workload Scheduling for the Exploration of Turbulence Simulations · SC 2010
Spatial and temporal data management
spatial query processing
0.112008
Organizing and indexing non-convex regions · Proc. VLDB Endow. 2008
Distributed and cloud data management › distributed database architecture
database cluster
0.112007
Data exploration of turbulence simulations using a database cluster · SC 2007
Query processing and optimization › query execution
batch query processing
0.012010
JAWS: Job-Aware Workload Scheduling for the Exploration of Turbulence Simulations · SC 2010
High-performance computing
data-intensive computing
0.012006
Poster reception - Engineering the 100 terabyte turbulence database (or how to track particles at home) · SC 2006

Methods — techniques the papers use, named apart from their topics

partial sums · 0.2distributed query evaluation · 0.2workload-aware batching · 0.2adaptive scheduling · 0.2data partitioning · 0.1cache-sensitive scheduling · 0.1spatial indexing · 0.1load balancing · 0.1computational geometry · 0.1
YearPublicationVenuePosition
2018 Building NDStore Through Hierarchical Storage Management and Microservice Processing
abstract
We describe NDStore, a scalable multi-hierarchical data storage deployment for spatial analysis of neuroscience data on the AWS cloud. The system design is inspired by the requirement to maintain high I/O throughput for workloads that build neural connectivity maps of the brain from peta-scale imaging data using computer vision algorithms. We store all our data on the AWS object store S3 to limit our deployment costs. S3 serves as our base-tier of storage. Redis, an in-memory key-value engine, is used as our caching tier. The data is dynamically moved between the different storage tiers based on user access. All programming interfaces to this system are RESTful web-services. We include a performance evaluation that shows that our production system provides good performance for a variety of workloads by combining the assets of multiple cloud services.
Kunal Lillaney, Dean Kleissas, Alexander Eusman, Eric A. Perlman, William R. Gray Roncal, Joshua T. Vogelstein, Randal C. Burns
eScience4
2013 The open connectome project data cluster: scalable analysis and vision for high-throughput neuroscience
abstract
- neural connectivity maps of the brain-using the parallel execution of computer vision algorithms on high-performance compute clusters. These services and open-science data sets are publicly available at openconnecto.me. The system design inherits much from NoSQL scale-out and data-intensive computing architectures. We distribute data to cluster nodes by partitioning a spatial index. We direct I/O to different systems-reads to parallel disk arrays and writes to solid-state storage-to avoid I/O interference and maximize throughput. All programming interfaces are RESTful Web services, which are simple and stateless, improving scalability and usability. We include a performance evaluation of the production system, highlighting the effec-tiveness of spatial data organization.
Randal C. Burns, Kunal Lillaney, Daniel R. Berger, Logan Grosenick, Karl Deisseroth, R. Clay Reid, William R. Gray Roncal, Priya Manavalan, Davi Bock, Narayanan Kasthuri, Michael M. Kazhdan, Stephen J. Smith, Dean Kleissas, Eric A. Perlman, Kwanghun Chung, Nicholas C. Weiler, Jeff Lichtman, Alex Szalay, Joshua T. Vogelstein, R. Jacob Vogelstein
SSDBM14
2011 I/O streaming evaluation of batch queries for data-intensive computational turbulence
abstract
We describe a method for evaluating computational turbulence queries, including Lagrange Polynomial interpolation, based on partial sums that allows the underlying data to be accessed in any order and in parts. We exploit these properties to stream data from disk in a single pass and concurrently evaluate batch queries. The combination of sequential I/O and data sharing improves performance by an order of magnitude when compared with direct evaluation of each query. The technique also supports distributed evaluation of queries in a database cluster, assembling the partial sums from each node at the query mediator. Interpolation is fundamental to computational turbulence, over 95% of queries use these routines, and the partial sums method allows the JHU Turbulence Database Cluster to realize scale and throughput for our scientists' data-intensive workloads.
Kalin Kanov, Eric A. Perlman, Randal C. Burns, Yanif Ahmad, Alex Szalay
SC2
2010 JAWS: Job-Aware Workload Scheduling for the Exploration of Turbulence Simulations
abstract
We present JAWS, a job-aware, data-driven batch scheduler that improves query throughput for data-intensive scientific database clusters. As datasets reach petabyte-scale, workloads that scan through vast amounts of data to extract features are gaining importance in the sciences. However, acute performance bottlenecks result when multiple queries execute simultaneously and compete for I/O resources. Our solution, JAWS, divides queries into I/O-friendly sub-queries for scheduling. It then identifies overlapping data requirements within the workload and executes sub-queries in batches to maximize data sharing and reduce redundant I/O. JAWS extends our previous work by supporting workflows in which queries exhibit data dependencies, exploiting workload knowledge to coordinate caching decisions, and combating starvation through adaptive and incremental trade-offs between query throughput and response time. Instrumenting JAWS in the Turbulence Database Cluster yields nearly three-fold improvement in query throughput when contention in the workload is high.
Eric A. Perlman, Randal C. Burns, Tanu Malik, Tamás Budavári, Charles Meneveau, Alex Szalay
SC2
2010 Organization of Data in Non-convex Spatial Domains
Eric A. Perlman, Randal C. Burns, Michael M. Kazhdan, Rebecca R. Murphy, William P. Ball, Nina Amenta
SSDBM1
2008 Organizing and indexing non-convex regions
abstract
We demonstrate data indexing and query processing techniques that improve the efficiency of comparing, correlating, and joining data contained in non-convex regions. We use computational geometry techniques to automatically characterize the region of space from which data are drawn, partition the region based on that characterization, and create an index from the partitions. Our motivating application performs distributed data analysis queries among federated database sites that store scientific data sets from the Chesapeake Bay. Our preliminary findings indicate that these techniques often reduce the number of I/Os needed to serve a query by a factor of five---depending on the geometry of the query region.
Eric A. Perlman, Randal C. Burns, Michael M. Kazhdan
Proc. VLDB Endow.1
2007 Data exploration of turbulence simulations using a database cluster
abstract
We describe a new environment for the exploration of turbulent flows that uses a cluster of databases to store complete histories of Direct Numerical Simulation (DNS) results. This allows for spatial and temporal exploration of high-resolution data that were traditionally too large to store and too computationally expensive to produce on demand. We perform analysis of these data directly on the databases nodes, which minimizes the volume of network traffic. The low network demands enable us to provide public access to this experimental platform and its datasets through Web services. This paper details the system design and implementation. Specifically, we focus on hierarchical spatial indexing, cache-sensitive spatial scheduling of batch workloads, localizing computation through data partitioning, and load balancing techniques that minimize data movement. We provide real examples of how scientists use the system to perform high-resolution turbulence research from standard desktop computing environments.
Eric A. Perlman, Randal C. Burns, Charles Meneveau
SC1
2006 Poster reception - Engineering the 100 terabyte turbulence database (or how to track particles at home)
abstract
We describe a new environment for large-scale turbulence simulations that uses a cluster of database nodes to store the complete space-time history of fluid velocities. This allows for rapid access to high resolution data that were traditionally too large to store and too computationally expensive to produce on demand.We perform the actual experimental analysis inside the database nodes, which allows for data-intensive computations to be performed across a large number of nodes with relatively little network traffic.We currently have a limited-scale prototype system running actual turbulence simulations and are in the process of establishing a production cluster with high-resolution data. We will discuss our design choices and initial results with load balancing a data-intensive, migratory workload.
Eric A. Perlman, Randal C. Burns
SC1