VLDB 2026 Research / reviewers in the wild / expert
Eric A. Perlman
dblp:93/2783
· DBLP profile ↗
8ranked-venue papers
4as first author
0since 2021 · last 2018
0000-0001-5542-1302ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
High-performance computing · 63% Cloud and datacenter computing · 18% Storage systems · 18% | |
| Databases, data mining, and information retrieval
3 papers |
Spatial and temporal data management · 65% Distributed and cloud data management · 26% Query processing and optimization · 9% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Computational science and engineering · 100% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing
scientific data management |
0.2 | 2 | 2011 | I/O streaming evaluation of batch queries for data-intensive computational turbulence · SC 2011 JAWS: Job-Aware Workload Scheduling for the Exploration of Turbulence Simulations · SC 2010 |
Spatial and temporal data management
spatial indexing |
0.2 | 2 | 2008 | Organizing and indexing non-convex regions · Proc. VLDB Endow. 2008 Data exploration of turbulence simulations using a database cluster · SC 2007 |
Computational science and engineering › computational fluid dynamics
turbulence simulation |
0.1 | 2 | 2007 | Data exploration of turbulence simulations using a database cluster · SC 2007 Poster reception - Engineering the 100 terabyte turbulence database (or how to track particles at home) · SC 2006 |
High-performance computing
streaming i/o |
0.1 | 1 | 2011 | I/O streaming evaluation of batch queries for data-intensive computational turbulence · SC 2011 |
Storage systems
i/o optimization |
0.1 | 1 | 2010 | JAWS: Job-Aware Workload Scheduling for the Exploration of Turbulence Simulations · SC 2010 |
Cloud and datacenter computing › job scheduling
query scheduling |
0.1 | 1 | 2010 | JAWS: Job-Aware Workload Scheduling for the Exploration of Turbulence Simulations · SC 2010 |
Spatial and temporal data management
spatial query processing |
0.1 | 1 | 2008 | Organizing and indexing non-convex regions · Proc. VLDB Endow. 2008 |
Distributed and cloud data management › distributed database architecture
database cluster |
0.1 | 1 | 2007 | Data exploration of turbulence simulations using a database cluster · SC 2007 |
Query processing and optimization › query execution
batch query processing |
0.0 | 1 | 2010 | JAWS: Job-Aware Workload Scheduling for the Exploration of Turbulence Simulations · SC 2010 |
High-performance computing
data-intensive computing |
0.0 | 1 | 2006 | Poster reception - Engineering the 100 terabyte turbulence database (or how to track particles at home) · SC 2006 |
Methods — techniques the papers use, named apart from their topics
partial sums · 0.2distributed query evaluation · 0.2workload-aware batching · 0.2adaptive scheduling · 0.2data partitioning · 0.1cache-sensitive scheduling · 0.1spatial indexing · 0.1load balancing · 0.1computational geometry · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Building NDStore Through Hierarchical Storage Management and Microservice ProcessingabstractWe describe NDStore, a scalable multi-hierarchical data storage deployment for spatial analysis of neuroscience data on the AWS cloud. The system design is inspired by the requirement to maintain high I/O throughput for workloads that build neural connectivity maps of the brain from peta-scale imaging data using computer vision algorithms. We store all our data on the AWS object store S3 to limit our deployment costs. S3 serves as our base-tier of storage. Redis, an in-memory key-value engine, is used as our caching tier. The data is dynamically moved between the different storage tiers based on user access. All programming interfaces to this system are RESTful web-services. We include a performance evaluation that shows that our production system provides good performance for a variety of workloads by combining the assets of multiple cloud services. Kunal Lillaney, Dean Kleissas, Alexander Eusman, Eric A. Perlman, William R. Gray Roncal, Joshua T. Vogelstein, Randal C. Burns |
eScience | 4 |
| 2013 | The open connectome project data cluster: scalable analysis and vision for high-throughput neuroscienceabstract- neural connectivity maps of the brain-using the parallel execution of computer vision algorithms on high-performance compute clusters. These services and open-science data sets are publicly available at openconnecto.me. The system design inherits much from NoSQL scale-out and data-intensive computing architectures. We distribute data to cluster nodes by partitioning a spatial index. We direct I/O to different systems-reads to parallel disk arrays and writes to solid-state storage-to avoid I/O interference and maximize throughput. All programming interfaces are RESTful Web services, which are simple and stateless, improving scalability and usability. We include a performance evaluation of the production system, highlighting the effec-tiveness of spatial data organization. Randal C. Burns, Kunal Lillaney, Daniel R. Berger, Logan Grosenick, Karl Deisseroth, R. Clay Reid, William R. Gray Roncal, Priya Manavalan, Davi Bock, Narayanan Kasthuri, Michael M. Kazhdan, Stephen J. Smith, Dean Kleissas, Eric A. Perlman, Kwanghun Chung, Nicholas C. Weiler, Jeff Lichtman, Alex Szalay, Joshua T. Vogelstein, R. Jacob Vogelstein |
SSDBM | 14 |
| 2011 | I/O streaming evaluation of batch queries for data-intensive computational turbulenceabstractWe describe a method for evaluating computational turbulence queries, including Lagrange Polynomial interpolation, based on partial sums that allows the underlying data to be accessed in any order and in parts. We exploit these properties to stream data from disk in a single pass and concurrently evaluate batch queries. The combination of sequential I/O and data sharing improves performance by an order of magnitude when compared with direct evaluation of each query. The technique also supports distributed evaluation of queries in a database cluster, assembling the partial sums from each node at the query mediator. Interpolation is fundamental to computational turbulence, over 95% of queries use these routines, and the partial sums method allows the JHU Turbulence Database Cluster to realize scale and throughput for our scientists' data-intensive workloads. Kalin Kanov, Eric A. Perlman, Randal C. Burns, Yanif Ahmad, Alex Szalay |
SC | 2 |
| 2010 | JAWS: Job-Aware Workload Scheduling for the Exploration of Turbulence SimulationsabstractWe present JAWS, a job-aware, data-driven batch scheduler that improves query throughput for data-intensive scientific database clusters. As datasets reach petabyte-scale, workloads that scan through vast amounts of data to extract features are gaining importance in the sciences. However, acute performance bottlenecks result when multiple queries execute simultaneously and compete for I/O resources. Our solution, JAWS, divides queries into I/O-friendly sub-queries for scheduling. It then identifies overlapping data requirements within the workload and executes sub-queries in batches to maximize data sharing and reduce redundant I/O. JAWS extends our previous work by supporting workflows in which queries exhibit data dependencies, exploiting workload knowledge to coordinate caching decisions, and combating starvation through adaptive and incremental trade-offs between query throughput and response time. Instrumenting JAWS in the Turbulence Database Cluster yields nearly three-fold improvement in query throughput when contention in the workload is high. Eric A. Perlman, Randal C. Burns, Tanu Malik, Tamás Budavári, Charles Meneveau, Alex Szalay |
SC | 2 |
| 2010 | Organization of Data in Non-convex Spatial Domains
Eric A. Perlman, Randal C. Burns, Michael M. Kazhdan, Rebecca R. Murphy, William P. Ball, Nina Amenta |
SSDBM | 1 |
| 2008 | Organizing and indexing non-convex regionsabstractWe demonstrate data indexing and query processing techniques that improve the efficiency of comparing, correlating, and joining data contained in non-convex regions. We use computational geometry techniques to automatically characterize the region of space from which data are drawn, partition the region based on that characterization, and create an index from the partitions. Our motivating application performs distributed data analysis queries among federated database sites that store scientific data sets from the Chesapeake Bay. Our preliminary findings indicate that these techniques often reduce the number of I/Os needed to serve a query by a factor of five---depending on the geometry of the query region. Eric A. Perlman, Randal C. Burns, Michael M. Kazhdan |
Proc. VLDB Endow. | 1 |
| 2007 | Data exploration of turbulence simulations using a database clusterabstractWe describe a new environment for the exploration of turbulent flows that uses a cluster of databases to store complete histories of Direct Numerical Simulation (DNS) results. This allows for spatial and temporal exploration of high-resolution data that were traditionally too large to store and too computationally expensive to produce on demand. We perform analysis of these data directly on the databases nodes, which minimizes the volume of network traffic. The low network demands enable us to provide public access to this experimental platform and its datasets through Web services. This paper details the system design and implementation. Specifically, we focus on hierarchical spatial indexing, cache-sensitive spatial scheduling of batch workloads, localizing computation through data partitioning, and load balancing techniques that minimize data movement. We provide real examples of how scientists use the system to perform high-resolution turbulence research from standard desktop computing environments. Eric A. Perlman, Randal C. Burns, Charles Meneveau |
SC | 1 |
| 2006 | Poster reception - Engineering the 100 terabyte turbulence database (or how to track particles at home)abstractWe describe a new environment for large-scale turbulence simulations that uses a cluster of database nodes to store the complete space-time history of fluid velocities. This allows for rapid access to high resolution data that were traditionally too large to store and too computationally expensive to produce on demand.We perform the actual experimental analysis inside the database nodes, which allows for data-intensive computations to be performed across a large number of nodes with relatively little network traffic.We currently have a limited-scale prototype system running actual turbulence simulations and are in the process of establishing a production cluster with high-resolution data. We will discuss our design choices and initial results with load balancing a data-intensive, migratory workload. Eric A. Perlman, Randal C. Burns |
SC | 1 |