VLDB 2026 Research / reviewers in the wild / expert
Ross G. Miller
dblp:270/3920
· DBLP profile ↗
7ranked-venue papers
0as first author
1since 2021 · last 2023
0000-0002-2179-495XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Storage systems · 60% High-performance computing · 14% Cloud and datacenter computing · 13% | |
| Databases, data mining, and information retrieval
1 paper |
Data integration and cleaning · 100% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › file systems › distributed file system
parallel file system |
0.7 | 3 | 2019 | End-to-end I/O portfolio for the summit supercomputing ecosystem · SC 2019 Best Practices and Lessons Learned from Deploying and Operating Large-Scale Data-Centric Parallel File Systems · SC 2014 Efficient Object Storage Journaling in a Distributed Parallel File System · FAST 2010 |
Storage systems
flash and SSD |
0.4 | 1 | 2019 | End-to-end I/O portfolio for the summit supercomputing ecosystem · SC 2019 |
High-performance computing
supercomputing |
0.3 | 2 | 2019 | Best Practices and Lessons Learned from Deploying and Operating Large-Scale Data-Centric Parallel File Systems · SC 2014 End-to-end I/O portfolio for the summit supercomputing ecosystem · SC 2019 |
Cloud and datacenter computing
log analysis |
0.3 | 1 | 2017 | GUIDE: a scalable information directory service to collect, federate, and analyze logs for operational insights into a leadership HPC facility · SC 2017 |
Storage systems › i/o architecture
i/o subsystem |
0.1 | 1 | 2019 | End-to-end I/O portfolio for the summit supercomputing ecosystem · SC 2019 |
Storage systems › file systems
distributed file system |
0.1 | 1 | 2010 | Efficient Object Storage Journaling in a Distributed Parallel File System · FAST 2010 |
Storage systems
storage reliability |
0.1 | 1 | 2014 | Best Practices and Lessons Learned from Deploying and Operating Large-Scale Data-Centric Parallel File Systems · SC 2014 |
Methods — techniques the papers use, named apart from their topics
log federation · 0.6data warehousing · 0.6end-to-end i/o solution · 0.4data movement · 0.4technology evaluation · 0.2benchmarking · 0.2journaling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | UnifyFS: A User-level Shared File System for Unified Access to Distributed Local StorageabstractWe introduce UnifyFS, a user-level file system that aggregates node-local storage tiers available on high performance computing (HPC) systems and makes them available to HPC applications under a unified namespace. UnifyFS employs transparent I/O interception, so it does not require changes to application code and is compatible with commonly used HPC I/O libraries. The design of UnifyFS supports the predominant HPC I/O workloads and is optimized for bulk-synchronous I/O patterns. Furthermore, UnifyFS provides customizable file system semantics to flexibly adapt its behavior for diverse I/O workloads and storage devices. In this paper, we discuss the unique design goals and architecture of UnifyFS and evaluate its performance on a leadership-class HPC system. In our experimental results, we demonstrate that UnifyFS exhibits excellent scaling performance for write operations and can improve the performance of application checkpoint operations by as much as 3× versus a tuned configuration. Michael J. Brim, Adam Moody, Seung-Hwan Lim, Ross G. Miller, Swen Böhm, Cameron Stanavige, Kathryn Mohror, Sarp Oral |
IPDPS | 4 |
| 2020 | Understanding the Interplay between Hardware Errors and User Job Characteristics on the Titan SupercomputerabstractDesigning dependable supercomputers begins with an understanding of errors in real-world, large-scale systems. The Titan supercomputer at Oak Ridge National Laboratory provides a unique opportunity to investigate errors when an actual system is actively used by multiple concurrent users and workloads from diverse domains at varying scales. This study presents a thorough analysis of 6, 908, 497 hardware errors from 18, 688 compute nodes of Titan for 312, 215 user jobs over a 3-year time period. Through careful joining of two system logs – the Machine Check Architecture (MCA) log and the job scheduler log – we show the correlated pattern of hardware errors for each job and user, in addition to individual descriptive statistics of errors, jobs, and users. Since the majority of hardware errors are memory errors, this study also shows the importance of error correcting in memory systems. Seung-Hwan Lim, Ross G. Miller, Sudharshan S. Vazhkudai |
IPDPS | 2 |
| 2019 | End-to-end I/O portfolio for the summit supercomputing ecosystemabstractThe I/O subsystem for the Summit supercomputer, No. 1 on the Top500 list, and its ecosystem of analysis platforms is composed of two distinct layers, namely the in-system layer and the center-wide parallel file system layer (PFS), Spider 3. The in-system layer uses node-local SSDs and provides 26.7 TB/s for reads, 9.7 TB/s for writes, and 4.6 billion IOPS to Summit. The Spider 3 PFS layer uses IBM's Spectrum Scale™ and provides 2.5 TB/s and 2.6 million IOPS to Summit and other systems. While deploying them as two distinct layers was operationally efficient, it also presented usability challenges in terms of multiple mount points and lack of transparency in data movement. To address these challenges, we have developed novel end-to-end I/O solutions for the concerted use of the two storage layers. We present the I/O subsystem architecture, the end-to-end I/O solution space, their design considerations and our deployment experience. Sarp Oral, Sudharshan S. Vazhkudai, Feiyi Wang, Christopher Zimmer 0001, Christopher Brumgard, Jesse Hanley, George Markomanolis, Ross G. Miller, Dustin Leverman, Scott Atchley, Verónica G. Vergara Larrea |
SC | 8 |
| 2017 | GUIDE: a scalable information directory service to collect, federate, and analyze logs for operational insights into a leadership HPC facilityabstractIn this paper, we describe the GUIDE framework used to collect, federate, and analyze log data from the Oak Ridge Leadership Computing Facility (OLCF), and how we use that data to derive insights into facility operations. We collect system logs and extract monitoring data at every level of the various OLCF subsystems, and have developed a suite of pre-processing tools to make the raw data consumable. The cleansed logs are then ingested and federated into a central, scalable data warehouse, Splunk, that offers storage, indexing, querying, and visualization capabilities. We have further developed and deployed a set of tools to analyze these multiple disparate log streams in concert and derive operational insights. We describe our experience from developing and deploying the GUIDE infrastructure, and deriving valuable insights on the various subsystems, based on two years of operations in the production OLCF environment. Sudharshan S. Vazhkudai, Ross G. Miller, Devesh Tiwari, Christopher Zimmer 0001, Feiyi Wang, Sarp Oral, Raghul Gunasekaran, Deryl Steinert |
SC | 2 |
| 2014 | Accelerating Data Acquisition, Reduction, and Analysis at the Spallation Neutron SourceabstractORNL operates the world's brightest neutron source, the Spallation Neutron Source (SNS). Funded by the US DOE Office of Basic Energy Science, this national user facility hosts hundreds of scientists from around the world, providing a platform to enable break-through research in materials science, sustainable energy, and basic science. While the SNS provides scientists with advanced experimental instruments, the deluge of data generated from these instruments represents both a big data challenge and a big data opportunity. For example, instruments at the SNS can now generate multiple millions of neutron events per second providing unprecedented experiment fidelity but leaving the user with a dataset that cannot be processed and analyzed in a timely fashion using legacy techniques. To address this big data challenge, ORNL has developed a near real-time streaming data reduction and analysis infrastructure. The Accelerating Data Acquisition, Reduction, and Analysis (ADARA) system provides a live streaming data infrastructure based on a high-performance publish subscribe system, in situ data reduction, visualization, and analysis tools, and integration with a high-performance computing and data storage infrastructure. ADARA allows users of the SNS instruments to analyze their experiment as it is run and make changes to the experiment in real-time and visualize the results of these changes. In this paper we describe ADARA, provide a high-level architectural overview of the system, and present a set of use-cases and real-world demonstrations of the technology. Galen M. Shipman, Stuart I. Campbell, David Dillow, Mathieu Doucet, Jim Kohl, Garrett E. Granroth, Ross G. Miller, Dale Stansberry, Thomas Proffen, Russel Taylor |
eScience | 7 |
| 2014 | Best Practices and Lessons Learned from Deploying and Operating Large-Scale Data-Centric Parallel File SystemsabstractThe Oak Ridge Leadership Computing Facility (OLCF) has deployed multiple large-scale parallel file systems (PFS) to support its operations. During this process, OLCF acquired significant expertise in large-scale storage system design, file system software development, technology evaluation, benchmarking, procurement, deployment, and operational practices. Based on the lessons learned from each new PFS deployment, OLCF improved its operating procedures, and strategies. This paper provides an account of our experience and lessons learned in acquiring, deploying, and operating large-scale parallel file systems. We believe that these lessons will be useful to the wider HPC community. Sarp Oral, James Simmons, Jason Hill, Dustin Leverman, Feiyi Wang, Matthew Ezell, Ross G. Miller, Douglas Fuller, Raghul Gunasekaran, Youngjae Kim 0001, Saurabh Gupta 0002, Devesh Tiwari, Sudharshan S. Vazhkudai, James H. Rogers, David Dillow, Galen M. Shipman, Arthur S. Bland |
SC | 7 |
| 2010 | Efficient Object Storage Journaling in a Distributed Parallel File System
Sarp Oral, Feiyi Wang, David Dillow, Galen M. Shipman, Ross G. Miller, Oleg Drokin |
FAST | 5 |