EDBT 2026 Demo / reviewers in the wild / expert
Michael J. Brim
dblp:01/8357
· DBLP profile ↗
11ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-7479-5526ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FitCache: A Transparent Drop-In Framework for Multi-Tier Caching to Accelerate Distributed Deep Learning Workloads
Guangxing Hu, Awais Khan 0002, Christopher Zimmer 0001, Michael J. Brim, Frank Mueller 0001 |
IPDPS | 4 |
| 2025 | Integrating scientific single-page applications with DevSecOps
Lance Drane, Marshall T. McDonnell, Randall Petras, Cody Stiner, Arthur J. Ruckman, Gavin M. Wiggins, Gregory Cage, Seth Hitefield, Jesse McGaha, Andrew Ayres, Michael J. Brim, Rick Archibald, Addi Malviya-Thakur |
Future Gener. Comput. Syst. | 12 |
| 2025 | Lustre Unveiled: Evolution, Design, Advancements, and Current TrendsabstractThe Lustre filesystem serves as a vital element in high-performance parallel storage, meeting the rising demands of scientific, research, and enterprise environments. Widely deployed across HPC environments, ranging from small-scale applications in AI/ML, to domains like oil and gas, drug discovery, and meteorology, and manufacturing, Lustre addresses the universal challenge of efficiently accessing vast and ever-increasing volumes of data. Lustre is the filesystem of choice on six out of the top 10 fastest supercomputers in the world today, over 65% of the top 100, and also for over 60% of the top 500. Despite its widespread popularity, there is a lack of a complete and up-to-date reference, covering Lustre’s evolution, design, and various advancements made over the years. In this journal, we aim to fill this gap by providing a comprehensive journey of Lustre, including its history with significant contributions to HPC, detailed architecture and design elements, exploration of advancements added through its evolution, and future directions. Additionally, we present a comparison of Lustre with other prominent storage technologies of the era. To illustrate the current state of Lustre, we analyze several filesystem trends, including utilization, performance, and usage patterns on Orion, the Lustre filesystem on the first exascale supercomputer Frontier. We hope that this journal serves as a comprehensive educational reference for the current and future generations interested in HPC filesystem storage aspects. Anjus George, Andreas Dilger, Michael J. Brim, Rick Mohr, Amir Shehata, Jong Choi 0001, Ahmad Maroof Karimi, Jesse Hanley, James Simmons, Dominic Manno, Verónica G. Vergara Larrea, Sarp Oral, Christopher Zimmer 0001 |
ACM Trans. Storage | 3 |
| 2024 | Privacy Preserving Federated Learning for Advanced Scientific EcosystemsabstractWe present a framework to provide privacy preserving (PP) federating learning (FL) across multiple computational and experimental facilities. This work joins the compute capabilities of National Energy Research Scientific Computing Center (NERSC) and Oak Ridge National Laboratory Research Cloud (ORC) with simulated experimental data, such as those produced at the SLAC National Accelerator Laboratory and Spallation Neutron Source (SNS). We describe the software infrastructure developed to provide privacy for computational and experimental networks. We developed algorithmic privacy across the federated system by embedding database security, computation, and communication into the federation architecture, utilizing scientific tools developed by the experimental community. Rick Archibald, Addi Malviya-Thakur, Marshall T. McDonnell, Gregory Cage, Cody Stiner, Lance Drane, M. Paul Laiu, Michael J. Brim, Mathieu Doucet, William T. Heller, Ryan Coffee |
IEEE Big Data | 8 |
| 2023 | UnifyFS: A User-level Shared File System for Unified Access to Distributed Local StorageabstractWe introduce UnifyFS, a user-level file system that aggregates node-local storage tiers available on high performance computing (HPC) systems and makes them available to HPC applications under a unified namespace. UnifyFS employs transparent I/O interception, so it does not require changes to application code and is compatible with commonly used HPC I/O libraries. The design of UnifyFS supports the predominant HPC I/O workloads and is optimized for bulk-synchronous I/O patterns. Furthermore, UnifyFS provides customizable file system semantics to flexibly adapt its behavior for diverse I/O workloads and storage devices. In this paper, we discuss the unique design goals and architecture of UnifyFS and evaluate its performance on a leadership-class HPC system. In our experimental results, we demonstrate that UnifyFS exhibits excellent scaling performance for write operations and can improve the performance of application checkpoint operations by as much as 3× versus a tuned configuration. Michael J. Brim, Adam Moody, Seung-Hwan Lim, Ross G. Miller, Swen Böhm, Cameron Stanavige, Kathryn Mohror, Sarp Oral |
IPDPS | 1 |
| 2023 | Frontier: Exploring ExascaleabstractAs the US Department of Energy (DOE) computing facilities began deploying petascale systems in 2008, DOE was already setting its sights on exascale. In that year, DARPA published a report on the feasibility of reaching exascale. The report authors identified several key challenges in the pursuit of exascale including power, memory, concurrency, and resiliency. That report informed the DOE's computing strategy for reaching exascale. With the deployment of Oak Ridge National Laboratory's Frontier supercomputer, we have officially entered the exascale era. In this paper, we discuss Frontier's architecture, how it addresses those challenges, and describe some early application results from Oak Ridge Leadership Computing Facility's Center of Excellence and the Exascale Computing Project. Scott Atchley, Christopher Zimmer 0001, Jack Lange, David E. Bernholdt, Verónica G. Vergara Larrea, Michael J. Brim, Reuben D. Budiardja, Sunita Chandrasekaran, Markus Eisenbach 0002, Thomas M. Evans 0001, Matthew Ezell, Nicholas Frontiere, Antigoni Georgiadou, Joseph Glenski, Philipp Grete, Steven P. Hamilton, John K. Holmen, Axel Huebl, Daniel A. Jacobson, Wayne Joubert, Kim H. McMahon, Elia Merzari, Stan G. Moore, Andrew Myers 0001, Stephen Nichols, Sarp Oral, Thomas Papatheodore, Danny Perez, David M. Rogers 0001, Evan Schneider, Jean-Luc Vay, P. K. Yeung |
SC | 7 |
| 2019 | Are we witnessing the spectre of an HPC meltdown?abstractSummary We measure and analyze the performance observed when running applications and benchmarks before and after the Meltdown and Spectre fixes have been applied to the Cray supercomputers and supporting systems at the Oak Ridge Leadership Computing Facility (OLCF). Of particular interest is the effect of these fixes on applications selected from the OLCF portfolio when running at scale. This comprehensive study presents results from experiments run on Titan, Eos, Cumulus, and Percival supercomputers at the OLCF. The results from this study are useful for HPC users running on Cray supercomputers and serve to better understand the impact that these two vulnerabilities have on diverse HPC workloads at scale. Verónica G. Vergara Larrea, Michael J. Brim, Wayne Joubert, Swen Böhm, Matthew B. Baker, Oscar R. Hernandez, Sarp Oral, James Simmons, Don E. Maxwell |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | I/O load balancing for big data HPC applicationsabstractHigh Performance Computing (HPC) big data problems require efficient distributed storage systems. However, at scale, such storage systems often experience load imbalance and resource contention due to two factors: the bursty nature of scientific application I/O; and the complex I/O path that is without centralized arbitration and control. For example, the extant Lustre parallel file system-that supports many HPC centers-comprises numerous components connected via custom network topologies, and serves varying demands of a large number of users and applications. Consequently, some storage servers can be more loaded than others, which creates bottlenecks and reduces overall application I/O performance. Existing solutions typically focus on per application load balancing, and thus are not as effective given their lack of a global view of the system. In this paper, we propose a data-driven approach to load balance the I/O servers at scale, targeted at Lustre deployments. To this end, we design a global mapper on Lustre Metadata Server, which gathers runtime statistics from key storage components on the I/O path, and applies Markov chain modeling and a minimum-cost maximum-flow algorithm to decide where data should be placed. Evaluation using a realistic system simulator and a real setup shows that our approach yields better load balancing, which in turn can improve end-to-end performance. Arnab Kumar Paul, Arpit Goyal, Feiyi Wang, Sarp Oral, Ali Raza Butt, Michael J. Brim, Sangeetha B. Srinivasa |
IEEE BigData | 6 |
| 2013 | Efficient and Scalable Retrieval Techniques for Global File PropertiesabstractLarge-scale systems typically mount many different file systems with distinct performance characteristics and capacity. Applications must efficiently use this storage in order to realize their full performance potential. Users must take into account potential file replication throughout the storage hierarchy as well as contention in lower levels of the I/O system, and must consider communicating the results of file I/O between application processes to reduce file system accesses. Addressing these issues and optimizing file accesses requires detailed runtime knowledge of file system performance characteristics and the location(s) of files on them. In this paper, we propose Fast Global File Status (FGFS), a scalable mechanism to retrieve file information, such as its degree of distribution or replication and consistency. We use a novel node-local technique that turns expensive, non-scalable file system calls into simple string comparison operations. FGFS raises the namespace of a locally-defined file path to a global namespace with little or no file system calls to obtain global file properties efficiently. Our evaluation on a large multi-physics application shows that most FGFS file status queries on its executable and 848 shared library files complete in 272 milliseconds or faster at 32,768 MPI processes. Even the most expensive operation, which checks global file consistency, completes in under 7 seconds at this scale, an improvement of several orders of magnitude over the traditional checksum technique. Dong H. Ahn, Michael J. Brim, Bronis R. de Supinski, Todd Gamblin, Gregory L. Lee, Matthew P. LeGendre, Barton P. Miller, Adam Moody, Martin Schulz 0001 |
IPDPS | 2 |
| 2009 | Group file operations for scalable tools and middlewareabstractGroup file operations are a new, intuitive idiom for tools and middleware - including parallel debuggers and runtimes, performance measurement and steering, and distributed resource management - that require scalable operations on large groups of distributed files. The idiom provides new semantics for using file groups in standard file operations to eliminate costly iteration. A file-based idiom promotes conciseness and portability, and eases adoption. With explicit semantics for aggregation of group results, the idiom addresses a key scalability barrier. We have designed TBON-FS, a new distributed file system that provides scalable group file operations by leveraging tree-based overlay networks (TBONs) for scalable communication and data aggregation. We integrated group file operations into several tools: parallel versions of common utilities including cp, grep, rsync, tail, and top, and the Ganglia Distributed Monitoring System. Our experience verifies the group file operation idiom is intuitive, easily adopted, and enables a wide variety of tools to run efficiently at scale. Michael J. Brim, Barton P. Miller |
HiPC | 1 |
| 2001 | M3C: Managing and Monitoring Multiple ClustersabstractPC clusters running Linux, provide the computational power of supercomputers of just a few years ago at a fraction of the purchase price. The lack of good administration and user level application management tools makes the operation of clusters more difficult than it need be and results in an increased operation cost. This paper describes an ongoing effort at Oak Ridge National Laboratory to develop tools to simplify the administration and use of computation clusters. The two tools integrated in this work include the M3C (Managing and Monitoring Multiple Clusters) and C3 (Cluster Command and Control) tool suite. M3C provides a Web-based graphical user interface for cluster administration. It is designed as an extendable framework that can work with different underlying back-end tools. C3 is an independent project that provides the underlying commands to effect an operation on a cluster. In this paper it serves as an example how a back-end tool can be integrated into the M3C framework. Michael J. Brim, Al Geist, Brian Luethke, Jens Schwidder, Stephen L. Scott |
CCGRID | 1 |