EDBT 2026 Demo / reviewers in the wild / expert
Philip C. Roth
dblp:57/6814
· DBLP profile ↗
21ranked-venue papers
9as first author
1since 2021 · last 2023
0000-0001-9583-1103ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 9 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
High-performance computing · 46% Storage systems · 22% Performance modeling and evaluation · 16% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 21 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › supercomputing
exascale computing |
0.7 | 1 | 2023 | Experiences readying applications for Exascale · SC 2023 |
High-performance computing
performance optimization at scale |
0.7 | 1 | 2023 | Experiences readying applications for Exascale · SC 2023 |
High-performance computing
supercomputing |
0.7 | 1 | 2023 | Experiences readying applications for Exascale · SC 2023 |
Memory systems › memory management › virtual memory
address translation |
0.3 | 1 | 2018 | Exploiting Internal Parallelism for Address Translation in Solid-State Drives · ACM Trans. Storage 2018 |
Storage systems › flash and SSD › flash memory management
flash translation layer |
0.3 | 1 | 2018 | Exploiting Internal Parallelism for Address Translation in Solid-State Drives · ACM Trans. Storage 2018 |
Storage systems › flash and SSD
solid-state drive |
0.3 | 1 | 2018 | Exploiting Internal Parallelism for Address Translation in Solid-State Drives · ACM Trans. Storage 2018 |
Storage systems › flash and SSD › SSD architecture
SSD parallelism |
0.3 | 1 | 2018 | Exploiting Internal Parallelism for Address Translation in Solid-State Drives · ACM Trans. Storage 2018 |
Performance modeling and evaluation › workload characterization › parallel workload analysis
communication pattern analysis |
0.2 | 1 | 2015 | Automated Characterization of Parallel Application Communication Patterns · HPDC 2015 |
Performance modeling and evaluation
workload characterization |
0.2 | 1 | 2015 | Automated Characterization of Parallel Application Communication Patterns · HPDC 2015 |
Performance modeling and evaluation › parallel performance evaluation
scalable performance tools |
0.1 | 2 | 2006 | On-line automated performance diagnosis on thousands of processes · PPoPP 2006 MRNet: A Software-Based Multicast/Reduction Network for Scalable Tools · SC 2003 |
Memory systems › cache management
cache replacement |
0.1 | 1 | 2018 | Exploiting Internal Parallelism for Address Translation in Solid-State Drives · ACM Trans. Storage 2018 |
Memory systems › cache management › cache replacement
LRU |
0.1 | 1 | 2018 | Exploiting Internal Parallelism for Address Translation in Solid-State Drives · ACM Trans. Storage 2018 |
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2008 | Early evaluation of IBM BlueGene/P · SC 2008 |
High-performance computing › supercomputing
supercomputing systems |
0.1 | 1 | 2008 | Early evaluation of IBM BlueGene/P · SC 2008 |
Parallel and multicore computing › MPI
MPI program analysis |
0.1 | 1 | 2015 | Automated Characterization of Parallel Application Communication Patterns · HPDC 2015 |
Performance modeling and evaluation › performance diagnosis
automated performance diagnosis |
0.1 | 1 | 2006 | On-line automated performance diagnosis on thousands of processes · PPoPP 2006 |
Parallel and multicore computing › parallel architecture
large-scale parallel systems |
0.1 | 1 | 2006 | On-line automated performance diagnosis on thousands of processes · PPoPP 2006 |
Performance modeling and evaluation
performance diagnosis |
0.1 | 1 | 2006 | On-line automated performance diagnosis on thousands of processes · PPoPP 2006 |
High-performance computing
collective communication |
0.0 | 1 | 2003 | MRNet: A Software-Based Multicast/Reduction Network for Scalable Tools · SC 2003 |
Parallel and multicore computing
parallel programming runtimes |
0.0 | 1 | 2003 | MRNet: A Software-Based Multicast/Reduction Network for Scalable Tools · SC 2003 |
Distributed systems
fault tolerance |
0.0 | 1 | 2003 | MRNet: A Software-Based Multicast/Reduction Network for Scalable Tools · SC 2003 |
Methods — techniques the papers use, named apart from their topics
performance tuning · 1.3early access system evaluation · 1.3trace-driven simulation · 0.3performance modeling · 0.3search-based analysis · 0.2mpip profiling · 0.2microbenchmarks · 0.1application benchmarks · 0.1distributed diagnosis · 0.1data aggregation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Experiences readying applications for ExascaleabstractThe advent of Exascale computing invites an assessment of existing best practices for developing application readiness on the world's largest supercomputers. This work details observations from the last four years in preparing scientific applications to run on the Oak Ridge Leadership Computing Facility's (OLCF) Frontier system. This paper addresses a range of topics in software including programmability, tuning, and portability considerations that are key to moving applications from existing systems to future installations. A set of representative workloads provides case studies for general system and software testing. We evaluate the use of early access systems for development across several generations of hardware. Finally, we discuss how best practices were identified and disseminated to the community through a wide range of activities including user-guides and trainings. We conclude with recommendations for ensuring application readiness on future leadership computing systems. Nicholas Malaya, O. E. Bronson Messer, Joseph Glenski, Antigoni Georgiadou, Justin Lietz, Kalyana C. Gottiparthi, Marcus S. Day, Jackie Chen, Jon S. Rood, Lucas Esclapez, James B. White III, Gustav R. Jansen, Nicholas Curtis, Stephen Nichols, Jakub Kurzak, Noel Chalmers, Chip Freitag, Paul T. Bauman, Alessandro Fanfarillo, Reuben D. Budiardja, Thomas Papatheodore, Nicholas Frontiere, Damon McDougall, Matthew R. Norman, Sarat Sreepathi, Philip C. Roth, Dmytro Bykov, Noah Wolfe, Paul Mullowney, Markus Eisenbach 0002, Marc T. Henry de Frahan, Wayne Joubert |
SC | 26 |
| 2019 | Nested Workflows for Loosely Coupled HPC SimulationsabstractThe increasing complexity of modern scientific simulations has given rise to the notion of re-usability as a means to reduce development time and effort. One form of re-usability involves hierarchical modelling, where the code and artifacts that model a physical phenomenon are used as the building blocks for more complex coupled physical systems. This reusability mode allows for improvements in constituents sub-models to impact the overall fidelity of the entire system with minimal effort. This benefit, however, hinges upon the ability to provide a fairly loose coupling across model boundaries, where changes in any sub-model do not result in wholesale changes to other sub-models, or in the structure and details of the code for the entire physical system. In this paper, we present the design, implementation, and case studies for incorporating sub-workflows as building blocks in a framework for loosely coupled high performance simulations. We outline the issues involved in providing lightweight customization points for full workflows that allows their use unchanged in a nested-workflow setting, while providing each sub-workflow with a separate execution context that ensures non-interference with other parts of the simulation. We present several use cases that demonstrate the successful use of the proposed design and its flexibility in accommodating different customization and scaling requirements. Wael R. Elwasif, Ane Lasa, Philip C. Roth, Timothy Reed Younkin, Mark R. Cianciosa |
AICCSA | 3 |
| 2018 | Exploiting Internal Parallelism for Address Translation in Solid-State DrivesabstractSolid-state Drives (SSDs) have changed the landscape of storage systems and present a promising storage solution for data-intensive applications due to their low latency, high bandwidth, and low power consumption compared to traditional hard disk drives. SSDs achieve these desirable characteristics using internal parallelism —parallel access to multiple internal flash memory chips—and a Flash Translation Layer (FTL) that determines where data are stored on those chips so that they do not wear out prematurely. However, current state-of-the-art cache-based FTLs like the Demand-based Flash Translation Layer (DFTL) do not allow IO schedulers to take full advantage of internal parallelism, because they impose a tight coupling between the logical-to-physical address translation and the data access. To address this limitation, we introduce a new FTL design called Parallel-DFTL that works with the DFTL to decouple address translation operations from data accesses. Parallel-DFTL separates address translation and data access operations into different queues, allowing the SSD to use concurrent flash accesses for both types of operations. We also present a Parallel-LRU cache replacement algorithm to improve the concurrency of address translation operations. To compare Parallel-DFTL against existing FTL approaches, we present a Parallel-DFTL performance model and compare its predictions against those for DFTL and an ideal page-mapping approach. We also implemented the Parallel-DFTL approach in an SSD simulator using real device parameters, and used trace-driven simulation to evaluate Parallel-DFTL’s efficacy. Our evaluation results show that Parallel-DFTL improved the overall performance by up to 32% for the real IO workloads we tested, and by up to two orders of magnitude with synthetic test workloads. We also found that Parallel-DFTL is able to achieve reasonable performance with a very small cache size and that it provides the best benefit for those workloads with large request size or with high write ratio. Wei Xie 0017, Yong Chen 0001, Philip C. Roth |
ACM Trans. Storage | 3 |
| 2017 | Special Issue on Data-Intensive Scalable Computing Systems
Philip C. Roth, Shane Canon |
Parallel Comput. | 1 |
| 2017 | ASA-FTL: An adaptive separation aware flash translation layer for solid state drives
Wei Xie 0017, Yong Chen 0001, Philip C. Roth |
Parallel Comput. | 3 |
| 2016 | Parallel-DFTL: A Flash Translation Layer That Exploits Internal Parallelism in Solid State DrivesabstractSolid State Drives (SSDs) using flash memory storage technology present a promising storage solution for data-intensive applications due to their low latency, high bandwidth, and low power consumption compared to traditional hard disk drives. SSDs achieve these desirable characteristics using internal parallelism - parallel access to multiple internal flash memory chips - and a Flash Translation Layer (FTL) that determines where data is stored on those chips so that they do not wear out prematurely. Unfortunately, current state-of- the-art cache-based FTLs like the Demand-based Flash Translation Layer (DFTL) do not allow IO schedulers to take full advantage of internal parallelism because they impose a tight coupling between the logical-to-physical address translation and the data access. In this work, we propose an innovative IO scheduling policy called Parallel-DFTL that works with the DFTL to break the coupled address translation operations from data accesses. Parallel-DFTL schedules address translation and data access operations separately, allowing the SSD to use its flash access channel resources concurrently and fully for both types of operations. We present a performance model of FTL schemes that predicts the benefit of Parallel-DFTL against DFTL. We implemented our approach in an SSD simulator using real SSD device parameters, and used trace-driven simulation to evaluate its efficacy. Parallel-DFTL improved overall performance by up to 32% for the real IO workloads we tested, and up to two orders of magnitude for our synthetic test workloads. It is also found that Parallel-DFTL is able to achieve reasonable performance with a very small cache size. Wei Xie 0017, Yong Chen 0001, Philip C. Roth |
NAS | 3 |
| 2015 | Automated Characterization of Parallel Application Communication PatternsabstractA concise description of an application's communication pattern is often useful, for example, as an efficient way to communicate application behavior to a system vendor. Several existing performance analysis tools can capture aspects of an application's communication behavior such as which processes communicated with which others and the communications operations they used. However, a human with a high degree of expertise is still required to recognize and characterize common communication idioms within the performance data collected by those tools. To simplify this characterization for non-experts, we have developed an approach for automatically characterizing a MPI application's communication behavior. We use the mpiP profiling tool to collect information about an application's communication topology and message volume. We then use a post-mortem search-based analysis to compare the application's observed communication pattern against a library of common communication patterns. By comparing the result of the various search paths, our approach identifies the combination of patterns that best matches the observed behavior. To evaluate our approach, we applied it to a synthetic example communication matrix and communication matrices obtained from two scientific applications. We determined that our automated approach was highly effective in characterizing the communication patterns represented in the matrices. Philip C. Roth, Jeremy S. Meredith, Jeffrey S. Vetter |
HPDC | 1 |
| 2014 | Value influence analysis for message passing applicationsabstractPeople who develop, debug, and optimize applications are most effective when they understand how those applications function. Value influence tracking is an on-line code analysis approach that provides a data-centric perspective on how a value contributes to later computation. Early work on value influence tracking focused on single-process applications. Building upon this early work, we have designed support for performing value influence tracking analyses with applications that use common MPI point-to-point and collective communication operations. In this paper, we describe the design and implementation of an approach for propagating value influence data between the processes of an MPI application that uses these types of operations. To demonstrate and evaluate our approach, we present case studies of using our value influence tracking implementation with the Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) and the Model for Prediction Across Scales (MPAS) ocean climate model running on the Keeneland Initial Delivery System (KIDS) Linux cluster. We also discuss how to extend our approach to support MPI one-sided operations and non-blocking collective communication operations. Philip C. Roth, Jeremy S. Meredith |
ICS | 1 |
| 2014 | Guest Editors' introduction to the special issue on "DISCS-2013"
Philip C. Roth, Yong Chen 0001 |
Parallel Comput. | 1 |
| 2013 | Using pattern-models to guide SSD deployment for Big Data applications in HPC systemsabstractFlash-memory based Solid State Drives (SSDs) embrace higher performance and lower power consumption compared to traditional storage devices (HDDs). These benefits are needed in HPC systems, especially with the growing demand of supporting Big Data applications. In this paper, we study placement and deployment strategies of SSDs in HPC systems to maximize the performance improvement, given a practical fixed hardware budget constraint. We propose a pattern-model approach to guide SSD deployment for HPC systems through two steps; characterizing workload and mapping deployment strategy. The first step is responsible for characterizing the access patterns of the workload and the second step contributes the actual deployment recommendation for Parallel File System (PFS) configuration combining with an analytical model. We have carried out initial experimental tests and the results confirmed that the proposed approach can guide placement of SSDs in HPC systems for accelerating data accesses. Our research will be helpful in guiding designs and developments for Big Data applications in current and projected HPC systems including exascale systems. Philip C. Roth, Yong Chen 0001 |
IEEE BigData | 2 |
| 2012 | DOSAS: Mitigating the Resource Contention in Active Storage SystemsabstractActive storage provides an effective method to mitigate the I/O bottleneck problem of data intensive high performance computing applications. It can reduce the amount of data transferred as the application runs by moving appropriate computations close to the data. Prior research has achieved considerable progress in developing several active storage prototypes. However, existing studies have neglected the impact of resource contention when concurrent processes request IOoperations from the same storage node simultaneously, which happens frequently in practice. In this paper, we analyze the impact of resource contention on active storage systems. Motivated by our analysis, we propose a novel Dynamic Operation Scheduling Active Storage architecture to address the resource contention issue. It offloads the active processing operations dynamically between storage nodes and compute nodes according to the system environment. By evaluating our architecture, we observed that: (1) resource contention is a critical problem for active storage systems, (2) the proposed dynamic operation scheduling method mitigates the problem, and (3) the new active storage architecture outperforms existing active storage systems. Yong Chen 0001, Philip C. Roth |
CLUSTER | 3 |
| 2011 | Probabilistic Communication and I/O Tracing with Deterministic Replay at ScaleabstractWith today's petascale supercomputers, applications often exhibit low efficiency, such as poor communication and I/O performance, that can be diagnosed by analysis tools. However, these tools either produce extremely large trace files that complicate performance analysis, or sacrifice accuracy to collect high-level statistical information using crude averaging. This work contributes Scala-H-Trace, which features more aggressive trace compression than any previous approach, particularly for applications that do not show strict regularity in SPMD behavior. Scala-H-Trace uses histograms expressing the probabilistic distribution of arbitrary communication and I/O parameters to capture variations. Yet, where other tools fail to scale, Scala-H-Trace guarantees trace files of near constant size, even for variable communication and I/O patterns, producing trace files orders of magnitudes smaller than using prior approaches. We demonstrate the ability to collect traces of applications running on thousands of processors with the potential to scale well beyond this level. We further present the first approach to deterministically replay such probabilistic traces (a) without deadlocks and (b) in a manner closely resembling the original applications. Our results show either near constant sized traces or only sub-linear increases in trace file sizes irrespective of the number of nodes utilized. Even with the aggressively compressed histogram-based traces, our replay times are within 12% to 15% of the runtime of original codes. Such concise traces resembling the behavior of production-style codes closely and our approach of deterministic replay of probabilistic traces are without precedence. Xing Wu 0004, Karthik Vijayakumar, Frank Mueller 0001, Xiaosong Ma, Philip C. Roth |
ICPP | 5 |
| 2011 | LACIO: A New Collective I/O Strategy for Parallel I/O SystemsabstractParallel applications benefit considerably from the rapid advance of processor architectures and the available massive computational capability, but their performance suffers from large latency of I/O accesses. The poor I/O performance has been attributed as a critical cause of the low sustained performance of parallel systems. Collective I/O is widely considered a critical solution that exploits the correlation among I/O accesses from multiple processes of a parallel application and optimizes the I/O performance. However, the conventional collective I/O strategy makes the optimization decision based on the logical file layout to avoid multiple file system calls and does not take the physical data layout into consideration. On the other hand, the physical data layout in fact decides the actual I/O access locality and concurrency. In this study, we propose a new collective I/O strategy that is aware of the underlying physical data layout. We confirm that the new Layout-Aware Collective I/O (LACIO) improves the performance of current parallel I/O systems effectively with the help of noncontiguous file system calls. It holds promise in improving the I/O performance for parallel systems. Yong Chen 0001, Xian-He Sun, Rajeev Thakur, Philip C. Roth, William Gropp |
IPDPS | 4 |
| 2008 | Early evaluation of IBM BlueGene/PabstractBlueGene/P (BG/P) is the second generation BlueGene architecture from IBM, succeeding BlueGene/L (BG/L). BG/P is a system-on-a-chip (SoC) design that uses four PowerPC 450 cores operating at 850 MHz with a double precision, dual pipe floating point unit per core. These chips are connected with multiple interconnection networks including a 3-D torus, a global collective network, and a global barrier network. The design is intended to provide a highly scalable, physically dense system with relatively low power requirements per flop. In this paper, we report on our examination of BG/P, presented in the context of a set of important scientific applications, and as compared to other major large scale supercomputers in use today. Our investigation confirms that BG/P has good scalability with an expected lower performance per processor when compared to the Cray XT4's Opteron. We also find that BG/P uses very low power per floating point operation for certain kernels, yet it has less of a power advantage when considering science driven metrics for mission applications. Sadaf R. Alam, Richard F. Barrett, M. Bast, Mark R. Fahey, Jeffery A. Kuehn, Collin McCurdy, James H. Rogers, Philip C. Roth, Ramanan Sankaran, Jeffrey S. Vetter, Patrick H. Worley, Weikuan Yu |
SC | 8 |
| 2007 | Modeling the Impact of Checkpoints on Next-Generation Systems
Ron A. Oldfield, Sarala Arunagiri, Patricia J. Teller, Seetharami R. Seelam, Maria Ruiz Varela, Rolf Riesen, Philip C. Roth |
MSST | 7 |
| 2006 | Early evaluation of the Cray XT3abstractOak Ridge National Laboratory recently received delivery of a 5,294 processor Cray XT3. The XT3 is Cray's third-generation massively parallel processing system. The system builds on a single processor node - built around the AMD Opteron - and uses a custom chip - called SeaStar - to provide interprocess or communication. In addition, the system uses a lightweight operating system on the compute nodes. This paper describes our initial experiences with the system, including micro-benchmark, kernel, and application benchmark results. In particular, we provide performance results for strategic Department of Energy applications areas including climate and fusion. We demonstrate experiments on the installed system, scaling applications up to 4,096 processors. Jeffrey S. Vetter, Sadaf R. Alam, Thomas H. Dunigan, Mark R. Fahey, Philip C. Roth, Patrick H. Worley |
IPDPS | 5 |
| 2006 | On-line automated performance diagnosis on thousands of processesabstractPerformance analysis tools are critical for the effective use of large parallel computing resources, but existing tools have failed to address three problems that limit their scalability: (1) management and processing of the volume of performance data generated when monitoring a large number of application processes, (2) communication between a large number of tool components, and (3) presentation of performance data and analysis results for applications with a large number of processes. In this paper, we present a novel approach for finding performance problems in applications with a large number of processes that leverages our multicast and data aggregation infrastructure to address these three performance tool scalability barriers. First, we show how to design a scalable, distributed performance diagnosis facility. We demonstrate this design with an on-line, Philip C. Roth, Barton P. Miller |
PPoPP | 1 |
| 2004 | Benchmarking the MRNet Distributed Tool Infrastructure: Lessons LearnedabstractSummary form only given. MRNet is an infrastructure that provides scalable multicast and data aggregation functionality for distributed tools. While evaluating MRNet's performance and scalability, we learned several important lessons about benchmarking large-scale, distributed tools and middleware. First, automation is essential for a successful benchmarking effort, and should be leveraged whenever possible during the benchmarking process. Second, micro-benchmarking is invaluable not only for establishing the performance of low-level functionality, but also for design verification and debugging. Third, resource management systems need substantial improvements in their support for running tools and applications together. Finally, the most demanding experiments should be attempted early and often during a benchmarking effort to increase the chances of detecting problems with the tool and experimental methodology. Philip C. Roth, Dorian C. Arnold, Barton P. Miller |
IPDPS | 1 |
| 2003 | MRNet: A Software-Based Multicast/Reduction Network for Scalable ToolsabstractWe present MRNet, a software-based multicast/reduction network for building scalable performance and system administration tools. MRNet supports multiple simultaneous, asynchronous collective communication operations. MRNet is flexible, allowing tool builders to tailor its process network topology to suit their tool's requirements and the underlying system's capabilities. MRNet is extensible, allowing tool builders to incorporate custom data reductions to augment its collection of built-in reductions. We evaluated MRNet in a simple test tool and also integrated into an existing, real-world performance tool with up to 512 tool back-ends. In the real-world tool, we used MRNet not only for multicast and simple data reductions but also with custom histogram and clock skew detection reductions. In our experiments, the MRNet-based tools showed significantly better performance than the tools without MRNet for average message latency and throughput, overall tool start-up latency, and performance data processing throughput. Philip C. Roth, Dorian C. Arnold, Barton P. Miller |
SC | 1 |
| 2003 | Deep Start: a hybrid strategy for automated performance problem searchesabstractAbstract To attack the problem of scalability of performance diagnosis tools with respect to application code size, we have developed the Deep Start search strategy—a new technique that uses stack sampling to augment an automated search for application performance problems. Our hybrid approach locates performance problems more quickly and finds performance problems hidden from a more straightforward search strategy. The Deep Start strategy uses stack samples collected as a by‐product of normal search instrumentation to selectdeep starters, functions that are likely to be application bottlenecks. With priorities and careful control of the search refinement, our strategy gives preference to experiments on the deep starters and their callees. This approach enables the Deep Start strategy to find application bottlenecks more efficiently and more effectively than a more straightforward search strategy. We implemented the Deep Start search strategy in the Performance Consultant, Paradyn's automated bottleneck detection component. In our tests, Deep Start found half of our test applications' known bottlenecks between 32% and 59% faster than the Performance Consultant's current search strategy, and finished finding bottlenecks between 10% and 61% faster. In addition to improving the search time, Deep Start often found more bottlenecks than the call graph search strategy. Copyright © 2003 John Wiley & Sons, Ltd. Philip C. Roth, Barton P. Miller |
Concurr. Comput. Pract. Exp. | 1 |
| 2002 | Deep Start: A Hybrid Strategy for Automated Performance Problem Searches
Philip C. Roth, Barton P. Miller |
Euro-Par | 1 |