EDBT 2026 Demo / reviewers in the wild / expert
Robert T. McLay
dblp:70/5619
· DBLP profile ↗
5ranked-venue papers
2as first author
0since 2021 · last 2014
0000-0002-9446-5639ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Storage systems · 36% High-performance computing · 33% Performance modeling and evaluation · 14% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › file systems › distributed file system
parallel file system |
0.2 | 1 | 2014 | A User-Friendly Approach for Tuning Parallel File Operations · SC 2014 |
High-performance computing
parallel i/o |
0.2 | 1 | 2014 | A User-Friendly Approach for Tuning Parallel File Operations · SC 2014 |
Storage systems › i/o optimization
parallel i/o optimization |
0.2 | 1 | 2014 | A User-Friendly Approach for Tuning Parallel File Operations · SC 2014 |
Distributed systems
resource monitoring |
0.2 | 1 | 2013 | Enabling comprehensive data-driven system management for large computational facilities · SC 2013 |
Performance modeling and evaluation
workload characterization |
0.2 | 1 | 2013 | Enabling comprehensive data-driven system management for large computational facilities · SC 2013 |
Storage systems › file systems › distributed file system › parallel file system
lustre file system |
0.1 | 1 | 2014 | A User-Friendly Approach for Tuning Parallel File Operations · SC 2014 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.0 | 1 | 2013 | Enabling comprehensive data-driven system management for large computational facilities · SC 2013 |
High-performance computing › scientific computing systems
computational fluid dynamics |
0.0 | 1 | 1997 | MPP Solution of Rayleigh - Bénard - Marangoni Flows · SC 1997 |
High-performance computing
domain decomposition |
0.0 | 1 | 1997 | MPP Solution of Rayleigh - Bénard - Marangoni Flows · SC 1997 |
High-performance computing
scientific computing systems |
0.0 | 1 | 1997 | MPP Solution of Rayleigh - Bénard - Marangoni Flows · SC 1997 |
Methods — techniques the papers use, named apart from their topics
performance modeling · 0.2auto-tuning library · 0.2predictive analytics · 0.2log integration · 0.2parallel iterative solvers · 0.0overlapping communication and computation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | A User-Friendly Approach for Tuning Parallel File OperationsabstractThe Lustre file system provides high aggregated I/O bandwidth and is in widespread use throughout the HPC community. Here we report on work (1) developing a model for understanding collective parallel MPI write operations on Lustre, and (2) producing a library that optimizes parallel write performance in a user-friendly way. We note that a system's default stripe count is rarely a good choice for parallel I/O, and that performance depends on a delicate balance between the number of stripes and the actual (not requested) number of collective writers. Unfortunate combinations of these parameters may degrade performance considerably. For the programmer, however, it's all about the stripe count: an informed choice of this single parameter allows MPI to assign writers in a way that achieves near-optimal performance. We offer recommendations for those who wish to tune performance manually and describe the easy-to-use T3PIO library that manages the tuning automatically. Robert T. McLay, Doug James, Si Liu 0008, John Cazes, William L. Barth |
SC | 1 |
| 2013 | Enabling comprehensive data-driven system management for large computational facilitiesabstractThis paper presents a tool chain, based on the open source tool TACC_Stats, for systematic and comprehensive job level resource use measurement for large cluster computers, and its incorporation into XDMoD, a reporting and analytics framework for resource management that targets meeting the information needs of users, application developers, systems administrators, systems management and funding managers. Accounting, scheduler and event logs are integrated with system performance data from TACC_Stats. TACC_Stats periodically records resource use including many hardware counters for each job running on each node. Furthermore, system level metrics are obtained through aggregation of the node (job) level data. Analysis of this data generates many types of standard and custom reports and even a limited predictive capability that has not previously been available for open-source, Linux-based software systems. This paper presents case studies of information that can be applied for effective resource management. We believe this system to be the first fully comprehensive system for supporting the information needs of all stakeholders in open-source software based HPC systems. James C. Browne, Robert L. DeLeon, Charng-Da Lu, Matthew D. Jones, Steven M. Gallo, Amin Ghadersohi, Abani K. Patra, William L. Barth, John L. Hammond, Thomas R. Furlani, Robert T. McLay |
SC | 11 |
| 2011 | Design and Evaluation of Network Topology-/Speed- Aware Broadcast Algorithms for InfiniBand ClustersabstractIt is an established fact that the network topology can have an impact on the performance of scientific parallel applications. However, little work has been done to design an easy to use solution inside a communication library supporting a parallel programming model where the complexities of making the application performance network topology agnostic is hidden from the end user. Similarly, the rapid improvements in networking technology and speed are resulting in many commodity clusters becoming heterogeneous, with respect to networking speed. For example, switches and adapters belonging to different generations (SDR - 8 Gbps, DDR - 16 Gbps and QDR - 36 Gbps speeds in InfiniBand) are integrated into a single system. This leads to an additional challenge to make the communication library aware of the performance implications of heterogeneous link speeds. Accordingly, the communication library can perform optimizations taking link speed into account. In this paper, we propose a framework to automatically detect the topology and speed of an InfiniBand network and make it available to users through an easy to use interface. We also make design changes inside the MPI library to dynamically query this topology detection service and to form a topology model of the underlying network. We have redesigned the broadcast algorithm to take into account this network topology information and dynamically adapt the communication pattern to best fit the characteristics of the underlying network. To the best of our knowledge, this is the first such work for InfiniBand clusters. Our experimental results show that, for large homogeneous systems and large message sizes, we get up to 14% improvement in the latency of the broadcast operation using our proposed network topology-aware scheme over the default scheme at the micro-benchmark level. At the application level, the proposed framework delivers up to 8% improvement in total application run-time especially as job size scales up. The proposed network speed-aware algorithms are able to attain micro-benchmark performance on the heterogeneous SDR-DDR InfiniBand cluster to perform on par with runs on the DDR only portion of the cluster for small to medium sized messages. We also demonstrate that the network speed aware algorithms perform 70% to 100% better than the naive algorithms when both are run on the heterogeneous SDR-DDR InfiniBand cluster. Hari Subramoni, Krishna Chaitanya Kandalla, Jérôme Vienne, Sayantan Sur, William L. Barth, Karen A. Tomko, Robert T. McLay, Karl W. Schulz, Dhabaleswar K. Panda 0001 |
CLUSTER | 7 |
| 1997 | MPP Solution of Rayleigh - Bénard - Marangoni FlowsabstractA domain decomposition strategy and parallel gradient-type iterative solution scheme have been developed and implemented for computation of complex 3D viscous flow problems involving heat transfer and surface tension effects. Special attention has been paid to the kernels for the computationally intensive matrix-vector products and dot products, to memory management, and to overlapping communication and computation. Details of these implementation issues are described together with associated performance and scalability studies. Representative Rayleigh- Bénard and microgravity Marangoni flow calculations on the Cray T3D are presented, and performance results verifying a sustained rate in excess of 16 gigaflops on 512 nodes of the T3D have been obtained. The work is currently being extended to the T3E and we have begun carrying out further performance benchmarks and scalability studies on this platform. Preliminary performance studies have recently been carried out and sustained rates above 50 gigaflops and 100 gigaflops have been achieved on the 512 node T3E-600 and 1024 node T3E-900 configurations respectively. Graham F. Carey, Christopher Harle, Robert T. McLay, Spencer Swift |
SC | 3 |
| 1996 | Maximizing Sparse Matrix--Vector Product Performance on RISC Based MIMD Computers
Robert T. McLay, Spencer Swift, Graham F. Carey |
J. Parallel Distributed Comput. | 1 |