EDBT 2026 Demo / reviewers in the wild / expert
William L. Barth
dblp:08/1485 · also Bill Barth, William Barth
· DBLP profile ↗
13ranked-venue papers
1as first author
0since 2021 · last 2017
0000-0003-3333-2952ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
High-performance computing · 40% Distributed systems · 23% Storage systems · 19% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 50% Rendering · 50% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems › communication optimization
communication-computation overlap |
0.2 | 1 | 2014 | Petascale High Order Dynamic Rupture Earthquake Simulations on Heterogeneous Supercomputers · SC 2014 |
High-performance computing › scientific computing systems
earthquake simulation |
0.2 | 1 | 2014 | Petascale High Order Dynamic Rupture Earthquake Simulations on Heterogeneous Supercomputers · SC 2014 |
Storage systems › file systems › distributed file system
parallel file system |
0.2 | 1 | 2014 | A User-Friendly Approach for Tuning Parallel File Operations · SC 2014 |
High-performance computing
parallel i/o |
0.2 | 1 | 2014 | A User-Friendly Approach for Tuning Parallel File Operations · SC 2014 |
Storage systems › i/o optimization
parallel i/o optimization |
0.2 | 1 | 2014 | A User-Friendly Approach for Tuning Parallel File Operations · SC 2014 |
High-performance computing
performance optimization at scale |
0.2 | 1 | 2014 | Petascale High Order Dynamic Rupture Earthquake Simulations on Heterogeneous Supercomputers · SC 2014 |
High-performance computing
scientific computing systems |
0.2 | 1 | 2014 | Petascale High Order Dynamic Rupture Earthquake Simulations on Heterogeneous Supercomputers · SC 2014 |
Distributed systems
resource monitoring |
0.2 | 1 | 2013 | Enabling comprehensive data-driven system management for large computational facilities · SC 2013 |
Performance modeling and evaluation
workload characterization |
0.2 | 1 | 2013 | Enabling comprehensive data-driven system management for large computational facilities · SC 2013 |
Parallel and multicore computing › parallel programming models › message passing
MPI runtime |
0.1 | 1 | 2012 | Design of a scalable InfiniBand topology service to enable network-topology-aware placement of processes · SC 2012 |
Distributed systems
topology discovery |
0.1 | 1 | 2012 | Design of a scalable InfiniBand topology service to enable network-topology-aware placement of processes · SC 2012 |
Visualization and visual analytics
flow visualization |
0.1 | 1 | 2007 | Virtual Rheoscopic Fluids for Flow Visualization · IEEE Trans. Vis. Comput. Graph. 2007 |
Rendering
volume rendering |
0.1 | 1 | 2007 | Virtual Rheoscopic Fluids for Flow Visualization · IEEE Trans. Vis. Comput. Graph. 2007 |
Storage systems › file systems › distributed file system › parallel file system
lustre file system |
0.1 | 1 | 2014 | A User-Friendly Approach for Tuning Parallel File Operations · SC 2014 |
Hardware accelerators and domain-specific architectures › many-core accelerator
many-core coprocessor |
0.1 | 1 | 2014 | Petascale High Order Dynamic Rupture Earthquake Simulations on Heterogeneous Supercomputers · SC 2014 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.0 | 1 | 2013 | Enabling comprehensive data-driven system management for large computational facilities · SC 2013 |
Distributed systems › communication optimization
topology-aware communication |
0.0 | 1 | 2012 | Design of a scalable InfiniBand topology service to enable network-topology-aware placement of processes · SC 2012 |
Methods — techniques the papers use, named apart from their topics
performance modeling · 0.4unstructured mesh · 0.2discontinuous galerkin method · 0.2auto-tuning library · 0.2predictive analytics · 0.2log integration · 0.2neighbor-joining algorithm · 0.1anisotropic reflectance model · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Enabling Dependability-Driven Resource Use and Message Log-Analysis for Cluster System DiagnosisabstractRecent work have used both failure logs and resource use data separately (and together) to detect system failure-inducing errors and to diagnose system failures. System failure occurs as a result of error propagation and the (unsuccessful) execution of error recovery mechanisms. Knowledge of error propagation patterns and unsuccessful error recovery is important for more accurate and detailed failure diagnosis, and knowledge of recovery protocols deployment is important for improving system reliability. This paper presents the CORRMEXT framework which carries failure diagnosis another significant step forward by analyzing and reporting error propagation patterns and degrees of success and failure of error recovery protocols. CORRMEXT uses both error messages and resource use data in its analyses. Application of CORRMEXT to data from the Ranger supercomputer have produced new insights. CORRMEXT has: (i) identified correlations between resource use counters that capture recovery attempts after an error, (ii) identified correlations between error events to capture error propagation patterns within the system, (iii) identified error propagation and recovery paths during system execution to explain system behaviour, (iv) showed that the earliest times of change in system behaviour can only be identified by analyzing both the correlated resource use counters and correlated errors. CORRMEXT will be installed on the HPC clusters at the Texas Advanced Computing Center in Autumn 2017. Edward Chuah, Arshad Jhumka, Samantha Alt, Theodoros Damoulas, Nentawe Gurumdimma, Marie-Christine Sawley, William L. Barth, Tommy Minyard, James C. Browne |
HiPC | 7 |
| 2017 | On Orienting Edges of Unstructured Two- and Three-Dimensional MeshesabstractFinite element codes typically use data structures that represent unstructured meshes as collections of cells, faces, and edges, each of which require associated coordinate systems. One then needs to store how the coordinate system of each edge relates to that of neighboring cells. However, we can simplify data structures and algorithms if we can a priori orient coordinate systems in such a way that the coordinate systems on the edges follow uniquely from those on the cells by rule. Such rules require that every unstructured mesh allow the assignment of directions to edges that satisfy the convention in adjacent cells. We show that the convention chosen for unstructured quadrilateral meshes in the deal.II library always allows to orient meshes. It can therefore be used to make codes simpler, faster, and less bug prone. We present an algorithm that orients meshes in O ( N ) operations. We then show that consistent orientations are not always possible for 3D hexahedral meshes. Thus, cells generally need to store the direction of adjacent edges, but our approach also allows the characterization of cases where this is not necessary. The 3D extension of our algorithm either orients edges consistently, or aborts, both within O ( N ) steps. Rainer Agelek, Wolfgang Bangerth, William L. Barth |
ACM Trans. Math. Softw. | 4 |
| 2016 | Using Message Logs and Resource Use Data for Cluster Failure DiagnosisabstractFailure diagnosis for large compute clusters using only message logs is known to be incomplete. Recent availability of resource use data provides another potentially useful source of data for failure detection and diagnosis. Early work combining message logs and resource use data for failure diagnosis has shown promising results. This paper describes the CRUMEL framework which implements a new approach to combining rationalized message logs and resource use data for failure diagnosis. CRUMEL identifies patterns of errors and resource use and correlates these patterns by time with system failures. Application of CRUMEL to data from the Ranger supercomputer has yielded improved diagnoses over previous research. CRUMEL has: (i) showed that more events correlated with system failures can only be identified by applying different correlation algorithms, (ii) confirmed six groups of errors, (iii) identified Lustre I/O resource use counters which are correlated with occurrence of Lustre faults which are potential flags for online detection of failures, (iv) matched the dates of correlated error events and correlated resource use with the dates of compute node hang-ups and (v) identified two more error groups associated with compute node hang-ups. The pre-processed data will be put on the public domain in September, 2016. Edward Chuah, Arshad Jhumka, James C. Browne, Nentawe Gurumdimma, Sai Narasimhamurthy, William L. Barth |
HiPC | 6 |
| 2014 | Petascale High Order Dynamic Rupture Earthquake Simulations on Heterogeneous SupercomputersabstractWe present an end-to-end optimization of the innovative Arbitrary high-order DERivative Discontinuous Galerkin (ADER-DG) software SeisSol targeting Intel® Xeon Phi coprocessor platforms, achieving unprecedented earthquake model complexity through coupled simulation of full frictional sliding and seismic wave propagation. SeisSol exploits unstructured meshes to flexibly adapt for complicated geometries in realistic geological models. Seismic wave propagation is solved simultaneously with earthquake faulting in a multiphysical manner leading to a heterogeneous solver structure. Our architecture aware optimizations deliver up to 50% of peak performance, and introduce an efficient compute-communication overlapping scheme shadowing the multiphysics computations. SeisSol delivers near-optimal weak scaling, reaching 8.6 DP-PFLOPS on 8,192 nodes of the Tianhe-2 supercomputer. Our performance model projects reaching 18 -- 20 DP-PFLOPS on the full Tianhe-2 machine. Of special relevance to modern civil engineering needs, our pioneering simulation of the 1992 Landers earthquake shows highly detailed rupture evolution and ground motion at frequencies up to 10 Hz. Alexander Heinecke, Alexander Breuer, Sebastian Rettenberger, Michael Bader, Alice-Agnes Gabriel, Christian Pelties, Arndt Bode, William L. Barth, Xiangke Liao, Karthikeyan Vaidyanathan, Mikhail Smelyanskiy, Pradeep Dubey |
SC | 8 |
| 2014 | A User-Friendly Approach for Tuning Parallel File OperationsabstractThe Lustre file system provides high aggregated I/O bandwidth and is in widespread use throughout the HPC community. Here we report on work (1) developing a model for understanding collective parallel MPI write operations on Lustre, and (2) producing a library that optimizes parallel write performance in a user-friendly way. We note that a system's default stripe count is rarely a good choice for parallel I/O, and that performance depends on a delicate balance between the number of stripes and the actual (not requested) number of collective writers. Unfortunate combinations of these parameters may degrade performance considerably. For the programmer, however, it's all about the stripe count: an informed choice of this single parameter allows MPI to assign writers in a way that achieves near-optimal performance. We offer recommendations for those who wish to tune performance manually and describe the easy-to-use T3PIO library that manages the tuning automatically. Robert T. McLay, Doug James, Si Liu 0008, John Cazes, William L. Barth |
SC | 5 |
| 2014 | Comprehensive, open-source resource usage measurement and analysis for HPC systemsabstractSUMMARY The important role high‐performance computing (HPC) resources play in science and engineering research, coupled with its high cost (capital, power and manpower), short life and oversubscription, requires us to optimize its usage – an outcome that is only possible if adequate analytical data are collected and used to drive systems management at different granularities – job, application, user and system. This paper presents a method for comprehensive job, application and system‐level resource use measurement, and analysis and its implementation. The steps in the method are system‐wide collection of comprehensive resource use and performance statistics at the job and node levels in a uniform format across all resources, mapping and storage of the resultant job‐wise data to a relational database, which enables further implementation and transformation of the data to the formats required by specific statistical and analytical algorithms. Analyses can be carried out at different levels of granularity: job, user, application or system‐wide. Measurements are based on a new lightweight job‐centric measurement tool ‘TACC_Stats’, which gathers a comprehensive set of resource use metrics on all compute nodes and data logged by the system scheduler. The data mapping and analysis tools are an extension of the XDMoD project. The method is illustrated with analyses of resource use for the Texas Advanced Computing Center's Lonestar4, Ranger and Stampede supercomputers and the HPC cluster at the Center for Computational Research. The illustrations are focused on resource use at the system, job and application levels and reveal many interesting insights into system usage patterns and also anomalous behavior due to failure/misuse. The method can be applied to any system that runs the TACC_Stats measurement tool and a tool to extract job execution environment data from the system scheduler. Copyright © 2014 John Wiley & Sons, Ltd. James C. Browne, Robert L. DeLeon, Abani K. Patra, William L. Barth, John L. Hammond, Matthew D. Jones, Thomas R. Furlani, Barry I. Schneider, Steven M. Gallo, Amin Ghadersohi, Ryan J. Gentner, Jeffrey T. Palmer, Nikolay Simakov, Martins Innus, Andrew E. Bruno, Joseph P. White, Cynthia D. Cornelius, Thomas Yearke, Kyle Marcus, Gregor von Laszewski, Fugang Wang |
Concurr. Comput. Pract. Exp. | 4 |
| 2013 | Design of network topology aware scheduling services for large InfiniBand clustersabstractThe goal of any scheduler is to satisfy user's demands for computation and achieve a good performance in overall system utilization by efficiently assigning jobs to resources. However, the current state-of-the-art scheduling techniques do not intelligently balance node allocation based on the total bandwidth available between switches - that leads to over subscription. Additionally, poor placement of processes can lead to network congestion and poor performance. In this paper, we explore the design of a network-topology-aware plugin for the SLURM job scheduler for modern InfiniBand-based clusters. We present designs to enhance the performance of applications with varying communication characteristics. Through our techniques, we are able to considerably reduce the amount of network contention observed during the Alltoall / FFT operations. The results of our experimental evaluation indicate that our proposed technique is able to deliver up to a 9% improvement in the communication time of P3DFFT at 512 processes. We also see that our techniques are able to increase the performance of microbenchmarks that rely on point-to-point operations up to 40% for all message sizes. Our techniques were also able to improve the throughput of a 512-core cluster by up to 8%. Hari Subramoni, Devendar Bureddy, Krishna Chaitanya Kandalla, Karl W. Schulz, William L. Barth, Jonathan L. Perkins, Mark Daniel Arnold, Dhabaleswar K. Panda 0001 |
CLUSTER | 5 |
| 2013 | Enabling comprehensive data-driven system management for large computational facilitiesabstractThis paper presents a tool chain, based on the open source tool TACC_Stats, for systematic and comprehensive job level resource use measurement for large cluster computers, and its incorporation into XDMoD, a reporting and analytics framework for resource management that targets meeting the information needs of users, application developers, systems administrators, systems management and funding managers. Accounting, scheduler and event logs are integrated with system performance data from TACC_Stats. TACC_Stats periodically records resource use including many hardware counters for each job running on each node. Furthermore, system level metrics are obtained through aggregation of the node (job) level data. Analysis of this data generates many types of standard and custom reports and even a limited predictive capability that has not previously been available for open-source, Linux-based software systems. This paper presents case studies of information that can be applied for effective resource management. We believe this system to be the first fully comprehensive system for supporting the information needs of all stakeholders in open-source software based HPC systems. James C. Browne, Robert L. DeLeon, Charng-Da Lu, Matthew D. Jones, Steven M. Gallo, Amin Ghadersohi, Abani K. Patra, William L. Barth, John L. Hammond, Thomas R. Furlani, Robert T. McLay |
SC | 8 |
| 2013 | Linking Resource Usage Anomalies with System Failures from Cluster Log DataabstractBursts of abnormally high use of resources are thought to be an indirect cause of failures in large cluster systems, but little work has systematically investigated the role of high resource usage on system failures, largely due to the lack of a comprehensive resource monitoring tool which resolves resource use by job and node. The recently developed TACC_Stats resource use monitor provides the required resource use data. This paper presents the ANCOR diagnostics system that applies TACC_Stats data to identify resource use anomalies and applies log analysis to link resource use anomalies with system failures. Application of ANCOR to first identify multiple sources of resource anomalies on the Ranger supercomputer, then correlate them with failures recorded in the message logs and diagnosing the cause of the failures, has identified four new causes of compute node soft lockups. ANCOR can be adapted to any system that uses a resource use monitor which resolves resource use by job. Edward Chuah, Arshad Jhumka, Sai Narasimhamurthy, John L. Hammond, James C. Browne, William L. Barth |
SRDS | 6 |
| 2012 | Design of a scalable InfiniBand topology service to enable network-topology-aware placement of processesabstractOver the last decade, InfiniBand has become an increasingly popular interconnect for deploying modern supercomputing systems. However, there exists no detection service that can discover the underlying network topology in a scalable manner and expose this information to runtime libraries and users of the high performance computing systems in a convenient way. In this paper, we design a novel and scalable method to detect the InfiniBand network topology by using Neighbor-Joining techniques (NJ). To the best of our knowledge, this is the first instance where the neighbor joining algorithm has been applied to solve the problem of detecting InfiniBand network topology. We also design a network-topology-aware MPI library that takes advantage of the network topology service. The library places processes taking part in the MPI job in a network-topology-aware manner with the dual aim of increasing intra-node communication and reducing the long distance inter-node communication across the InfiniBand fabric. Hari Subramoni, Sreeram Potluri, Krishna Chaitanya Kandalla, William L. Barth, Jérôme Vienne, Jeff Keasler, Karen A. Tomko, Karl W. Schulz, Adam Moody, Dhabaleswar K. Panda 0001 |
SC | 4 |
| 2011 | Design and Evaluation of Network Topology-/Speed- Aware Broadcast Algorithms for InfiniBand ClustersabstractIt is an established fact that the network topology can have an impact on the performance of scientific parallel applications. However, little work has been done to design an easy to use solution inside a communication library supporting a parallel programming model where the complexities of making the application performance network topology agnostic is hidden from the end user. Similarly, the rapid improvements in networking technology and speed are resulting in many commodity clusters becoming heterogeneous, with respect to networking speed. For example, switches and adapters belonging to different generations (SDR - 8 Gbps, DDR - 16 Gbps and QDR - 36 Gbps speeds in InfiniBand) are integrated into a single system. This leads to an additional challenge to make the communication library aware of the performance implications of heterogeneous link speeds. Accordingly, the communication library can perform optimizations taking link speed into account. In this paper, we propose a framework to automatically detect the topology and speed of an InfiniBand network and make it available to users through an easy to use interface. We also make design changes inside the MPI library to dynamically query this topology detection service and to form a topology model of the underlying network. We have redesigned the broadcast algorithm to take into account this network topology information and dynamically adapt the communication pattern to best fit the characteristics of the underlying network. To the best of our knowledge, this is the first such work for InfiniBand clusters. Our experimental results show that, for large homogeneous systems and large message sizes, we get up to 14% improvement in the latency of the broadcast operation using our proposed network topology-aware scheme over the default scheme at the micro-benchmark level. At the application level, the proposed framework delivers up to 8% improvement in total application run-time especially as job size scales up. The proposed network speed-aware algorithms are able to attain micro-benchmark performance on the heterogeneous SDR-DDR InfiniBand cluster to perform on par with runs on the DDR only portion of the cluster for small to medium sized messages. We also demonstrate that the network speed aware algorithms perform 70% to 100% better than the naive algorithms when both are run on the heterogeneous SDR-DDR InfiniBand cluster. Hari Subramoni, Krishna Chaitanya Kandalla, Jérôme Vienne, Sayantan Sur, William L. Barth, Karen A. Tomko, Robert T. McLay, Karl W. Schulz, Dhabaleswar K. Panda 0001 |
CLUSTER | 5 |
| 2010 | Quantifying performance benefits of overlap using MPI-2 in a seismic modeling applicationabstractAWM-Olsen is a widely used ground motion simulation code based on a parallel finite difference solution of the 3-D velocity-stress wave equation. This application runs on tens of thousands of cores and consumes several million CPU hours on the TeraGrid Clusters every year. A significant portion of its run-time (37% in a 4,096 process run), is spent in MPI communication routines. Hence, it demands an optimized communication design coupled with a low-latency, high-bandwidth network and an efficient communication subsystem for good performance. In this paper, we analyze the performance bottlenecks of the application with regard to the time spent in MPI communication calls. We find that much of this time can be overlapped with computation using MPI non-blocking calls. We use both two-sided and MPI-2 one-sided communication semantics to re-design the communication in AWM-Olsen. We find that with our new design, using MPI-2 one-sided communication semantics, the entire application can be sped up by 12% at 4K processes and by 10% at 8K processes on a state-of-the-art InfiniBand cluster, Ranger at the Texas Advanced Computing Center (TACC). Sreeram Potluri, Ping Lai, Karen A. Tomko, Sayantan Sur, Yifeng Cui, Mahidhar Tatineni, Karl W. Schulz, William L. Barth, Amitava Majumdar 0001, Dhabaleswar K. Panda 0001 |
ICS | 8 |
| 2007 | Virtual Rheoscopic Fluids for Flow VisualizationabstractPhysics-based flow visualization techniques seek to mimic laboratory flow visualization methods with virtual analogues. In this work we describe the rendering of a virtual rheoscopic fluid to produce images with results strikingly similar to laboratory experiments with real-world rheoscopic fluids using products such as Kalliroscope. These fluid additives consist of microscopic, anisotropic particles which, when suspended in the flow, align with both the flow velocity and the local shear to produce high-quality depictions of complex flow structures. Our virtual rheoscopic fluid is produced by defining a closed-form formula for the orientation of shear layers in the flow and using this orientation to volume render the flow as a material with anisotropic reflectance and transparency. Examples are presented for natural convection, thermocapillary convection, and Taylor-Couette flow simulations. The latter agree well with photographs of experimental results of Taylor-Couette flows from the literature. William L. Barth, Christopher Burns 0004 |
IEEE Trans. Vis. Comput. Graph. | 1 |