VLDB 2026 Research / reviewers in the wild / expert
Al Geist
dblp:g/AlGeist · also George Al Geist II
· DBLP profile ↗
27ranked-venue papers
4as first author
0since 2021 · last 2020
0000-0001-9350-1688ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 4 first-authorDatabases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Security and privacy · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
High-performance computing · 73% Performance modeling and evaluation · 14% Distributed systems · 9% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 12 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › supercomputing
supercomputer deployment |
0.3 | 1 | 2018 | The design, deployment, and evaluation of the CORAL pre-exascale systems · SC 2018 |
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2018 | The design, deployment, and evaluation of the CORAL pre-exascale systems · SC 2018 |
High-performance computing › scientific computing systems
molecular dynamics simulation |
0.1 | 1 | 2006 | Performance characterization of molecular dynamics techniques for biomolecular simulations · PPoPP 2006 |
High-performance computing
scientific computing |
0.1 | 1 | 2006 | Performance characterization of molecular dynamics techniques for biomolecular simulations · PPoPP 2006 |
Bioinformatics and computational biology
protein-protein interaction prediction |
0.0 | 1 | 2003 | Inference of Protein-Protein Interactions by Unlikely Profile Pair · ICDM 2003 |
High-performance computing
scientific computing systems |
0.0 | 1 | 1999 | Toward a Common Component Architecture for High-Performance Scientific Computing · HPDC 1999 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 1998 | HARNESS: Heterogeneous Adaptable Reconfigurable NEtworked SystemS · HPDC 1998 |
Bioinformatics and computational biology › molecular informatics › molecular modeling
biomolecular simulation |
0.0 | 1 | 2006 | Performance characterization of molecular dynamics techniques for biomolecular simulations · PPoPP 2006 |
Distributed systems › replication
data replication |
0.0 | 1 | 1997 | Scalable Networked Information Processing Environment (SNIPE) · SC 1997 |
Distributed systems › fault tolerance
fault-tolerant distributed systems |
0.0 | 1 | 1997 | Scalable Networked Information Processing Environment (SNIPE) · SC 1997 |
High-performance computing › distributed computing infrastructure
metacomputing |
0.0 | 1 | 1997 | Scalable Networked Information Processing Environment (SNIPE) · SC 1997 |
Distributed systems
grid computing |
0.0 | 1 | 1998 | HARNESS: Heterogeneous Adaptable Reconfigurable NEtworked SystemS · HPDC 1998 |
Methods — techniques the papers use, named apart from their topics
statistical simulation · 0.0bootstrapping · 0.0port connection model · 0.0interface definition language · 0.0pluggable architecture · 0.0dynamic reconfiguration · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Application health monitoring for extreme-scale resiliency using cooperative fault managementabstractsupercomputers, and beyond. Applications oblivious to and incapable of handling transient soft and hard errors could waste supercomputing resources or, worse, yield misleading scientific insights. We introduce a novel application-driven silent error detection and recovery strategy based on application health monitoring. Our methodology uses application output that follows known patterns as indicators of an application's health, and knowledge that violation of these patterns could be indication of faults. Information from system monitors that report hardware and software health status is used to corroborate faults. Collectively, this information is used by a fault coordinator agent to take preventive and corrective measures by applying computational steering to an application between checkpoints. This cooperative fault management system uses the Fault Tolerance Backplane as a communication channel. The benefits of this framework are demonstrated with two real application case studies, molecular dynamics and quantum chemistry simulations, on scalable clusters with simulated memory and I/O corruptions. The developed approach is general and can be easily applied to other applications. Pratul K. Agarwal, Thomas J. Naughton, Byung H. Park, David E. Bernholdt, Joshua Hursey, Al Geist |
Concurr. Comput. Pract. Exp. | 6 |
| 2018 | The design, deployment, and evaluation of the CORAL pre-exascale systems
Sudharshan S. Vazhkudai, Bronis R. de Supinski, Arthur S. Bland, Al Geist, James C. Sexton, James A. Kahle, Christopher Zimmer 0001, Scott Atchley, Sarp Oral, Don E. Maxwell, Verónica G. Vergara Larrea, Adam Bertsch, Robin Goldstone, Wayne Joubert, Christopher M. Chambreau, David Appelhans, Robert Blackmore, Ben Casses, George Chochia, Gene Davison, Matthew Ezell, Thomas Gooding, Elsa Gonsiorowski, Leopold Grinberg, Bill Hanson, Bill Hartner, Ian Karlin, Matthew L. Leininger, Dustin Leverman, Chris Marroquin, Adam Moody, Martin Ohmacht, Ramesh Pankajakshan, Fernando Pizzano, James H. Rogers, Bryan S. Rosenburg, Drew Schmidt, Mallikarjun Shankar, Feiyi Wang, Py Watson, Bob Walkup, Lance D. Weems, Junqi Yin |
SC | 4 |
| 2009 | System log pre-processing to improve failure predictionabstractLog preprocessing, a process applied on the raw log before applying a predictive method, is of paramount importance to failure prediction and diagnosis. While existing filtering methods have demonstrated good compression rate, they fail to preserve important failure patterns that are crucial for failure analysis. To address the problem, in this paper we present a log preprocessing method. It consists of three integrated steps: (1) event categorization to uniformly classify system events and identify fatal events; (2) event filtering to remove temporal and spatial redundant records, while also preserving necessary failure patterns for failure analysis; (3) causality-related filtering to combine correlated events for filtering through apriori association rule mining. We demonstrate the effectiveness of our preprocessing method by using real failure logs collected from the Cray XT4 at ORNL and the Blue Gene/L system at SDSC. Experiments show that our method can preserve more failure patterns for failure analysis, thereby improving failure prediction by up to 174%. Ziming Zheng, Zhiling Lan, Al Geist |
DSN | 4 |
| 2009 | CIFTS: A Coordinated Infrastructure for Fault-Tolerant SystemsabstractConsiderable work has been done on providing fault tolerance capabilities for different software components on large-scale high-end computing systems. Thus far, however, these fault-tolerant components have worked insularly and independently and information about faults is rarely shared. Such lack of system-wide fault tolerance is emerging as one of the biggest problems on leadership-class systems. In this paper, we propose a coordinated infrastructure, named CIFTS, that enables system software components to share fault information with each other and adapt to faults in a holistic manner. Central to the CIFTS infrastructure is a Fault Tolerance Backplane (FTB) that enables fault notification and awareness throughout the software stack, including fault-aware libraries, middleware, and applications. We present details of the CIFTS infrastructure and the interface specification that has allowed various software programs, including MPICH2, MVAPICH, Open MPI, and PVFS, to plug into the CIFTS infrastructure. Further, through a detailed evaluation we demonstrate the nonintrusive low-overhead capability of CIFTS that lets applications run with minimal performance degradation. Rinku Gupta, Pete Beckman, Ewing L. Lusk, Paul Hargrove, Al Geist, Dhabaleswar K. Panda 0001, Andrew Lumsdaine, Jack J. Dongarra |
ICPP | 6 |
| 2008 | Rapid and robust ranking of text documents in a dynamically changing corpusabstractRanking documents in a selected corpus plays an important role in information retrieval systems. Despite notable advances in this direction, with continuously accumulating text documents, maintaining up-to-date ordering among documents in the domains of interest is a challenging task. Conventional approaches can produce an ordering that is only valid within a given corpus. Thus, with such approaches, ordering should be completely redone as documents are added to or deleted from the corpus. In this paper, we introduce a corpus- independent framework for rapid ordering of documents in a dynamically changing corpus. Like in many practical approaches, our framework suggests utilizing a similarity measure in some metric space indicating the degree of relevance of a document to the domain of interest. However, unlike in corpus- dependent approaches, the relevance score of a document remains valid with changes being introduced into the corpus (insertion of new documents, for example), thus allowing a rapid ordering within the corpus. This paper particularly details a statistical approach to compute such relevance scores. Nagiza F. Samatova, Rajesh Munavalli, Ramya Krishnamurthy, Houssain Kettani, Al Geist |
AICCSA | 6 |
| 2008 | Virtualized Environments for the Harness High Performance Computing WorkbenchabstractThis paper describes recent accomplishments in providing a virtualized environment concept and prototype for scientific application development and deployment as part of the Harness High Performance Computing (HPC) Workbench research effort. The presented work focuses on tools and mechanisms that simplify scientific application development and deployment tasks, such that only minimal adaptation is needed when moving from one HPC system to another or after HPC system upgrades. The overall technical approach focuses on the concept of adapting the HPC system environment to the actual needs of individual scientific applications instead of the traditional scheme of adapting scientific applications to individual HPC system environment properties. The presented prototype implementation is based on the mature and lightweight chroot virtualization approach for Unix-type systems with a focus on virtualized file system structure and virtualized shell environment variables utilizing virtualized environment configuration descriptions in Extensible Markup Language (XML) format. The presented work can be easily extended to other virtualization technologies, such as system-level virtualization solutions using hypervisors. Björn Könning, Christian Engelmann, Stephen L. Scott, Al Geist |
PDP | 4 |
| 2007 | The Neutron Science TeraGrid Gateway: a TeraGrid science gateway to support the Spallation Neutron SourceabstractAbstract The National Science Foundation's Extensible Terascale Facility (ETF), or TeraGrid ( http://www.teragrid.org/ ), is entering its operational phase. An example of an ETF science gateway effort is the Neutron Science TeraGrid Gateway (NSTG). The Oak Ridge National Laboratory (ORNL) resource provider effort (ORNL‐RP) now in operation is bridging the gap between a large‐scale experimental community and the TeraGrid as a large‐scale national cyberinfrastructure. Of particular importance here is the collaboration with the Spallation Neutron Source (SNS) at ORNL. The U.S. Department of Energy's SNS ( http://www.sns.gov/ ) at ORNL will be commissioned in the spring of 2006 as the world's brightest source of neutrons. Neutron science users can run experiments, generate datasets, perform data reduction, analysis, visualize results, collaborate with remotes users, and archive long‐term data in repositories with curation services. The ORNL‐RP and the SNS data analysis group have spent 18 months developing and exploring user requirements, including the creation of prototypical services such as a facility portal, data, and application execution services. We describe results from these efforts and discuss implications for science gateway creation. Finally, we show incorporation into implementation planning for the NSTG and SNS architectures. The plan is for a primarily portal‐based user interaction supported by a service‐oriented architecture for functional implementation. Published in 2006 by John Wiley & Sons, Ltd. John Cobb, Al Geist, James Arthur Kohl, Stephen D. Miller, Peter F. Peterson, Gregory G. Pike, Michael A. Reuter, Tom Swain, Sudharshan S. Vazhkudai, Nithya N. Vijayakumar |
Concurr. Comput. Pract. Exp. | 2 |
| 2006 | A Parallel Plug-In Programming Paradigm
Ronald Baumann, Christian Engelmann, Al Geist |
HPCC | 3 |
| 2006 | Performance characterization of molecular dynamics techniques for biomolecular simulationsabstractLarge-scale simulations and computational modeling using molecular dynamics (MD) continues to make significant impacts in the field of biology. It is well known that simulations of biological events at native time and length scales requires computing power several orders of magnitude beyond today's commonly available systems. Supercomputers, such as IBM Blue Gene/L and Cray XT3, will soon make tens to hundreds of teraFLOP/s of computing power available by utilizing thousands of processors. The popular algorithms and MD applications, however, were not initially designed to run on thousands of processors. In this paper, we present detailed investigations of the performance issues, which are crucial for improving the scalability of the MD-related algorithms and applications on massively parallel processing (MPP) architectures. Due to the varying characteristics of biological input problems, we study two prototypical biological complexes that use the MD algorithm: an explicit solvent and an implicit solvent. In particular, we study the AMBER application, which supports a variety of these types of input problems. For the explicit solvent problem, we focused on the particle mesh Ewald (PME) method for calculating the electrostatic energy, and for the implicit solvent model, we targeted the Generalized Born (GB) calculation. We uncovered and subsequently modified a limitation in AMBER that restricted the scaling beyond 128 processors. We collected performance data for experiments on up to 2048 Blue Gene/L and XT3 processors and subsequently identified that the scaling is largely limited by the underlying algorithmic characteristics and also by the implementation of the algorithms. Furthermore, we found that the input problem size of biological system is constrained by memory available per node. In conclusion, our results indicate that MD codes can significantly benefit from the current generation architectures with relatively modest optimization efforts. Nevertheless, the key for enabling scientific breakthroughs lies in exploiting the full potential of these new architectures. Sadaf R. Alam, Jeffrey S. Vetter, Pratul K. Agarwal, Al Geist |
PPoPP | 4 |
| 2005 | The Design and Prototype of RUDA, a Distributed Grid Accounting System
Meili Chen, Al Geist, David E. Bernholdt, Kasidit Chanchio, Daniel L. Million |
ICCSA (3) | 2 |
| 2004 | Reservoir-Based Random Sampling with Replacement from Data StreamabstractRandom sampling is a widely accepted basis for estimation from large data sets that outstrip available computer memory. When the data comes as a stream, its total size is potentially infinite and usually only one pass through the data is possible. Reservoir sampling is a method of maintaining a fixed size random sample from streaming data. All reservoir schemes that have been introduced in the past are random sampling without replacement; no duplicates are allowed in a sample. This paper introduces a new method for reservoir sampling with replacement. We first prove that the proposed method indeed maintains a random sample with replacement at any given time. Then we introduce a refined version that significantly speeds up the overall sampling procedure. George Ostrouchov, Nagiza F. Samatova, Al Geist |
SDM | 4 |
| 2003 | Inference of Protein-Protein Interactions by Unlikely Profile PairabstractWe note that a set of statistically "unusual" protein-profile pairs in experimentally determined database of protein-protein interactions can typify protein-protein interactions, and propose a novel method called PICUPP that sifts such protein-profile pairs using a statistical simulation. It is demonstrated that unusual Pfam and InterPro profile pairs can be extracted from the DIP database using a bootstrapping approach. We particularly illustrate that such protein-profile pairs can be used for predicting putative pairs of interacting proteins. Their prediction accuracies are around 86% and 90% when InterPro and Pfam profiles are used, respectively at 75% confidence level. George Ostrouchov, Gong-Xin Yu, Al Geist, Andrey Gorin, Nagiza F. Samatova |
ICDM | 4 |
| 2002 | RACHET: An Efficient Cover-Based Merging of Clustering Hierarchies from Distributed Datasets
Nagiza F. Samatova, George Ostrouchov, Al Geist, Anatoli V. Melechko |
Distributed Parallel Databases | 3 |
| 2001 | M3C: Managing and Monitoring Multiple ClustersabstractPC clusters running Linux, provide the computational power of supercomputers of just a few years ago at a fraction of the purchase price. The lack of good administration and user level application management tools makes the operation of clusters more difficult than it need be and results in an increased operation cost. This paper describes an ongoing effort at Oak Ridge National Laboratory to develop tools to simplify the administration and use of computation clusters. The two tools integrated in this work include the M3C (Managing and Monitoring Multiple Clusters) and C3 (Cluster Command and Control) tool suite. M3C provides a Web-based graphical user interface for cluster administration. It is designed as an extendable framework that can work with different underlying back-end tools. C3 is an independent project that provides the underlying commands to effect an operation on a cluster. In this paper it serves as an example how a back-end tool can be integrated into the M3C framework. Michael J. Brim, Al Geist, Brian Luethke, Jens Schwidder, Stephen L. Scott |
CCGRID | 2 |
| 2000 | ORNL M3C tool
Al Geist, Brian Luethke, Jens Schwidder, Stephen L. Scott |
CLUSTER | 1 |
| 1999 | Toward a Common Component Architecture for High-Performance Scientific ComputingabstractDescribes work in progress to develop a standard for interoperability among high-performance scientific components. This research stems from the growing recognition that the scientific community needs to better manage the complexity of multidisciplinary simulations and better address scalable performance issues on parallel and distributed architectures. The driving force for this is the need for fast connections among components that perform numerically intensive work and for parallel collective interactions among components that use multiple processes or threads. This paper focuses on the areas we believe are most crucial in this context, namely an interface definition language that supports scientific abstractions for specifying component interfaces and a port connection model for specifying component interactions. Robert C. Armstrong, Dennis Gannon, Al Geist, Kate Keahey, Scott R. Kohn, Lois C. McInnes, Steven G. Parker, Brent A. Smolinski |
HPDC | 3 |
| 1999 | HARNESS: a next generation distributed virtual machine
Micah D. Beck, Jack J. Dongarra, Graham E. Fagg, Al Geist, Paul Gray, James Arthur Kohl, Mauro Migliardi, Keith Moore, Terry Moore, Philip Papadopoulous |
Future Gener. Comput. Syst. | 4 |
| 1999 | Heterogeneous parallel and distributed computing
Vaidy S. Sunderam, Al Geist |
Parallel Comput. | 2 |
| 1998 | HARNESS: Heterogeneous Adaptable Reconfigurable NEtworked SystemSabstractWe describe our vision, goals and plans for HARNESS, a distributed, reconfigurable and heterogeneous computing environment that supports dynamically adaptable parallel applications. HARNESS builds on the core concept of the personal virtual machine as an abstraction for distributed parallel programming, but fundamentally extends this idea, greatly enhancing dynamic capabilities. HARNESS is being designed to embrace dynamics at every level through a pluggable model that allows multiple distributed virtual machines (DVMs) to merge, split and interact with each other. It provides mechanisms for new and legacy applications to collaborate with each other using the HARNESS infrastructure, and defines and implements new plug-in interfaces and modules so that applications can dynamically customize their virtual environment. HARNESS fits well within the larger picture of computational grids as a dynamic mechanism to hide the heterogeneity and complexity of the nationally distributed infrastructure. HARNESS DVMs allow programmers and users to construct personal subsets of an existing computational grid and treat them as unified network computers, providing a familiar and comfortable environment that provides easy-to-understand scoping. Jack J. Dongarra, Graham E. Fagg, Al Geist, James Arthur Kohl, Philip M. Papadopoulos, Stephen L. Scott, Vaidy S. Sunderam, M. Magliardi |
HPDC | 3 |
| 1998 | LPVM: a step towards multithread PVMabstractLPVM (lightweight-process PVM) is an experimental PVM version that supports the use of lightweight processes or threads as the basic unit of parallelism. It was designed to study potential performance improvements and implementation issues required by multithread message-passing systems. The current version of LPVM was implemented on SMPs (shared memory processors) using POSIX threads and is designed to be thread-safe. Initial test results on a SUN SMP are compared with the performance of standard PVM. Task spawning is an order of magnitude faster in LPVM, but message-passing between threads using shared memory was found to be slower than standard PVM due to overheads in getting and releasing shared memory locks. © 1998 John Wiley & Sons, Ltd. Honbo Zhou, Al Geist |
Concurr. Pract. Exp. | 2 |
| 1997 | Scalable Networked Information Processing Environment (SNIPE)abstractSNIPE is a metacomputing system that aims to provide a reliable, secure, fault-tolerant environment for long-term distributed computing applications and data stores across the global InterNet. This system combines global naming and replication of both processing and data to support large scale information processing applications leading to better availablity and reliability than currently available with typical cluster computing and/or distributed computer environments. Graham E. Fagg, Keith Moore, Jack J. Dongarra, Al Geist |
SC | 4 |
| 1994 | The PVM Concurrent Computing System: Evolution, Experiences, and Trends
Vaidy S. Sunderam, Al Geist, Jack J. Dongarra, Robert Manchek |
Parallel Comput. | 2 |
| 1993 | Panel - Software Tools for High-Performance Distributed Computing
Vaidy S. Sunderam, Geoffrey C. Fox, Al Geist, William Gropp, Bob Harrison, Adam Kolawa, Michael J. Quinn, Anthony Skjellum |
HPDC | 3 |
| 1992 | Network-based concurrent computing on the PVM systemabstractAbstract Concurrent computing environments based on loosely coupled networks have proven effective as resources for multiprocessing. Experiences with and enhancements to version 1.0 of PVM (Parallel Virtual Machine) are described in this paper. PVM is a software package that allows the utilization of a heterogeneous network of parallel and serial computers as a single computational resource. This report also describes an interactive graphical interface to PVM, and porting and performance results from production applications. Al Geist, Vaidy S. Sunderam |
Concurr. Pract. Exp. | 1 |
| 1992 | Parallel superconductor code on the iPSC/860
Al Geist, B. Ginatempo, William A. Shelton, G. Malcolm Stocks |
J. Supercomput. | 1 |
| 1992 | Algorithm 710: FORTRAN subroutines for computing the eigenvalues and eigenvectors of a general matrix by reduction to general tridiagonal formabstractThis paper describes programs to reduce a nonsymmetric matrix to tridiagonal form, to compute the eigenvalues of the tridiagonal matrix, to improve the accuracy of an eigenvalue, and to compute the corresponding eigenvector. The intended purpose of the software is to find a few eigenpairs of a dense nonsymmetric matrix faster and more accurately than previous methods. The performance and accuracy of the new routines are compared to two EISPACK paths: RG and HQR-INVIT. The results show that the new routines are more accurate and also faster if less than 20 percent of the eigenpairs are needed. Jack J. Dongarra, Al Geist, Charles H. Romine |
ACM Trans. Math. Softw. | 2 |
| 1990 | Finding eigenvalues and eigenvectors of unsymmetric matrices using a distributed-memory multiprocessor
Al Geist, George J. Davis |
Parallel Comput. | 1 |