EDBT 2026 Demo / reviewers in the wild / expert
Gerald Baumgartner
dblp:10/1300 · also Gerald B. Baumgartner
· DBLP profile ↗
28ranked-venue papers
4as first author
3since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 1 first-authorSoftware engineering, systems software and programming languages · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
High-performance computing · 97% Parallel and multicore computing · 3% | |
| Software engineering, system software, and programming languages
4 papers |
Programming languages and type systems · 47% Compilers and program optimization · 18% Software testing · 12% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computational science and engineering · 100% |
Topics — the 15 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing
quantum chemistry |
0.1 | 3 | 2005 | Performance modeling and optimization of parallel out-of-core tensor contractions · PPoPP 2005 A high-level approach to synthesis of high-performance codes for quantum chemistry · SC 2002 Space-Time Trade-Off Optimization for a Class of Electronic Structure Calculations · PLDI 2002 |
High-performance computing
scientific computing systems |
0.1 | 2 | 2002 | A high-level approach to synthesis of high-performance codes for quantum chemistry · SC 2002 Space-Time Trade-Off Optimization for a Class of Electronic Structure Calculations · PLDI 2002 |
High-performance computing › tensor computation
tensor contractions |
0.1 | 1 | 2005 | Performance modeling and optimization of parallel out-of-core tensor contractions · PPoPP 2005 |
Compilers and program optimization › code generation
high-performance code generation |
0.0 | 1 | 2002 | A high-level approach to synthesis of high-performance codes for quantum chemistry · SC 2002 |
Programming languages and type systems › object-oriented programming
object protocols |
0.0 | 1 | 2000 | Compiler and tool support for debugging object protocols · SIGSOFT FSE 2000 |
Software testing › protocol testing
protocol conformance checking |
0.0 | 1 | 2000 | Compiler and tool support for debugging object protocols · SIGSOFT FSE 2000 |
Debugging and program repair
runtime debugging |
0.0 | 1 | 2000 | Compiler and tool support for debugging object protocols · SIGSOFT FSE 2000 |
Program analysis
static analysis |
0.0 | 1 | 2000 | Compiler and tool support for debugging object protocols · SIGSOFT FSE 2000 |
Programming languages and type systems
type systems |
0.0 | 2 | 2000 | Implementing Signatures for C++ · ACM Trans. Program. Lang. Syst. 1997 Compiler and tool support for debugging object protocols · SIGSOFT FSE 2000 |
Programming languages and type systems
language design |
0.0 | 1 | 1997 | Implementing Signatures for C++ · ACM Trans. Program. Lang. Syst. 1997 |
Programming languages and type systems › language design
language extension |
0.0 | 1 | 1997 | Implementing Signatures for C++ · ACM Trans. Program. Lang. Syst. 1997 |
Programming languages and type systems › type systems › polymorphism
subtype polymorphism |
0.0 | 1 | 1997 | Implementing Signatures for C++ · ACM Trans. Program. Lang. Syst. 1997 |
Computational science and engineering › computational chemistry
electronic structure calculation |
0.0 | 1 | 2002 | Space-Time Trade-Off Optimization for a Class of Electronic Structure Calculations · PLDI 2002 |
Parallel and multicore computing › parallelizing compiler
parallel code generation |
0.0 | 1 | 2002 | A high-level approach to synthesis of high-performance codes for quantum chemistry · SC 2002 |
Compilers and program optimization
compiler construction |
0.0 | 1 | 1997 | Implementing Signatures for C++ · ACM Trans. Program. Lang. Syst. 1997 |
Methods — techniques the papers use, named apart from their topics
program synthesis · 0.1performance optimization · 0.1operation-minimal form analysis · 0.1memory-constrained optimization · 0.1tensor contraction · 0.1architecture-specific code generation · 0.1performance modeling · 0.1loop optimization · 0.1static checking · 0.0runtime checking · 0.0predicate association · 0.0preprocessor · 0.0compiler front-end · 0.0back-end support · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Variance Reduction of Resampling for Sequential Monte Carlo
Xiongming Dai, Gerald Baumgartner |
ICAART (3) | 2 |
| 2023 | Optimal Camera Configuration for Large-Scale Motion Capture Systems
Xiongming Dai, Gerald Baumgartner |
BMVC | 2 |
| 2022 | Latency-based Vector Scheduling of Many-task Applications for a Hybrid CloudabstractA centralized scheduler can become a bottleneck for placing the tasks of a many-task application on heterogeneous cloud resources. We have previously demonstrated that a de-centralized vector scheduling approach based on performance measurements can be used successfully for this task placement scenario. In this paper, we extend this approach to task placement based on latency measurements. Each node collects the performance measurements from its neighbors on an overlay graph, measures the communication latency, and then makes local decisions on where to move tasks. We present a centralized algorithm for configuring the overlay graph based on latency measurements and extend the vector scheduling approach to take latency into considerations. Our experiments in CloudLab demonstrate that this approach results in better performance and resource utilization than without latency information. Shifat P. Mithila, Gerald Baumgartner |
CLOUD | 2 |
| 2015 | A Hybrid Cloud Framework for Scientific ComputingabstractCloud services are transforming many computing tasks, but the unique requirements of scientific computing have caused it to lag behind in cloud adoption because of the performance variation of cloud resources. Based on our experience with the Organic Grid, we propose a framework for a hybrid cloud that will intelligently distribute work to appropriate computing resources to mitigate the impact of performance variation. We describe a cloud framework that integrates with specialized hardware and distributes work intelligently among heterogeneous computing resources. Our approach is to organize a set of computing nodes in an overlay network, to allow each node as an individual agent to position itself within the network to maximize its productivity. An application finds the resources and decides which task to run on which cloud nodes. Our simulations demonstrate that our methods can significantly reduce communication burdens of the most overworked nodes, especially on networks with the highest task to node ratios. Brian Peterson, Gerald Baumgartner, Qingyang Wang 0001 |
CLOUD | 2 |
| 2012 | Empirical performance model-driven data layout optimization and library call selection for tensor contraction expressions
Qingda Lu, Xiaoyang Gao 0002, Sriram Krishnamoorthy, Gerald Baumgartner, J. Ramanujam, P. Sadayappan |
J. Parallel Distributed Comput. | 4 |
| 2011 | Memory-optimal evaluation of expression trees involving large objects
Chi-Chung Lam, Thomas Rauber, Gerald Baumgartner, Daniel Cociorva, P. Sadayappan |
Comput. Lang. Syst. Struct. | 3 |
| 2007 | Efficient search-space pruning for integrated fusion and tiling transformationsabstractAbstract Compile‐time optimizations involve a number of transformations such as loop permutation, fusion, tiling, array contraction etc. The selection of the appropriate transformation to minimize the execution time is a challenging task. We address this problem in the context of tensor contraction expressions involving arrays too large to fit in main memory. Domain‐specific features of the computation are exploited to develop an integrated framework that facilitates the exploration of the entire search space of optimizations. In this paper, we discuss the exploration of the space of loop fusion and tiling transformations in order to minimize the disk I/O cost. These two transformations are integrated and pruning strategies are presented that significantly reduce the number of loop structures to be evaluated for subsequent transformations. The evaluation of the framework using representative contraction expressions from quantum chemistry shows a dramatic reduction in the size of the search space using the strategies presented. Copyright © 2007 John Wiley & Sons, Ltd. Xiaoyang Gao 0002, Sriram Krishnamoorthy, Swarup Kumar Sahoo, Chi-Chung Lam, Gerald Baumgartner, J. Ramanujam, P. Sadayappan |
Concurr. Comput. Pract. Exp. | 5 |
| 2006 | Memory minimization for tensor contractions using integer linear programmingabstractThis paper presents a technique for memory optimization for a class of computations that arises in the field of correlated electronic structure methods such as coupled cluster and configuration interaction methods in quantum chemistry. In this class of computations, loop computations perform a multi-dimensional sum of product of input arrays. There are many different ways to get the same final results that differ in the required number of arithmetic operations required. In addition, for a given number of arithmetic operations, different expressions of the loop have different memory requirements. Loop fusion is a plausible solution for reducing memory usage. By fusing loops between producer loop nest and consumer loop nest, the required storage of intermediate array is reduced by the range of the fused loop. Because resultant loops have to be legal after fusion, some loops can not be fused at the same time. In this paper, we have developed a novel integer linear programming (ILP) formulation that is shown to be highly effective on a number of test cases producing the optimal solutions using very small execution times. The main idea in the ILP formulation is the encoding of legality rules for loop fusion of a special class of loops using logical constraints over binary decision variables and a highly effective approximation of memory usage A. Allam, J. Ramanujam, Gerald Baumgartner, P. Sadayappan |
IPDPS | 3 |
| 2006 | Efficient synthesis of out-of-core algorithms using a nonlinear optimization solver
Sandhya Krishnan, Sriram Krishnamoorthy, Gerald Baumgartner, Chi-Chung Lam, J. Ramanujam, P. Sadayappan, Venkatesh Choppella |
J. Parallel Distributed Comput. | 3 |
| 2006 | Layout transformation support for the disk resident arrays framework
Sriram Krishnamoorthy, Gerald Baumgartner, Chi-Chung Lam, Jarek Nieplocha, P. Sadayappan |
J. Supercomput. | 2 |
| 2005 | Performance modeling and optimization of parallel out-of-core tensor contractionsabstractThe Tensor Contraction Engine (TCE) is a domain-specific compiler for implementing complex tensor contraction expressions arising in quantum chemistry applications modeling electronic structure. This paper develops a performance model for tensor contractions, considering both disk I/O as well as inter-processor communication costs, to facilitate performance-model driven loop optimization for this domain. Experimental results are provided that demonstrate the accuracy and effectiveness of the model. Xiaoyang Gao 0002, Swarup Kumar Sahoo, Chi-Chung Lam, J. Ramanujam, Qingda Lu, Gerald Baumgartner, P. Sadayappan |
PPoPP | 6 |
| 2005 | Synthesis of High-Performance Parallel Programs for a Class of ab Initio Quantum Chemistry ModelsabstractThis paper provides an overview of a program synthesis system for a class of quantum chemistry computations. These computations are expressible as a set of tensor contractions and arise in electronic structure modeling. The input to the system is a a high-level specification of the computation, from which the system can synthesize high-performance parallel code tailored to the characteristics of the target architecture. Several components of the synthesis system are described, focusing on performance optimization issues that they address. Gerald Baumgartner, Alexander A. Auer, David E. Bernholdt, Alina Bibireata, Venkatesh Choppella, Daniel Cociorva, Xiaoyang Gao 0002, Robert J. Harrison, So Hirata, Sriram Krishnamoorthy, Sandhya Krishnan, Chi-Chung Lam, Qingda Lu, Marcel Nooijen, Russell M. Pitzer, J. Ramanujam, P. Sadayappan, Alexander Sibiryakov |
Proc. IEEE | 1 |
| 2005 | The organic grid: self-organizing computation on a peer-to-peer networkabstractDesktop grids have been used to perform some of the largest computations in the world and have the potential to grow by several more orders of magnitude. However, current approaches to utilizing desktop resources require either centralized servers or extensive knowledge of the underlying system, limiting their scalability. We propose a new design for desktop grids that relies on a self-organizing, fully decentralized approach to the organization of the computation. Our approach, called the organic grid, is a radical departure from current approaches and is modeled after the way complex biological systems organize themselves. Similar to current desktop grids, a large computational task is broken down into sufficiently small subtasks. Each subtask is encapsulated into a mobile agent, which is then released on the grid and discovers computational resources using autonomous behavior. In the process of "colonization" of available resources, the judicious design of the agent behavior produces the emergence of crucial properties of the computation that can be tailored to specific classes of applications. We demonstrate this concept with a reduced-scale proof-of-concept implementation that executes a data-intensive independent-task application on a set of heterogeneous, geographically distributed machines. We present a detailed exploration of the design space of our system and a performance evaluation of our implementation using metrics appropriate for assessing self-organizing desktop grids. Arjav J. Chakravarti, Gerald Baumgartner, Mario Lauria |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2004 | Application-specific scheduling for the organic gridabstractSummary form only given. We propose a biologically inspired and fully-decentralized approach to the organization of computation that is based on the autonomous scheduling of strongly mobile agents on a peer-to-peer network. Our approach achieves the following design objectives: near-zero knowledge of network topology, zero knowledge of system status, autonomous scheduling, distributed computation, lack of specialized nodes. Every node is equally responsible for scheduling and computation, both of which are performed with practically no information about the system. We believe that this model is ideally suited for large-scale unstructured grids such as desktop grids. This model avoids the extensive system knowledge requirements of traditional grid scheduling approaches. Contrary to the popular master/worker organization of current desktop grids, our approach does not rely on specialized super-servers or on application-specific clients. By encapsulating computation and scheduling behavior into mobile agents, we decouple both application code and scheduling functionality from the underlying infrastructure. The resulting system is one where every node can start a large grid job, and where the computation naturally organizes itself around available resources. Through the careful design of agent behavior, the resulting global organization of the computation can be customized for different classes of applications. In a previous paper, we described a proof-of-concept prototype for an independent task application. We generalize the scheduling framework and demonstrate that our approach is applicable to a computation with a highly synchronous communication pattern, namely Cannon's matrix multiplication. Arjav J. Chakravarti, Gerald Baumgartner, Mario Lauria |
CLUSTER | 2 |
| 2004 | Efficient Layout Transformation for Disk-Based Multidimensional Arrays
Sriram Krishnamoorthy, Gerald Baumgartner, Chi-Chung Lam, Jarek Nieplocha, P. Sadayappan |
HiPC | 2 |
| 2004 | Efficient Synthesis of Out-of-Core Algorithms Using a Nonlinear Optimization SolverabstractSummary form only given. We address the problem of efficient out-of-core code generation for a special class of imperfectly nested loops encoding tensor contractions. These loops operate on arrays too large to fit in physical memory. The problem involves determining optimal tiling and placement of disk I/O statements. This entails a search in an explosively large parameter space. We formulate the problem as a nonlinear optimization problem and use a discrete constraint solver to generate optimized out-of-core code. Measurements on sequential and parallel versions of the generated code demonstrate the effectiveness of the proposed approach. Sandhya Krishnan, Sriram Krishnamoorthy, Gerald Baumgartner, Chi-Chung Lam, J. Ramanujam, P. Sadayappan, Venkatesh Choppella |
IPDPS | 3 |
| 2003 | Efficient Parallel Out-of-Core Matrix TranspositionabstractThis paper addresses the problem of parallel transposition of large out-of-core arrays. Although algorithms for out-of-core matrix transposition have been widely studied, previously proposed algorithms have sought to minimize the number of I/O operations and the in-memory permutation time. We propose an algorithm that directly targets the improvement of overall transposition time. The I/O characteristics of the system are used to determine the read, write and communication block sizes such that the total execution time is minimized. We also provide a solution to the array redistribution problem for arrays on disk. The solution to the sequential transposition problem and the parallel array redistribution problem are then combined to obtain an algorithm for the parallel out-of-core transposition problem. Sriram Krishnamoorthy, Gerald Baumgartner, Daniel Cociorva, Chi-Chung Lam, P. Sadayappan |
CLUSTER | 2 |
| 2003 | Data Locality Optimization for Synthesis of Efficient Out-of-Core Algorithms
Sandhya Krishnan, Sriram Krishnamoorthy, Gerald Baumgartner, Daniel Cociorva, Chi-Chung Lam, P. Sadayappan, J. Ramanujam, David E. Bernholdt, Venkatesh Choppella |
HiPC | 3 |
| 2003 | Implementation of Strong Mobility for Multi-Threaded Agents in JavaabstractStrong mobility, which allows multithreaded agents to be migrated transparently at any time, is a powerful mechanism for implementing a peer-to-peer computing environment, in which agents carrying a computational payload find available computing resources. Existing approaches to strong mobility either modify the Java virtual machine or do not correctly preserve the Java semantics when migrating multithreaded agents. We give an overview of our implementation strategy for strong mobility in which each agent thread maintains its own serializable execution state at all times, while thread states are captured just before a move. We explain how to solve the synchronization problems involved in migrating a multithreaded agent and how to cleanly terminate the Java threads in the originating virtual machine. We present experimental results that indicate that our implementation approach is feasible in practice. Arjav J. Chakravarti, Xiaojin Wang, Jason O. Hallstrom, Gerald Baumgartner |
ICPP | 4 |
| 2002 | Space-Time Trade-Off Optimization for a Class of Electronic Structure CalculationsabstractThe accurate modeling of the electronic structure of atoms and molecules is very computationally intensive. Many models of electronic structure, such as the Coupled Cluster approach, involve collections of tensor contractions. There are usually a large number of alternative ways of implementing the tensor contractions, representing different trade-offs between the space required for temporary intermediates and the total number of arithmetic operations. In this paper, we present an algorithm that starts with an operation-minimal form of the computation and systematically explores the possible space-time trade-offs to identify the form with lowest cost that fits within a specified memory limit. Its utility is demonstrated by applying it to a computation representative of a component in the CCSD(T) formulation in the NWChem quantum chemistry suite from Pacific Northwest National Laboratory. Daniel Cociorva, Gerald Baumgartner, Chi-Chung Lam, P. Sadayappan, J. Ramanujam, Marcel Nooijen, David E. Bernholdt, Robert J. Harrison |
PLDI | 2 |
| 2002 | A high-level approach to synthesis of high-performance codes for quantum chemistryabstractThis paper discusses an approach to the synthesis of high-performance parallel programs for a class of computations encountered in quantum chemistry and physics. These computations are expressible as a set of tensor contractions and arise in electronic structure modeling. An overview is provided of the synthesis system, that transforms a high-level specification of the computation into high-performance parallel code, tailored to the characteristics of the target architecture. An example from computational chemistry is used to illustrate how different code structures are generated under different assumptions of available memory on the target computer. Gerald Baumgartner, David E. Bernholdt, Daniel Cociorva, Robert J. Harrison, So Hirata, Chi-Chung Lam, Marcel Nooijen, Russell M. Pitzer, J. Ramanujam, P. Sadayappan |
SC | 1 |
| 2001 | Towards Automatic Synthesis of High-Performance Codes for Electronic Structure Calculations: Data Locality Optimization
Daniel Cociorva, J. W. Wilkins, Gerald Baumgartner, P. Sadayappan, J. Ramanujam, Marcel Nooijen, David E. Bernholdt, Robert J. Harrison |
HiPC | 3 |
| 2001 | Loop optimization for a class of memory-constrained computationsabstractCompute-intensive multi-dimensional summations that involve products of several arrays arise in the modeling of electronic structure of materials. Sometimes several alternative formulations of a computation, representing different space-time trade-offs, are possible. By computing and storing some intermediate arrays, reduction of the number of arithmetic operations is possible, but the size of intermediate temporary arrays may be prohibitively large. Loop fusion can be applied to reduce memory requirements, but that could impede effective tiling to minimize memory access costs. This paper develops an integrated model combining loop tiling for enhancing data reuse, and loop fusion for reduction of memory for intermediate temporary arrays. An algorithm is presented that addresses the selection of tile sizes and choice of loops for fusion, with the objective of minimizing cache misses while keeping the total memory usage within a given limit. Experimental results are reported that demonstrate the effectiveness of the combined loop tiling and fusion transformations performed by using the developed framework. Daniel Cociorva, J. W. Wilkins, Chi-Chung Lam, Gerald Baumgartner, J. Ramanujam, P. Sadayappan |
ICS | 4 |
| 2000 | Compiler and tool support for debugging object protocolsabstractWe describe an extension to the Java programming language that supports static conformance checking and dynamic debugging of object “protocols,” i.e., sequencing constraints on the order in which methods may be called. Our Java protocols have a statically checkable subset embedded in richer descriptions that can be checked at run time. The statically checkable subtype conformance relation is based on Nierstrasz' proposal for regular (finite-state) process types, and is also very close to the conformance relation for architectural connectors in the Wright architectural description language by Allen and Garlan. Richer sequencing properties, which cannot be expressed by regular types alone, can be specified and checked at run time by associating predicates with object states. We describe the language extensions and their rationale, and the design of tool support for static and dynamic checking and debugging. Sergey Butkevich, Marco Renedo, Gerald Baumgartner, Michal Young |
SIGSOFT FSE | 3 |
| 2000 | Safe Structural Conformance for JavaabstractIn Java, an interface specifies public abstract methods and associated public constants. Conformance of a class to an interface is by name. We propose to allow structural conformance to interfaces: any class or interface that declares or implements each method in a target interface conforms structurally to the interface, and any expression of the source class or interface type can be used where a value of the target interface type is expected. We argue that structural conformance results in a major gain in flexibility in situations that require retroactive abstraction over types. Structural conformance requires no additional syntax and only small modifications to the Java compiler and optionally, for performance reasons, the virtual machine, resulting in a minor performance penalty. Our extension is type-safe: a cast-free program that compiles without errors will not have any type errors at run time. Our extension is conservative: existing Java programs still compile and run in the same manner as under the original language definition. Finally, structural conformance works well with recent extensions such as Java remote method invocation. We have implemented our extension of Java with structural interface conformance by modifying the Java Developers Kit 1.1.5 source release for Solaris and Windows 95/NT. We have also created a test suite for the extension. Konstantin Läufer, Gerald Baumgartner, Vincent F. Russo |
Comput. J. | 2 |
| 1999 | Memory-Optimal Evaluation of Expression Trees Involving Large Objects
Chi-Chung Lam, Daniel Cociorva, Gerald Baumgartner, P. Sadayappan |
HiPC | 3 |
| 1997 | Implementing Signatures for C++abstractWe outline the design and detail the implementation of a language extension for abstracting types and for decoupling subtyping and inheritance in C++. This extension gives the user more of the flexibility of dynamic typing while retaining the efficiency and security of static typing. After a brief discussion of syntax and semantics of this language extension and examples of its use, we present and analyze three different implementation techniques: a preprocessor to a C++ compiler, an implementation in the front end of a C++ compiler, and a low-level implementation with back-end support. We follow with an analysis of the performance of the three implementation techniques and show that our extension actually allows subtype polymorphism to be implemented more efficiently than with virtual functions. We conclude with a discussion of the lessons we learned for future programming language design. Gerald Baumgartner, Vincent F. Russo |
ACM Trans. Program. Lang. Syst. | 1 |
| 1995 | Signatures: A Language Extension for Improving Type Abstraction and Subtype Polymorphism in C++abstractAbstract C++ uses inheritance as a substitute for subtype polymorphism. We give examples where this makes the type system too inflexible. We then describe a conservative language extension that allows a programmer to define an abstract type hierarchy independent of any implementation hierarchies, to retroactively abstract over an implementation, and to decouple subtyping from inheritance. This extension gives the user more of the flexibility of dynamic typing while retaining the efficiency and security of static typing. Withdefault implementationsandviewsflexible mechanisms are provided for implementing an abstract type by different concrete class types. We first show how the language extension can be implemented in a preprocessor to a C++ compiler, and then detail and analyse the efficiency of an implementation we directly incorporated in the GNU C++ compiler. Gerald Baumgartner, Vincent F. Russo |
Softw. Pract. Exp. | 1 |