VLDB 2026 Research / reviewers in the wild / expert
Charles Koelbel
dblp:c/CharlesKoelbel · also Chuck Koelbel
· DBLP profile ↗
20ranked-venue papers
5as first author
0since 2021 · last 2009
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Cloud and datacenter computing · 44% Distributed systems · 20% Parallel and multicore computing · 17% | |
| Software engineering, system software, and programming languages
5 papers |
Compilers and program optimization · 85% Programming languages and type systems · 15% |
Topics — the 22 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing
workflow scheduling |
0.2 | 3 | 2009 | VGrADS: enabling e-Science workflows on grids and clouds with fault tolerance · SC 2009 Grid scheduling and protocols - Evaluation of a workflow scheduler using integrated performance modelling and batch queue wait time prediction · SC 2006 Scheduling strategies for mapping application workflows onto the grid · HPDC 2005 |
Distributed systems
fault tolerance |
0.1 | 1 | 2009 | VGrADS: enabling e-Science workflows on grids and clouds with fault tolerance · SC 2009 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.1 | 1 | 2006 | Grid scheduling and protocols - Evaluation of a workflow scheduler using integrated performance modelling and batch queue wait time prediction · SC 2006 |
Performance modeling and evaluation
performance prediction |
0.1 | 1 | 2006 | Grid scheduling and protocols - Evaluation of a workflow scheduler using integrated performance modelling and batch queue wait time prediction · SC 2006 |
Compilers and program optimization
domain-specific compilation |
0.1 | 1 | 2005 | Telescoping Languages: A System for Automatic Generation of Domain Languages · Proc. IEEE 2005 |
Distributed systems
grid computing |
0.1 | 1 | 2005 | Scheduling strategies for mapping application workflows onto the grid · HPDC 2005 |
Parallel and multicore computing
load balancing |
0.1 | 1 | 2005 | Scheduling strategies for mapping application workflows onto the grid · HPDC 2005 |
Cloud and datacenter computing
resource management |
0.1 | 1 | 2005 | Scheduling strategies for mapping application workflows onto the grid · HPDC 2005 |
Programming languages and type systems › dynamic languages
scripting language |
0.0 | 1 | 2005 | Telescoping Languages: A System for Automatic Generation of Domain Languages · Proc. IEEE 2005 |
Electronic design automation › high-level synthesis
scheduling |
0.0 | 1 | 2005 | Scheduling strategies for mapping application workflows onto the grid · HPDC 2005 |
Parallel and multicore computing › data-parallel programming
data-parallel compilation |
0.0 | 1 | 1995 | A Model and Compilation Strategy for Out-of-Core Data Parallel Programs · PPoPP 1995 |
High-performance computing
parallel i/o |
0.0 | 1 | 1995 | A Model and Compilation Strategy for Out-of-Core Data Parallel Programs · PPoPP 1995 |
Parallel and multicore computing
parallel programming models |
0.0 | 2 | 1993 | Compiling Global Name-Space Parallel Loops for Distributed Execution · IEEE Trans. Parallel Distributed Syst. 1991 High performance Fortran: implementor and users workshop · SC 1993 |
Parallel and multicore computing › parallel programming models › data-parallel language
high performance fortran |
0.0 | 1 | 1993 | High performance Fortran: implementor and users workshop · SC 1993 |
Parallel and multicore computing › parallel computing
parallel programming languages |
0.0 | 1 | 1993 | Common runtime support for high-performance parallel languages · SC 1993 |
Parallel and multicore computing
parallel programming runtimes |
0.0 | 1 | 1993 | Common runtime support for high-performance parallel languages · SC 1993 |
Compilers and program optimization
parallelizing compiler |
0.0 | 1 | 1991 | Compiling Global Name-Space Parallel Loops for Distributed Execution · IEEE Trans. Parallel Distributed Syst. 1991 |
Parallel and multicore computing › parallel computing
distributed execution |
0.0 | 1 | 1991 | Compiling Global Name-Space Parallel Loops for Distributed Execution · IEEE Trans. Parallel Distributed Syst. 1991 |
Memory systems › shared memory
distributed shared memory |
0.0 | 1 | 1990 | Supporting Shared Data Structures on Distributed Memory Architectures · PPoPP 1990 |
Parallel and multicore computing
shared data structures |
0.0 | 1 | 1990 | Supporting Shared Data Structures on Distributed Memory Architectures · PPoPP 1990 |
Compilers and program optimization
parallel language compilation |
0.0 | 1 | 1993 | Common runtime support for high-performance parallel languages · SC 1993 |
Parallel and multicore computing › parallel programming models
message passing |
0.0 | 1 | 1991 | Compiling Global Name-Space Parallel Loops for Distributed Execution · IEEE Trans. Parallel Distributed Syst. 1991 |
Methods — techniques the papers use, named apart from their topics
virtualized reservations · 0.1resource selection · 0.1performance modeling · 0.1performance model based scheduling · 0.1library preprocessing · 0.1heuristic scheduling · 0.1annotation · 0.1runtime system · 0.0compile-time analysis · 0.0i/o optimization · 0.0data-parallel compilation · 0.0runtime code generation · 0.0run-time code generation · 0.0program analysis · 0.0message-passing transformation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2009 | Hybrid Re-scheduling Mechanisms for Workflow Applications on Multi-cluster GridabstractGrid computing is now a viable computational paradigm for executing large scale workflow applications. However, many aspects of performance optimization remain challenging. In this paper, we focus on the workflow scheduling mechanism. While there is much work on static scheduling approaches for workflow applications in parallel environments, little work has been done on a real-world multi-cluster grid environment. Since a typical grid environment is dynamic, we propose a new cluster-based scheduling mechanism that dynamically executes a top-down static scheduling algorithm using the real-time feedback from the execution monitor. We also propose a novel two phase migration mechanism that mitigates the effect of a possible bad reschedule decision. Our experimental results show that this approach achieves the best performance among all the scheduling approaches we implemented on both reserved resources and those with external loads. Charles Koelbel, Keith D. Cooper |
CCGRID | 2 |
| 2009 | Combined Fault Tolerance and Scheduling Techniques for Workflow Applications on Computational GridsabstractComplex scientific workflows are now Increasingly executed on computational grids. In addition to the challenges of managing and scheduling these workflows, reliability challenges arise because of the unreliable nature of large-scale grid infrastructure. Fault tolerance mechanisms like over-provisioning and checkpoint-recovery are used in current grid application management systems to address these reliability challenges. In this work, we propose new approaches that combine these fault tolerance techniques with existing workflow scheduling algorithms. We present a study on the effectiveness of the combined approaches by analyzing their impact on the reliability of workflow execution, workflow performance and resource usage under different reliability models, failure prediction accuracies and workflow application types. Anirban Mandal, Charles Koelbel, Keith D. Cooper |
CCGRID | 3 |
| 2009 | Batch queue resource scheduling for workflow applicationsabstractWorkflow computations have become a major programming paradigm for scientific applications. However, acquiring enough computational resources to execute a workflow poses a challenge in a batch queue controlled resource due to the space-sharing nature of the resource management policy. This paper introduces a scheduling technique that aggregates a workflow application into several subcomponents. It then uses the batch queue to acquire resources for each subcomponent, overlapping resource provisioning overhead (wait time) of one with the execution of others. We implemented a prototype of this technique and tested it using five high performance computing centers job submission logs. The results show that our approach can eliminate as much as 70% of the wait time over more traditional techniques that request resources for individual workflow nodes or that acquire all the resources for the whole workflow at once. Charles Koelbel, Keith D. Cooper |
CLUSTER | 2 |
| 2009 | VGrADS: enabling e-Science workflows on grids and clouds with fault toleranceabstractToday's scientific workflows use distributed heterogeneous resources through diverse grid and cloud interfaces that are often hard to program. In addition, especially for time-sensitive critical applications, predictable quality of service is necessary across these distributed resources. VGrADS' virtual grid execution system (vgES) provides an uniform qualitative resource abstraction over grid and cloud systems. We apply vgES for scheduling a set of deadline sensitive weather forecasting workflows. Specifically, this paper reports on our experiences with (1) virtualized reservations for batchqueue systems, (2) coordinated usage of TeraGrid (batch queue), Amazon EC2 (cloud), our own clusters (batch queue) and Eucalyptus (cloud) resources, and (3) fault tolerance through automated task replication. The combined effect of these techniques was to enable a new workflow planning method to balance performance, reliability and cost considerations. The results point toward improved resource selection and execution management support for a variety of e-Science applications over grids and cloud systems. Lavanya Ramakrishnan, Charles Koelbel, Yang-Suk Kee, Richard Wolski, Daniel Nurmi, Dennis Gannon, Graziano Obertelli, Asim YarKhan, Anirban Mandal, T. Mark Huang, Kiran Thyagaraja, Dmitrii Zagorodnov |
SC | 2 |
| 2008 | Cluster-Based Hybrid Scheduling Mechanisms for Workflow Applications on the GridabstractThanks to advances in wide-area network technologies and the decreasing cost of computing resources, Grid computing is now a viable computational paradigm. However, many aspects of successfully using the Grid remain research topics. Among them, we identify scheduling of workflow applications as a key problem. While there is much work on static scheduling approaches for workflow applications in parallel environments, little work has been done on a real-world Grid environment. In this paper, we launch four model workflow applications with different configurations on a multi-cluster Grid testbed. By observing the applications' performance, we propose a new cluster-based hybrid scheduling mechanism that dynamically executes a top-down static scheduling algorithm using the real-time feedback from the execution monitor. Our experimental results show that this approach achieves the best performance among all the scheduling approaches we implemented on both reserved resources and those with external loads. Charles Koelbel, Keith D. Cooper |
eScience | 2 |
| 2007 | Relative Performance of Scheduling Algorithms in Grid EnvironmentsabstractEffective scheduling is critical for the performance of an application launched onto the Grid environment. Finding effective scheduling algorithms for this problem is a challenging research area. Many scheduling algorithms have been proposed, studied and compared on heterogeneous parallel computers but there are few studies comparing the performance of scheduling algorithms in Grid environments. The Grid is unique because of the drastic cost differences between inter-cluster and the intra-cluster data transfers. In this paper, we compare several scheduling algorithms that represent two classes of schedulers used for Grid computing. We analyze the results to explain how different resource environments and workflow application structures affect the performance of these algorithms. Based on our experiments, we introduce a new measurement called effective aggregated computing power (EACP) that could drastically improve the performance of some schedulers. Charles Koelbel, Ken Kennedy |
CCGRID | 2 |
| 2006 | Scalable Grid Application Scheduling via Decoupled Resource Selection and SchedulingabstractOver the past years grid infrastructures have been deployed at larger and larger scales, with envisioned deployments incorporating tens of thousands of resources. Therefore, application scheduling algorithms can become unscalable (albeit polynomial) and thus unusable in large-scale environments. One reason for unscalability is that these algorithms perform implicit resource selection. One can achieve better scalability by performing explicit resource selection independently from scheduling in a "decoupled' approach. Furthermore, we hypothesize that one can achieve similar or even better performance as with the non-decoupled approach, which we call the "one step" approach, by selecting resources judiciously. Leveraging the Virtual Grid abstraction, we demonstrate that the decoupled approach is indeed both scalable and effective in large-scale and highly heterogeneous resource environments. Anirban Mandal, Henri Casanova, Andrew A. Chien, Yang-Suk Kee, Ken Kennedy, Charles Koelbel |
CCGRID | 7 |
| 2006 | Grid scheduling and protocols - Evaluation of a workflow scheduler using integrated performance modelling and batch queue wait time predictionabstractLarge-scale distributed systems offer computational power at unprecedented levels. In the past, HPC users typically had access to relatively few individual supercomputers and, in general, would assign a one-to-one mapping of applications to machines. Modern HPC users have simultaneous access to a large number of individual machines and are beginning to make use of all of them for single-application execution cycles. One method that application developers have devised in order to take advantage of such systems is to organize an entire application execution cycle as a workflow. The scheduling of such workflows has been the topic of a great deal of research in the past few years and, although very sophisticated algorithms have been devised, a very specific aspect of these distributed systems, namely that most supercomputing resources employ batch queue scheduling software, has heretofore been omitted from consideration, presumably because it is difficult to model accurately. In this work, we augment an existing workflow scheduler through the introduction of methods which make accurate predictions of both the performance of the application on specific hardware, and the amount of time individual workflow tasks will spend waiting in batch queues. Our results show that although a workflow scheduler alone may choose correct task placement based on data locality or network connectivity, this benefit is often compromised by the fact that most jobs submitted to current systems must wait in overcommited batch queues for a significant portion of time. However, incorporating the enhancements we describe improves workflow execution time in settings where batch queues impose significant delays on constituent workflow tasks. Daniel Nurmi, Anirban Mandal, John Brevik, Charles Koelbel, Richard Wolski, Ken Kennedy |
SC | 4 |
| 2005 | Scheduling strategies for mapping application workflows onto the gridabstractIn this work, we describe new strategies for scheduling and executing workflow applications on grid resources using the GrADS [Ken Kennedy et al., 2002] infrastructure. Workflow scheduling is based on heuristic scheduling strategies that use application component performance models. The workflow is executed using a novel strategy to bind and launch the application onto heterogeneous resources. We apply these strategies in the context of executing EMAN, a bio-imaging workflow application, on the grid. The results of our experiments show that our strategy of performance model based, in-advance heuristic workflow scheduling results in 1.5 to 2.2 times better makespan than other existing scheduling strategies. This strategy also achieves optimal load balance across the different grid sites for this application. Anirban Mandal, Ken Kennedy, Charles Koelbel, Gabriel Marin, John M. Mellor-Crummey, S. Lennart Johnsson |
HPDC | 3 |
| 2005 | Telescoping Languages: A System for Automatic Generation of Domain LanguagesabstractThe software gap - the discrepancy between the need for new software and the aggregate capacity of the workforce to produce it - is a serious problem for scientific software. Although users appreciate the convenience (and, thus, improved productivity) of using relatively high-level scripting languages, the slow execution speeds of these languages remain a problem. Lower level languages, such as C and Fortran, provide better performance for production applications, but at the cost of tedious programming and optimization by experts. If applications written in scripting languages could be routinely compiled into highly optimized machine code, a huge productivity advantage would be possible. It is not enough, however, to simply develop excellent compiler technologies for scripting languages (as a number of projects have succeeded in doing for MATLAB). In practice, scientists typically extend these languages with their own domain-centric components, such as the MATLAB signal processing toolbox. Doing so effectively defines a new domain-specific language. If we are to address efficiency problems for such extended languages, we must develop a framework for automatically generating optimizing compilers for them. To accomplish this goal, we have been pursuing an innovative strategy that we call telescoping languages. Our approach calls for using a library-preprocessing phase to extensively analyze and optimize collections of libraries that define an extended language. Results of this analysis are collected into annotated libraries and used to generate a library-aware optimizer. The generated library-aware optimizer uses the knowledge gathered during preprocessing to carry out fast and effective optimization of high-level scripts. This enables script optimization to benefit from the intense analysis performed during preprocessing without repaying its price. Since library preprocessing is performed only at infrequent "language-generation" times, its cost is amortized over many Ken Kennedy, Bradley Broom, Arun Chauhan 0001, Robert J. Fowler, John Garvin, Charles Koelbel, Cheryl McCosh, John M. Mellor-Crummey |
Proc. IEEE | 6 |
| 2004 | Scheduling workflow applications in GrADSabstractIn this work, we describe new strategies for scheduling and executing workflow applications on Grid resources using the GrADS infrastructure. Workflow scheduling is based on heuristic scheduling strategies that use combined computational and memory hierarchy application component performance models. The workflow is executed using a novel strategy to bind and launch the application onto heterogeneous resources. We apply these strategies in the context of launching EMAN, a bio-imaging workflow application, onto the Grid. Anirban Mandal, Anshuman Dasgupta, Ken Kennedy, Mark Mazina, Charles Koelbel, Gabriel Marin, Keith D. Cooper, John M. Mellor-Crummey, S. Lennart Johnsson |
CCGRID | 5 |
| 1995 | A Model and Compilation Strategy for Out-of-Core Data Parallel ProgramsabstractIt is widely acknowledged in high-performance computing circles that parallel input/output needs substantial improvement in order to make scalable computers truly usable. We present a data storage model that allows processors independent access to their own data and a corresponding compilation strategy that integrates data-parallel computation with data distribution for out-of-core problems. Our results compare several communication methods and I/O optimizations using two out-of-core problems, Jacobi iteration and LU factorization. Rajesh Bordawekar, Alok N. Choudhary, Ken Kennedy, Charles Koelbel, Michael H. Paleczny |
PPoPP | 4 |
| 1993 | High performance Fortran: implementor and users workshopabstractArticle High performance Fortran: implementor and users workshop Share on Authors: A. Choudhary Syracuse University Syracuse, NY Syracuse University Syracuse, NYView Profile , C. Koelbel Rice University, Houston, TX Rice University, Houston, TXView Profile , M. Zosel Lawrence Livermore, National Lab Lawrence Livermore, National LabView Profile Authors Info & Claims Supercomputing '93: Proceedings of the 1993 ACM/IEEE conference on SupercomputingDecember 1993 Pages 610–613https://doi.org/10.1145/169627.169808Online:01 December 1993Publication History 1citation133DownloadsMetricsTotal Citations1Total Downloads133Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Alok N. Choudhary, Charles Koelbel, Mary Zosel |
SC | 2 |
| 1993 | Common runtime support for high-performance parallel languagesabstractNo abstract available. Geoffrey C. Fox, Sanjay Ranka, Michael L. Scott, Allen D. Malony, James C. Browne, Marina C. Chen, Alok N. Choudhary, Thomas E. Cheatham, Janice E. Cuny, Rudolf Eigenmann, Amr F. Fahmy, Ian T. Foster, Dennis Gannon, Tomasz Haupt, Carl Kesselman, Charles Koelbel, Wei Li 0015, Monica S. Lam, Thomas J. LeBlanc, Jim Openshaw, David A. Padua, Constantine D. Polychronopoulos, Joel H. Saltz, Alan Sussman, Gil Weigand, Katherine A. Yelick |
SC | 16 |
| 1991 | Programming data parallel algorithms on distributed memory using KaliabstractCurrent languages for distributed memory machines tend to directly reflect the underlying hardware and thus provide little support for implementing data parallel algorithms.This paper describes a programming environment, Kali, which provides a global name space and allows direct access to remote data values.In order to retain efficiency, Kali provides a system of annotations allowing the user to control aspects of the program critical to performance, such as data distribution and load balancing.We present a series of examples showing how Kali can easily express data parallel algorithms.We also discuss some of the issues raised in translating such programs for execution on distributed memory systems and show performance results for well-known numerical algorithms written in Kali. Charles Koelbel, Piyush Mehrotra |
ICS | 1 |
| 1991 | Compile-time generation of regular communications patterns
Charles Koelbel |
SC | 1 |
| 1991 | Compiling Global Name-Space Parallel Loops for Distributed ExecutionabstractCompiler support required to allow programmers to express their algorithms using a global name-space is discussed. A general method for the analysis of a high-level source program and its translation into a set of independently executing tasks that communicate using messages is presented. It is shown that if the compiler has enough information, the translation can be carried out at compile time; otherwise; run-time code is generated to implement the required data movement. The analysis required in both situations is described, and the performance of the generated code on the Intel iPSC/2 hypercube is presented.> Charles Koelbel, Piyush Mehrotra |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1990 | Supporting Shared Data Structures on Distributed Memory ArchitecturesabstractProgramming nonshared memory systems is more difficult than programming shared memory systems, since there is no support for shared data structures. Current programming languages for distributed memory architectures force the user to decompose all data structures into separate pieces, with each piece “owned” by one of the processors in the machine, and with all communication explicitly specified by low-level message-passing primitives. This paper presents a new programming environment for distributed memory architectures, providing a global name space and allowing direct access to remote parts of data values. We describe the analysis and program transformations required to implement this environment, and present the efficiency of the resulting code on the NCUBE/7 and IPSC/2 hypercubes. Charles Koelbel, Piyush Mehrotra, John Van Rosendale |
PPoPP | 1 |
| 1987 | Semi-Automatic Domain Decomposition in BLAZE
Charles Koelbel, Piyush Mehrotra, John Van Rosendale |
ICPP | 1 |
| 1985 | Prep-P: A Mapping Preprocessor for CHiP Architectures
Francine Berman, Michael A. Goodrich, Charles Koelbel, W. J. Robison III, Karen Showell |
ICPP | 3 |