Charles Koelbel

dblp:c/CharlesKoelbel · also Chuck Koelbel · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
0since 2021 · last 2009
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Cloud and datacenter computing · 44% Distributed systems · 20% Parallel and multicore computing · 17%
Software engineering, system software, and programming languages
5 papers
Compilers and program optimization · 85% Programming languages and type systems · 15%

Topics — the 22 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
workflow scheduling
0.232009
VGrADS: enabling e-Science workflows on grids and clouds with fault tolerance · SC 2009
Grid scheduling and protocols - Evaluation of a workflow scheduler using integrated performance modelling and batch queue wait time prediction · SC 2006
Scheduling strategies for mapping application workflows onto the grid · HPDC 2005
Distributed systems
fault tolerance
0.112009
VGrADS: enabling e-Science workflows on grids and clouds with fault tolerance · SC 2009
Cloud and datacenter computing
cluster resource management and scheduling
0.112006
Grid scheduling and protocols - Evaluation of a workflow scheduler using integrated performance modelling and batch queue wait time prediction · SC 2006
Performance modeling and evaluation
performance prediction
0.112006
Grid scheduling and protocols - Evaluation of a workflow scheduler using integrated performance modelling and batch queue wait time prediction · SC 2006
Compilers and program optimization
domain-specific compilation
0.112005
Telescoping Languages: A System for Automatic Generation of Domain Languages · Proc. IEEE 2005
Distributed systems
grid computing
0.112005
Scheduling strategies for mapping application workflows onto the grid · HPDC 2005
Parallel and multicore computing
load balancing
0.112005
Scheduling strategies for mapping application workflows onto the grid · HPDC 2005
Cloud and datacenter computing
resource management
0.112005
Scheduling strategies for mapping application workflows onto the grid · HPDC 2005
Programming languages and type systems › dynamic languages
scripting language
0.012005
Telescoping Languages: A System for Automatic Generation of Domain Languages · Proc. IEEE 2005
Electronic design automation › high-level synthesis
scheduling
0.012005
Scheduling strategies for mapping application workflows onto the grid · HPDC 2005
Parallel and multicore computing › data-parallel programming
data-parallel compilation
0.011995
A Model and Compilation Strategy for Out-of-Core Data Parallel Programs · PPoPP 1995
High-performance computing
parallel i/o
0.011995
A Model and Compilation Strategy for Out-of-Core Data Parallel Programs · PPoPP 1995
Parallel and multicore computing
parallel programming models
0.021993
Compiling Global Name-Space Parallel Loops for Distributed Execution · IEEE Trans. Parallel Distributed Syst. 1991
High performance Fortran: implementor and users workshop · SC 1993
Parallel and multicore computing › parallel programming models › data-parallel language
high performance fortran
0.011993
High performance Fortran: implementor and users workshop · SC 1993
Parallel and multicore computing › parallel computing
parallel programming languages
0.011993
Common runtime support for high-performance parallel languages · SC 1993
Parallel and multicore computing
parallel programming runtimes
0.011993
Common runtime support for high-performance parallel languages · SC 1993
Compilers and program optimization
parallelizing compiler
0.011991
Compiling Global Name-Space Parallel Loops for Distributed Execution · IEEE Trans. Parallel Distributed Syst. 1991
Parallel and multicore computing › parallel computing
distributed execution
0.011991
Compiling Global Name-Space Parallel Loops for Distributed Execution · IEEE Trans. Parallel Distributed Syst. 1991
Memory systems › shared memory
distributed shared memory
0.011990
Supporting Shared Data Structures on Distributed Memory Architectures · PPoPP 1990
Parallel and multicore computing
shared data structures
0.011990
Supporting Shared Data Structures on Distributed Memory Architectures · PPoPP 1990
Compilers and program optimization
parallel language compilation
0.011993
Common runtime support for high-performance parallel languages · SC 1993
Parallel and multicore computing › parallel programming models
message passing
0.011991
Compiling Global Name-Space Parallel Loops for Distributed Execution · IEEE Trans. Parallel Distributed Syst. 1991

Methods — techniques the papers use, named apart from their topics

virtualized reservations · 0.1resource selection · 0.1performance modeling · 0.1performance model based scheduling · 0.1library preprocessing · 0.1heuristic scheduling · 0.1annotation · 0.1runtime system · 0.0compile-time analysis · 0.0i/o optimization · 0.0data-parallel compilation · 0.0runtime code generation · 0.0run-time code generation · 0.0program analysis · 0.0message-passing transformation · 0.0
YearPublicationVenuePosition
2009 Hybrid Re-scheduling Mechanisms for Workflow Applications on Multi-cluster Grid
abstract
Grid computing is now a viable computational paradigm for executing large scale workflow applications. However, many aspects of performance optimization remain challenging. In this paper, we focus on the workflow scheduling mechanism. While there is much work on static scheduling approaches for workflow applications in parallel environments, little work has been done on a real-world multi-cluster grid environment. Since a typical grid environment is dynamic, we propose a new cluster-based scheduling mechanism that dynamically executes a top-down static scheduling algorithm using the real-time feedback from the execution monitor. We also propose a novel two phase migration mechanism that mitigates the effect of a possible bad reschedule decision. Our experimental results show that this approach achieves the best performance among all the scheduling approaches we implemented on both reserved resources and those with external loads.
Charles Koelbel, Keith D. Cooper
CCGRID2
2009 Combined Fault Tolerance and Scheduling Techniques for Workflow Applications on Computational Grids
abstract
Complex scientific workflows are now Increasingly executed on computational grids. In addition to the challenges of managing and scheduling these workflows, reliability challenges arise because of the unreliable nature of large-scale grid infrastructure. Fault tolerance mechanisms like over-provisioning and checkpoint-recovery are used in current grid application management systems to address these reliability challenges. In this work, we propose new approaches that combine these fault tolerance techniques with existing workflow scheduling algorithms. We present a study on the effectiveness of the combined approaches by analyzing their impact on the reliability of workflow execution, workflow performance and resource usage under different reliability models, failure prediction accuracies and workflow application types.
Anirban Mandal, Charles Koelbel, Keith D. Cooper
CCGRID3
2009 Batch queue resource scheduling for workflow applications
abstract
Workflow computations have become a major programming paradigm for scientific applications. However, acquiring enough computational resources to execute a workflow poses a challenge in a batch queue controlled resource due to the space-sharing nature of the resource management policy. This paper introduces a scheduling technique that aggregates a workflow application into several subcomponents. It then uses the batch queue to acquire resources for each subcomponent, overlapping resource provisioning overhead (wait time) of one with the execution of others. We implemented a prototype of this technique and tested it using five high performance computing centers job submission logs. The results show that our approach can eliminate as much as 70% of the wait time over more traditional techniques that request resources for individual workflow nodes or that acquire all the resources for the whole workflow at once.
Charles Koelbel, Keith D. Cooper
CLUSTER2
2009 VGrADS: enabling e-Science workflows on grids and clouds with fault tolerance
abstract
Today's scientific workflows use distributed heterogeneous resources through diverse grid and cloud interfaces that are often hard to program. In addition, especially for time-sensitive critical applications, predictable quality of service is necessary across these distributed resources. VGrADS' virtual grid execution system (vgES) provides an uniform qualitative resource abstraction over grid and cloud systems. We apply vgES for scheduling a set of deadline sensitive weather forecasting workflows. Specifically, this paper reports on our experiences with (1) virtualized reservations for batchqueue systems, (2) coordinated usage of TeraGrid (batch queue), Amazon EC2 (cloud), our own clusters (batch queue) and Eucalyptus (cloud) resources, and (3) fault tolerance through automated task replication. The combined effect of these techniques was to enable a new workflow planning method to balance performance, reliability and cost considerations. The results point toward improved resource selection and execution management support for a variety of e-Science applications over grids and cloud systems.
Lavanya Ramakrishnan, Charles Koelbel, Yang-Suk Kee, Richard Wolski, Daniel Nurmi, Dennis Gannon, Graziano Obertelli, Asim YarKhan, Anirban Mandal, T. Mark Huang, Kiran Thyagaraja, Dmitrii Zagorodnov
SC2
2008 Cluster-Based Hybrid Scheduling Mechanisms for Workflow Applications on the Grid
abstract
Thanks to advances in wide-area network technologies and the decreasing cost of computing resources, Grid computing is now a viable computational paradigm. However, many aspects of successfully using the Grid remain research topics. Among them, we identify scheduling of workflow applications as a key problem. While there is much work on static scheduling approaches for workflow applications in parallel environments, little work has been done on a real-world Grid environment. In this paper, we launch four model workflow applications with different configurations on a multi-cluster Grid testbed. By observing the applications' performance, we propose a new cluster-based hybrid scheduling mechanism that dynamically executes a top-down static scheduling algorithm using the real-time feedback from the execution monitor. Our experimental results show that this approach achieves the best performance among all the scheduling approaches we implemented on both reserved resources and those with external loads.
Charles Koelbel, Keith D. Cooper
eScience2
2007 Relative Performance of Scheduling Algorithms in Grid Environments
abstract
Effective scheduling is critical for the performance of an application launched onto the Grid environment. Finding effective scheduling algorithms for this problem is a challenging research area. Many scheduling algorithms have been proposed, studied and compared on heterogeneous parallel computers but there are few studies comparing the performance of scheduling algorithms in Grid environments. The Grid is unique because of the drastic cost differences between inter-cluster and the intra-cluster data transfers. In this paper, we compare several scheduling algorithms that represent two classes of schedulers used for Grid computing. We analyze the results to explain how different resource environments and workflow application structures affect the performance of these algorithms. Based on our experiments, we introduce a new measurement called effective aggregated computing power (EACP) that could drastically improve the performance of some schedulers.
Charles Koelbel, Ken Kennedy
CCGRID2
2006 Scalable Grid Application Scheduling via Decoupled Resource Selection and Scheduling
abstract
Over the past years grid infrastructures have been deployed at larger and larger scales, with envisioned deployments incorporating tens of thousands of resources. Therefore, application scheduling algorithms can become unscalable (albeit polynomial) and thus unusable in large-scale environments. One reason for unscalability is that these algorithms perform implicit resource selection. One can achieve better scalability by performing explicit resource selection independently from scheduling in a "decoupled' approach. Furthermore, we hypothesize that one can achieve similar or even better performance as with the non-decoupled approach, which we call the "one step" approach, by selecting resources judiciously. Leveraging the Virtual Grid abstraction, we demonstrate that the decoupled approach is indeed both scalable and effective in large-scale and highly heterogeneous resource environments.
Anirban Mandal, Henri Casanova, Andrew A. Chien, Yang-Suk Kee, Ken Kennedy, Charles Koelbel
CCGRID7
2006 Grid scheduling and protocols - Evaluation of a workflow scheduler using integrated performance modelling and batch queue wait time prediction
abstract
Large-scale distributed systems offer computational power at unprecedented levels. In the past, HPC users typically had access to relatively few individual supercomputers and, in general, would assign a one-to-one mapping of applications to machines. Modern HPC users have simultaneous access to a large number of individual machines and are beginning to make use of all of them for single-application execution cycles. One method that application developers have devised in order to take advantage of such systems is to organize an entire application execution cycle as a workflow. The scheduling of such workflows has been the topic of a great deal of research in the past few years and, although very sophisticated algorithms have been devised, a very specific aspect of these distributed systems, namely that most supercomputing resources employ batch queue scheduling software, has heretofore been omitted from consideration, presumably because it is difficult to model accurately. In this work, we augment an existing workflow scheduler through the introduction of methods which make accurate predictions of both the performance of the application on specific hardware, and the amount of time individual workflow tasks will spend waiting in batch queues. Our results show that although a workflow scheduler alone may choose correct task placement based on data locality or network connectivity, this benefit is often compromised by the fact that most jobs submitted to current systems must wait in overcommited batch queues for a significant portion of time. However, incorporating the enhancements we describe improves workflow execution time in settings where batch queues impose significant delays on constituent workflow tasks.
Daniel Nurmi, Anirban Mandal, John Brevik, Charles Koelbel, Richard Wolski, Ken Kennedy
SC4
2005 Scheduling strategies for mapping application workflows onto the grid
abstract
In this work, we describe new strategies for scheduling and executing workflow applications on grid resources using the GrADS [Ken Kennedy et al., 2002] infrastructure. Workflow scheduling is based on heuristic scheduling strategies that use application component performance models. The workflow is executed using a novel strategy to bind and launch the application onto heterogeneous resources. We apply these strategies in the context of executing EMAN, a bio-imaging workflow application, on the grid. The results of our experiments show that our strategy of performance model based, in-advance heuristic workflow scheduling results in 1.5 to 2.2 times better makespan than other existing scheduling strategies. This strategy also achieves optimal load balance across the different grid sites for this application.
Anirban Mandal, Ken Kennedy, Charles Koelbel, Gabriel Marin, John M. Mellor-Crummey, S. Lennart Johnsson
HPDC3
2005 Telescoping Languages: A System for Automatic Generation of Domain Languages
abstract
The software gap - the discrepancy between the need for new software and the aggregate capacity of the workforce to produce it - is a serious problem for scientific software. Although users appreciate the convenience (and, thus, improved productivity) of using relatively high-level scripting languages, the slow execution speeds of these languages remain a problem. Lower level languages, such as C and Fortran, provide better performance for production applications, but at the cost of tedious programming and optimization by experts. If applications written in scripting languages could be routinely compiled into highly optimized machine code, a huge productivity advantage would be possible. It is not enough, however, to simply develop excellent compiler technologies for scripting languages (as a number of projects have succeeded in doing for MATLAB). In practice, scientists typically extend these languages with their own domain-centric components, such as the MATLAB signal processing toolbox. Doing so effectively defines a new domain-specific language. If we are to address efficiency problems for such extended languages, we must develop a framework for automatically generating optimizing compilers for them. To accomplish this goal, we have been pursuing an innovative strategy that we call telescoping languages. Our approach calls for using a library-preprocessing phase to extensively analyze and optimize collections of libraries that define an extended language. Results of this analysis are collected into annotated libraries and used to generate a library-aware optimizer. The generated library-aware optimizer uses the knowledge gathered during preprocessing to carry out fast and effective optimization of high-level scripts. This enables script optimization to benefit from the intense analysis performed during preprocessing without repaying its price. Since library preprocessing is performed only at infrequent "language-generation" times, its cost is amortized over many
Ken Kennedy, Bradley Broom, Arun Chauhan 0001, Robert J. Fowler, John Garvin, Charles Koelbel, Cheryl McCosh, John M. Mellor-Crummey
Proc. IEEE6
2004 Scheduling workflow applications in GrADS
abstract
In this work, we describe new strategies for scheduling and executing workflow applications on Grid resources using the GrADS infrastructure. Workflow scheduling is based on heuristic scheduling strategies that use combined computational and memory hierarchy application component performance models. The workflow is executed using a novel strategy to bind and launch the application onto heterogeneous resources. We apply these strategies in the context of launching EMAN, a bio-imaging workflow application, onto the Grid.
Anirban Mandal, Anshuman Dasgupta, Ken Kennedy, Mark Mazina, Charles Koelbel, Gabriel Marin, Keith D. Cooper, John M. Mellor-Crummey, S. Lennart Johnsson
CCGRID5
1995 A Model and Compilation Strategy for Out-of-Core Data Parallel Programs
abstract
It is widely acknowledged in high-performance computing circles that parallel input/output needs substantial improvement in order to make scalable computers truly usable. We present a data storage model that allows processors independent access to their own data and a corresponding compilation strategy that integrates data-parallel computation with data distribution for out-of-core problems. Our results compare several communication methods and I/O optimizations using two out-of-core problems, Jacobi iteration and LU factorization.
Rajesh Bordawekar, Alok N. Choudhary, Ken Kennedy, Charles Koelbel, Michael H. Paleczny
PPoPP4
1993 High performance Fortran: implementor and users workshop
abstract
Article High performance Fortran: implementor and users workshop Share on Authors: A. Choudhary Syracuse University Syracuse, NY Syracuse University Syracuse, NYView Profile , C. Koelbel Rice University, Houston, TX Rice University, Houston, TXView Profile , M. Zosel Lawrence Livermore, National Lab Lawrence Livermore, National LabView Profile Authors Info & Claims Supercomputing '93: Proceedings of the 1993 ACM/IEEE conference on SupercomputingDecember 1993 Pages 610–613https://doi.org/10.1145/169627.169808Online:01 December 1993Publication History 1citation133DownloadsMetricsTotal Citations1Total Downloads133Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Alok N. Choudhary, Charles Koelbel, Mary Zosel
SC2
1993 Common runtime support for high-performance parallel languages
abstract
No abstract available.
Geoffrey C. Fox, Sanjay Ranka, Michael L. Scott, Allen D. Malony, James C. Browne, Marina C. Chen, Alok N. Choudhary, Thomas E. Cheatham, Janice E. Cuny, Rudolf Eigenmann, Amr F. Fahmy, Ian T. Foster, Dennis Gannon, Tomasz Haupt, Carl Kesselman, Charles Koelbel, Wei Li 0015, Monica S. Lam, Thomas J. LeBlanc, Jim Openshaw, David A. Padua, Constantine D. Polychronopoulos, Joel H. Saltz, Alan Sussman, Gil Weigand, Katherine A. Yelick
SC16
1991 Programming data parallel algorithms on distributed memory using Kali
abstract
Current languages for distributed memory machines tend to directly reflect the underlying hardware and thus provide little support for implementing data parallel algorithms.This paper describes a programming environment, Kali, which provides a global name space and allows direct access to remote data values.In order to retain efficiency, Kali provides a system of annotations allowing the user to control aspects of the program critical to performance, such as data distribution and load balancing.We present a series of examples showing how Kali can easily express data parallel algorithms.We also discuss some of the issues raised in translating such programs for execution on distributed memory systems and show performance results for well-known numerical algorithms written in Kali.
Charles Koelbel, Piyush Mehrotra
ICS1
1991 Compile-time generation of regular communications patterns
Charles Koelbel
SC1
1991 Compiling Global Name-Space Parallel Loops for Distributed Execution
abstract
Compiler support required to allow programmers to express their algorithms using a global name-space is discussed. A general method for the analysis of a high-level source program and its translation into a set of independently executing tasks that communicate using messages is presented. It is shown that if the compiler has enough information, the translation can be carried out at compile time; otherwise; run-time code is generated to implement the required data movement. The analysis required in both situations is described, and the performance of the generated code on the Intel iPSC/2 hypercube is presented.>
Charles Koelbel, Piyush Mehrotra
IEEE Trans. Parallel Distributed Syst.1
1990 Supporting Shared Data Structures on Distributed Memory Architectures
abstract
Programming nonshared memory systems is more difficult than programming shared memory systems, since there is no support for shared data structures. Current programming languages for distributed memory architectures force the user to decompose all data structures into separate pieces, with each piece “owned” by one of the processors in the machine, and with all communication explicitly specified by low-level message-passing primitives. This paper presents a new programming environment for distributed memory architectures, providing a global name space and allowing direct access to remote parts of data values. We describe the analysis and program transformations required to implement this environment, and present the efficiency of the resulting code on the NCUBE/7 and IPSC/2 hypercubes.
Charles Koelbel, Piyush Mehrotra, John Van Rosendale
PPoPP1
1987 Semi-Automatic Domain Decomposition in BLAZE
Charles Koelbel, Piyush Mehrotra, John Van Rosendale
ICPP1
1985 Prep-P: A Mapping Preprocessor for CHiP Architectures
Francine Berman, Michael A. Goodrich, Charles Koelbel, W. J. Robison III, Karen Showell
ICPP3