EDBT 2026 Demo / reviewers in the wild / expert
André Merzky
dblp:58/3981
· DBLP profile ↗
32ranked-venue papers
4as first author
7since 2021 · last 2024
0000-0002-7228-4327ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Workflow Mini-Apps: Portable, Scalable, Tunable & Faithful Representations of Scientific WorkflowsabstractWorkflows are critical for scientific discovery. However, the sophistication, heterogeneity, and scale of workflows make building, testing, and optimizing them increasingly challenging. Furthermore, their complexity and heterogeneity make performance reproducibility hard. In this paper, we propose workflow mini-apps as a tool to address the challenges in building and testing workflows while controlling the fidelity of representing real-world workflows. Workflow mini-apps are deployed and run on various HPC systems and architectures without workflow-specific constraints. We offer insight into their design and implementation, providing an analysis of their performance and reproducibility. Workflow mini-apps thus advance the science of workflows by providing simple, portable, and managed (fidelity) representations of otherwise complex and difficult-to-control real workflows. Ozgur O. Kilic, Tianle Wang 0001, Matteo Turilli, Mikhail Titov, André Merzky, Line C. Pouchard, Shantenu Jha |
CCGrid | 5 |
| 2024 | Radical-Cylon: A Heterogeneous Data Pipeline for Scientific Computing
Arup Kumar Sarker, Aymen Alsaadi, Niranda Perera, Mills Staylor, Gregor von Laszewski, Matteo Turilli, Ozgur O. Kilic, Mikhail Titov, André Merzky, Shantenu Jha, Geoffrey C. Fox |
JSSPP | 9 |
| 2023 | PSI/J: A Portable Interface for Submitting, Monitoring, and Managing JobsabstractIt is generally desirable for high-performance computing (HPC) applications to be portable between HPC systems, for example to make use of more performant hardware, make effective use of allocations, and to co-locate compute jobs with large datasets. Unfortunately, moving scientific applications between HPC systems is challenging for various reasons, most notably that HPC systems have different HPC schedulers. We introduce PSI/J, a job management abstraction API intended to simplify the construction of software components and applications that are portable over various HPC scheduler implementations. We argue that such a system is both necessary and that no viable alternative currently exists. We analyze similar notable APIs and attempt to determine the factors that influenced their evolution and adoption by the HPC community. We base the design of PSI/J on that analysis. We describe how PSI/J has been integrated in three workflow systems and one application, and also show via experiments that PSI/J imposes minimal overhead. Mihael Hategan, André Merzky, Nicholson T. Collier, Ketan Maheshwari, Jonathan Ozik, Matteo Turilli, Andreas Wilke, Justin M. Wozniak, Kyle Chard, Ian T. Foster, Rafael Ferreira da Silva, Shantenu Jha, Daniel E. Laney |
e-Science | 2 |
| 2022 | RAPTOR: Ravenous Throughput ComputingabstractWe describe the design, implementation and performance of the RADICAL-Pilot task overlay (RAPTOR). RAPTOR enables the execution of heterogeneous tasks-i.e., functions and executables with arbitrary duration-on HPC platforms, pro-viding high throughput and high resource utilization. RAPTOR supports the high throughput virtual screening requirements of DOE's National Virtual Biotechnology Laboratory effort to find therapeutic solutions for COVID-19. RAPTOR has been used on 8300 compute nodes to sustain 144M/hour docking hits, and to screen 1011 ligands. To the best of our knowledge, both the throughput rate and aggregated number of executed tasks are a factor of two greater than previously reported in literature. RAPTOR represents important progress towards improvement of computational drug discovery, in terms of size of libraries screened, and for the possibility of generating training data fast enough to serve the last generation of docking surrogate models. André Merzky, Matteo Turilli, Shantenu Jha |
CCGRID | 1 |
| 2022 | RADICAL-Pilot and PMIx/PRRTE: Executing Heterogeneous Workloads at Large Scale on Partitioned HPC Resources
Mikhail Titov, Matteo Turilli, André Merzky, Thomas J. Naughton, Wael R. Elwasif, Shantenu Jha |
JSSPP | 3 |
| 2022 | Design and Performance Characterization of RADICAL-Pilot on Leadership-Class PlatformsabstractMany extreme scale scientific applications have workloads comprised of a large number of individual high-performance tasks. The Pilot abstraction decouples workload specification, resource management, and task execution via job placeholders and late-binding. As such, suitable implementations of the Pilot abstraction can support the collective execution of large number of tasks on supercomputers. We introduce RADICAL-Pilot (RP) as a portable, modular and extensible pilot-enabled runtime system. We describe RP's design, architecture and implementation. We characterize its performance and show its ability to scalably execute workloads comprised of tens of thousands heterogeneous tasks on DOE and NSF leadership-class HPC platforms. Specifically, we investigate RP's weak/strong scaling with CPU/GPU, single/multi core, (non)MPI tasks and Python functions when using most of ORNL Summit and TACC Frontera. RADICAL-Pilot can be used stand-alone, as well as the runtime for third-party workflow systems. André Merzky, Matteo Turilli, Mikhail Titov, Aymen Alsaadi, Shantenu Jha |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | IMPECCABLE: Integrated Modeling PipelinE for COVID Cure by Assessing Better LEadsabstractThe drug discovery process currently employed in the pharmaceutical industry typically requires about 10 years and $2–3 billion to deliver one new drug. This is both too expensive and too slow, especially in emergencies like the COVID-19 pandemic. In silico methodologies need to be improved both to select better lead compounds, so as to improve the efficiency of later stages in the drug discovery protocol, and to identify those lead compounds more quickly. No known methodological approach can deliver this combination of higher quality and speed. Here, we describe an Integrated Modeling PipEline for COVID Cure by Assessing Better LEads (IMPECCABLE) that employs multiple methodological innovations to overcome this fundamental limitation. We also describe the computational framework that we have developed to support these innovations at scale, and characterize the performance of this framework in terms of throughput, peak performance, and scientific results. We show that individual workflow components deliver 100 × to 1000 × improvement over traditional methods, and that the integration of methods, supported by scalable infrastructure, speeds up drug discovery by orders of magnitudes. IMPECCABLE has screened ∼ 1011 ligands and has been used to discover a promising drug candidate. These capabilities have been used by the US DOE National Virtual Biotechnology Laboratory and the EU Centre of Excellence in Computational Biomedicine. Aymen Alsaadi, Dario Alfè, Yadu N. Babuji, Agastya Bhati, Ben Blaiszik, Alex Brace, Thomas S. Brettin, Kyle Chard, Ryan Chard, Austin Clyde, Peter V. Coveney, Ian T. Foster, Tom Gibbs, Shantenu Jha, Kristopher Keipert, Dieter Kranzlmüller, Thorsten Kurth, Hyungro Lee, Zhuozhao Li, Gerald Mathias, André Merzky, Alexander Partin, Arvind Ramanathan, Ashka Shah, Abraham C. Stern, Rick L. Stevens, Mikhail Titov, Anda Trifan, Aristeidis Tsaris, Matteo Turilli, Huub J. J. Van Dam, Shunzhou Wan, David Wifling, Junqi Yin |
ICPP | 22 |
| 2018 | Building Blocks for Workflow System MiddlewareabstractWe suggest there is a need for a fresh perspective on the design and development of middleware for high-performance workflows and workflow systems. We argue for a building blocks approach, outline a description of this approach and define their properties. We discuss RADICAL-Cybertools as one implementation of the building blocks concept, showing how they have been designed and developed in accordance with this approach. We discuss three case-studies where RADICAL-Cybertools have been used to develop new workflow systems capabilities and in-tegrated to enhance existing ones, illustrating the potential and promise of the building blocks approach. Matteo Turilli, André Merzky, Vivekanandan Balasubramanian, Shantenu Jha |
CCGrid | 2 |
| 2018 | Towards Exascale Computing for High Energy Physics: The ATLAS Experience at ORNLabstractTraditionally, the ATLAS experiment at Large Hadron Collider (LHC) has utilized distributed resources as provided by the Worldwide LHC Computing Grid (WLCG) to support data distribution, data analysis and simulations. For example, the ATLAS experiment uses a geographically distributed grid of approximately 200,000 cores continuously (250 000 cores at peak), (over 1,000 million core-hours per year) to process, simulate, and analyze its data (todays total data volume of ATLAS is more than 300 PB). After the early success in discovering a new particle consistent with the long-awaited Higgs boson, ATLAS is continuing the precision measurements necessary for further discoveries. Planned high-luminosity LHC upgrade and related ATLAS detector upgrades, that are necessary for physics searches beyond Standard Model, pose serious challenge for ATLAS computing. Data volumes are expected to increase at higher energy and luminosity, causing the storage and computing needs to grow at a much higher pace than the flat budget technology evolution (see Fig. 1). The need for simulation and analysis will overwhelm the expected capacity of WLCG computing facilities unless the range and precision of physics studies will be curtailed. V. Ananthraj, Kaushik De, Shantenu Jha, Alexei Klimentov, Danila Oleynik, Sarp Oral, André Merzky, Ruslan Mashinistov, Sergey Panitkin, P. Svirin, Matteo Turilli, Jack C. Wells, Sean R. Wilkinson |
eScience | 7 |
| 2018 | Using Pilot Systems to Execute Many Task Workloads on Supercomputers
André Merzky, Matteo Turilli, Manuel Maldonado, Mark Santcroos, Shantenu Jha |
JSSPP | 1 |
| 2017 | Evaluating Distributed Execution of WorkloadsabstractResource selection and task placement for distributed execution poses conceptual and implementation difficulties. Although resource selection and task placement are at the core of many tools and workflow systems, the methods are ad hoc rather than being based on models. Consequently, partial and non-interoperable implementations proliferate. We address both the conceptual and implementation difficulties by experimentally characterizing diverse modalities of resource selection and task placement. We compare the architectures and capabilities of two systems: the AIMES middleware and Swift workflow scripting language and runtime. We integrate these systems to enable the distributed execution of Swift workflows on Pilot-Jobs managed by the AIMES middleware. Our experiments characterize and compare alternative execution strategies by measuring the time to completion of heterogeneous uncoupled workloads executed at diverse scale and on multiple resources. We measure the adverse effects of pilot fragmentation and early binding of tasks to resources and the benefits of backfill scheduling across pilots on multiple resources. We then use this insight to execute a multi-stage workflow across five production-grade resources. We discuss the importance and implications for other tools and workflow systems Matteo Turilli, Yadu N. Babuji, André Merzky, Ming Tai Ha, Michael Wilde, Daniel S. Katz, Shantenu Jha |
eScience | 3 |
| 2016 | RepEx: A Flexible Framework for Scalable Replica Exchange Molecular Dynamics SimulationsabstractReplica Exchange (RE) simulations have emerged as an important algorithmic tool for the molecular sciences. Typically RE functionality is integrated into the molecular simulation software package. A primary motivation of the tight integration of RE functionality with simulation codes has been performance. This is limiting at multiple levels. First, advances in the RE methodology are tied to the molecular simulation code for which they were developed. Second, it is difficult to extend or experiment with novel RE algorithms, since expertise in the molecular simulation code is required. The tight integration results in difficulty to gracefully handle failures, and other runtime fragilities. We propose the RepEx framework which is addressing aforementioned shortcomings, while striking the balance between flexibility (any RE scheme) and scalability (several thousand replicas) over a diverse range of HPC platforms. The primary contributions of the RepEx framework are: (i) its ability to support different Replica Exchange schemes independent of molecular simulation codes, (ii) provide the ability to execute different exchange schemes and replica counts independent of the specific availability of resources, (iii) provide a runtime system that has first-class support for task-level parallelism, and (iv) provide a required scalability along multiple dimensions. Antons Treikalis, André Merzky, Tai-Sung Lee, Darrin M. York, Shantenu Jha |
ICPP | 2 |
| 2016 | Integrating Abstractions to Enhance the Execution of Distributed ApplicationsabstractOne of the factors that limits the scale, performance, and sophistication of distributed applications is the difficulty of concurrently executing them on multiple distributed computing resources. In part, this is due to a poor understanding of the general properties and performance of the coupling between applications and dynamic resources. This paper addresses this issue by integrating abstractions representing distributed applications, resources, and execution processes into a pilot-based middleware. The middleware provides a platform that can specify distributed applications, execute them on multiple resource and for different configurations, and is instrumented to support investigative analysis. We analyzed the execution of distributed applications using experiments that measure the benefits of using multiple resources, the late-binding of scheduling decisions, and the use of backfill scheduling. Matteo Turilli, Zhao Zhang 0007, André Merzky, Michael Wilde, Jon B. Weissman, Daniel S. Katz, Shantenu Jha |
IPDPS | 4 |
| 2016 | Application skeletons: Construction and use in eScience
Daniel S. Katz, André Merzky, Zhao Zhang 0007, Shantenu Jha |
Future Gener. Comput. Syst. | 2 |
| 2014 | Towards Standardized Job Submission and Control in Infrastructure Clouds
Peter Tröger, André Merzky |
J. Grid Comput. | 2 |
| 2012 | P∗: A model of pilot-abstractionsabstractPilot-Jobs support effective distributed resource utilization, and are arguably one of the most widely-used distributed computing abstractions - as measured by the number and types of applications that use them, as well as the number of production distributed cyberinfrastructures that support them. In spite of broad uptake, there does not exist a well-defined, unifying conceptual model of Pilot-Jobs which can be used to define, compare and contrast different implementations. Often Pilot-Job implementations are strongly coupled to the distributed cyber-infrastructure they were originally designed for. These factors present a barrier to extensibility and interoperability. This paper is an attempt to (i) provide a minimal but complete model (P*) of Pilot-Jobs, (ii) establish the generality of the P* Model by mapping various existing and well known Pilot-Job frameworks such as Condor and DIANE to P*, (iii) derive an interoperable and extensible API for the P* Model (Pilot-API), (iv) validate the implementation of the Pilot-API by concurrently using multiple distinct Pilot-Job frameworks on distinct production distributed cyberinfrastructures, and (v) apply the P* Model to Pilot-Data. André Luckow, Mark Santcroos, André Merzky, Ole Weidner, Pradeep Kumar Mantha, Shantenu Jha |
eScience | 3 |
| 2012 | Towards a common model for pilot-jobsabstractPilot-Jobs have become one of the most successful abstractions in distributed computing. In spite of extensive uptake, there does not exist a well defined, unifying conceptual model of pilot-jobs which can be used to define, compare and contrast different implementations. This presents a barrier to extensibility and interoperability. This paper is an attempt to, (i) provide a minimal but complete model (P*) of pilot-jobs, (ii) establish the generality of the P* Model by mapping various existing and well known pilot-jobs frameworks such as Condor and DIANE to P*, (iii) demonstrate the interoperable and concurrent usage of distinct pilot-job frameworks on different production distributed cyberinfrastructures via the use of an extensible API for the P* Model (Pilot-API). André Luckow, Mark Santcroos, Ole Weidner, André Merzky, Sharath Maddineni, Shantenu Jha |
HPDC | 4 |
| 2011 | Understanding application-level interoperability: Scaling-out MapReduce over high-performance grids and clouds
Saurabh Sehgal, Miklós Erdélyi, André Merzky, Shantenu Jha |
Future Gener. Comput. Syst. | 3 |
| 2010 | What Is the Price of Simplicity? - A Cross-Platform Evaluation of the SAGA API
Mathijs den Burger, Ceriel J. H. Jacobs, Thilo Kielmann, André Merzky, Ole Weidner, Hartmut Kaiser |
Euro-Par (1) | 4 |
| 2009 | Programming Abstractions for Data Intensive Computing on Clouds and GridsabstractMapReduce has emerged as an important data-parallel programming model for data-intensive computing - for Clouds and Grids. However most if not all implementations of MapReduce are coupled to a specific infrastructure. SAGA is a high-level programming interface which provides the ability to create distributed applications in an infrastructure independent way. In this paper, we show how MapReduce has been implemented using SAGA and demonstrate its interoperability across different distributed platforms - Grids, Cloud-like infrastructure and Clouds. We discuss the advantages of programmatically developing MapReduce using SAGA, by demonstrating that the SAGA-based implementation is infrastructure independent whilst still providing control over the deployment, distribution and runtime decomposition. The ability to control the distribution and placement of the computation units (workers) is critical in order to implement the ability to move computational work to the data. This is required to keep data network transfer low and in the case of commercial Clouds the monetary cost of computing the solution low. Using data-sets of size up to 10GB, and upto 10 workers, we provide detailed performance analysis of the SAGA-MapReduce implementation, and show how controllingthe distribution of computation and the payload per worker helps enhance performance. Chris Miceli, Michael Miceli, Shantenu Jha, Hartmut Kaiser, André Merzky |
CCGRID | 5 |
| 2009 | A Fresh Perspective on Developing and Executing DAG-Based Distributed Applications: A Case-Study of SAGA-Based MontageabstractMost workflow based applications currently have to adapt to available tools. While this keeps the cost of development low, it can lead to performance and flexibility tradeoffs that the application developer and deployer must make. In this paper, we use the Montage astronomical image mosaicking application as prototypical DAG-based workflow application to layout the development and deployment decisions for distributed applications. We discuss and explain the lack of simple (easy-to-use), scalable, and extensible distributed applications. We then introduce SAGA as a technology that permits the construction of abstractions that aid the development and execution of the applications, and thus addresses some of common shortcomings of traditional distributed applications development. We use Montage together with SAGA to examine how legacy applications can be made to run on distributed infrastructures, to see if our reasons are valid, and to compare potential new methods for creating distributed applications with existing technologies that are currently used. We demonstrate the ability to (i) scale-out and (ii) use different production infrastructure, while maintaining performance comparable to established systems. Our hope is that by demonstrating the simplicity of development along with other advantages (performance, scalability, extensibility, and infrastructure independence), this example will encourage others to think more broadly about how distributed applications are created and how new programming models such as Dryad can be supported in an infrastructure independent way, thus eventually leading to more applications that can seamlessly scale-out. André Merzky, Katerina Stamou, Shantenu Jha, Daniel S. Katz |
eScience | 1 |
| 2009 | Using clouds to provide grids with higher levels of abstraction and explicit support for usage modesabstractAbstract Grids in their current form of deployment and implementation have not been as successful as hoped in engendering distributed applications. Among other reasons, the level of detail that needs to be controlled for the successful development and deployment of applications remains too high. We argue that there is a need for higher levels of abstractions for current Grids. By introducing the relevant terminology, we try to understand Grids and Clouds as systems; we find this leads to a natural role for the concept of Affinity, and argue that this is a missing element in current Grids. Providing these affinities and higher‐level abstractions is consistent with the common concepts of Clouds. Thus this paper establishes how Clouds can be viewed as a logical and next higher‐level abstraction from Grids. Copyright © 2009 John Wiley & Sons, Ltd. Shantenu Jha, André Merzky, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 2 |
| 2008 | Distributed Replica-Exchange Simulations on Production Environments Using SAGA and MigolabstractThere exists a class of scientific applications for which utilizing distributed resources is critical for reducing the time-to-solution. In this paper, we discuss a specific class of applications - Replica-Exchange simulations - where the orchestration of many distributed jobs in a dynamic and inherently unreliable distributed environment is essential for a successful completion. We describe the design, development and deployment of a unique framework for constructing fault-tolerant distributed simulations. The framework consists of two primary components - SAGA and Migol. SAGA is a high-level programmatic abstraction layer that provides a standardised interface for the primary distributed functionality required for application development. We present details of a newly developed functionality in SAGA - the Checkpoint and Recovery (CPR) API. Migol is an adaptive middleware, which supports the fault-tolerance of distributed applications by providing the capability to recover applications from checkpoint files transparently. In addition to describing the integration of SAGA-CPR with the Migol infrastructure, we outline our experiences with running a large scale, general-purpose Replica-Exchange application in a production distributed environment. André Luckow, Shantenu Jha, Joohyun Kim 0001, André Merzky, Bettina Schnor |
eScience | 4 |
| 2007 | Grid Interoperability at the Application Level Using SAGAabstractSAGA is a high-level programming abstraction, which significantly facilitates the development and deployment of Grid-aware applications. The primary aim of this paper is to discuss how each of the three main components of the SAGA landscape - interface specification, specific implementation and the different adaptors for middleware distribution - facilitate application-level interoperability. We discuss SAGA in relation to the ongoing GIN Community Group efforts and show the consistency of the SAGA approach with the GIN Group efforts. We demonstrate how interoperability can be enabled by the use of SAGA, by discussing two simple, yet meaningful applications: in the first, SAGA enables applications to utilize interoperability and in the second example SAGA adaptors provide the basis for interoperability. Shantenu Jha, Hartmut Kaiser, André Merzky, Ole Weidner |
eScience | 3 |
| 2006 | Poster reception - The SAGA C++ reference implementation: a milestone toward new high-level grid applicationsabstractIn Grid computing multiple, incompatible middleware frameworks exist and are widely used in large research and production environments. Standard specifications for such Grid middleware are rarely available, and often still unstable. This hinders the ability of application programmers to write portable Grid application code.The Simple API for Grid Applications (SAGA) is an ongoing standardization effort within the Open Grid Forum (OGF). The SAGA API provides a simple, uniform interface for applications which utilize very dynamic and heterogeneous Grid environments. With SAGA, programmers of high-level Grid application do not have to learn about underlying Grid middleware layers and frameworks.Our newly developed SAGA C++ reference implementation makes this API available for real-world applications, providing a flexible, extensible, portable, and generic framework usable in dynamic environments.This poster describes the key features of SAGA, and provides examples of its use from the C++ reference implementation. Hartmut Kaiser, André Merzky, Stephan Hirmer, Gabrielle Allen, Edward Seidel |
SC | 2 |
| 2005 | The Grid Application Toolkit: Toward Generic and Easy Application Programming Interfaces for the GridabstractCore Grid technologies are rapidly maturing, but there remains a shortage of real Grid applications. One important reason is the lack of a simple and high-level application programming toolkit, bridging the gap between existing Grid middleware and application-level needs. The Grid Application Toolkit (GAT), as currently developed by the EC-funded project GridLab, provides this missing functionality. As seen from the application, the GAT provides a unified simple programming interface to the Grid infrastructure, tailored to the needs of Grid application programmers and users. A uniform programming interface will be needed for application developers to create a new generation of "Grid-aware" applications. The GAT implementation handles both the complexity and the variety of existing Grid middleware services via so-called adaptors. Complementing existing Grid middleware, GridLab also provides high-level services to implement the GAT functionality. We present the GridLab software architecture, consisting of the GAT, environment-specific adaptors, and GridLab services. We elaborate the concepts underlying the GAT and outline the corresponding application programming interface. We present the functionality of GridLab's high-level services and demonstrate how a dynamic Grid application can easily benefit from the GAT. All GridLab software is open source and can be downloaded from the project Web site. Gabrielle Allen, Kelly Davis, Tom Goodale, Andrei Hutanu, Hartmut Kaiser, Thilo Kielmann, André Merzky, Rob van Nieuwpoort, Alexander Reinefeld, Florian Schintke, Thorsten Schütt, Edward Seidel, Brygg Ullmer |
Proc. IEEE | 7 |
| 2004 | Remote partial file access using compact pattern descriptionsabstractWe present a method for the efficient access to parts of remote files. The efficiency is achieved by using a file format independent compact pattern description, that allows us to request several parts of a file in a single operation. This results in a drastically reduced number of remote operations and network latencies if compared to common solutions. We measured the time to access parts of remote files with compact patterns, compared it with normal GridFTP remote partial file access and observed a significant performance increase. Further we discuss how the presented pattern access can be used for an efficient read from multiple replicas and how this can be integrated into a data management system to support the storage of partial replicas for large scale simulations. Thorsten Schütt, André Merzky, Andrei Hutanu, Florian Schintke |
CCGRID | 2 |
| 2002 | Meta- and Grid-Computing
Michel Cosnard, André Merzky |
Euro-Par | 2 |
| 2002 | GridLab--a grid application toolkit and testbed
Edward Seidel, Gabrielle Allen, André Merzky, Jarek Nabrzyski |
Future Gener. Comput. Syst. | 3 |
| 2001 | Early Experiences with the EGrid TestbedabstractThe Testbed and Applications working group of the European Grid Forum (EGrid) is actively building and experimenting with a grid infrastructure connecting several research-based supercomputing sites located in Europe. The paper reports on our first feasibility study: running a self-migrating version of the Cactus simulation code across the European grid testbed, including "live" remote data visualization and steering from different demonstration booths at Supercomputing 2000, in Dallas, TX. We report on the problems that had to be resolved for this endeavour and identify open research challenges for building production-grade grid environments. Gabrielle Allen, Thomas Dramlitsch, Tom Goodale, Gerd Lanfermann, Thomas Radke, Edward Seidel, Thilo Kielmann, Kees Verstoep, Zoltán Balaton, Péter Kacsuk, Ferenc Szalai, Jörn Gehring, Axel Keller, Achim Streit, Ludek Matyska, Miroslav Ruda, Ales Krenek, Harald Knipp, André Merzky, Alexander Reinefeld, Florian Schintke, Bogdan Ludwiczak, Jarek Nabrzyski, Juliusz Pukacki, Hans-Peter Kersken, Giovanni Aloisio, Massimo Cafaro, Wolfgang Ziegler, Michael Russell |
CCGRID | 19 |
| 2001 | Cactus Grid Computing: Review of Current Development
Gabrielle Allen, Werner Benger, Thomas Dramlitsch, Tom Goodale, Hans-Christian Hege, Gerd Lanfermann, André Merzky, Thomas Radke, Edward Seidel |
Euro-Par | 7 |
| 2000 | The Cactus Code: A Problem Solving Environment for the GridabstractCactus is an open source problem solving environment designed for scientists and engineers. Its modular structure facilitates parallel computation across different architectures and collaborative code development between different groups. The Cactus Code originated in the academic research community, where it has been developed and used over many years by a large international collaboration of physicists and computational scientists. We discuss how the intensive computing requirements of physics applications now using the Cactus Code encourage the use of distributed and metacomputing, describe the development and experiments which have already been performed with Cactus, and detail how its design makes it an ideal application test-bed for Grid computing. Gabrielle Allen, Werner Benger, Tom Goodale, Hans-Christian Hege, Gerd Lanfermann, André Merzky, Thomas Radke, Edward Seidel, John Shalf |
HPDC | 6 |