David Abramson 0001

dblp:09/6203 · also David A. Abramson · DBLP profile ↗
← Back
104ranked-venue papers
29as first author
7since 2021 · last 2025
0000-0003-0441-4596ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 57 · 17 first-author · 1 since 2021Software engineering, systems software and programming languages · 36 · 10 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 33 · 9 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorArtificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
YearPublicationVenuePosition
2025 Institutional Research Computing Capabilities in Australia: 2024
abstract
Institutional research computing infrastructure plays a vital role in Australia’s research ecosystem, complementing and extending national-level facilities. This paper presents an analysis of research computing capabilities across Australian universities and research organisations, examining how institutional infrastructure supports research excellence through localised compute resources, specialised hardware, and cluster solutions. Our study reveals that institutional computing resources of nearly 112,258 CPU cores and 2,241 GPUs serve as essential bridges between desktop computing and national facilities for over 6,000 researchers, enabling research workflows that span from development to large-scale computations. We estimate the total replacement value of this infrastructure to be approximately $144M AUD. Based on detailed infrastructure data provided by research computing facilities across multiple institutions, we identify key patterns in infrastructure deployment, utilisation metrics, and strategic alignment with research priorities. Our findings demonstrate that institutional computing resources not only provide critical support for data-intensive research but also facilitate training and higher-degree research student projects, enable prototyping and development, and ensure data sovereignty compliance when necessary. The analysis shows how these facilities leverage national infrastructure investments while addressing institution-specific needs that cannot be met by national facilities alone. We present evidence that strategic investment in institutional research computing capabilities yields significant returns through increased research productivity, enhanced graduate training, and improved research outcomes. This study provides valuable insights for research organisations planning their computing infrastructure strategies and highlights the importance of maintaining robust institutional computing capabilities alongside national facilities.
Slava Kitaeff, Luc Betbeder-Matibet, Jake Carroll, Stephen Giugni, David Abramson 0001, John Zaitseff, Sarah Walters, David Powell, Chris Bording, Angus Macoustra, Fabien Voisin, Jarrod Hurley
eScience5
2024 An Analysis of Research Data Storage Systems
abstract
The scientific protocols, experiments and instruments that generate data are an integral part of the research lifecycle. Consequently, almost every scientific research institution requires a Research Data Storage System (RDSS). However, RDSS implementations vary significantly due to factors that include cost, geography, workloads, policy, risk tolerance and available technical skills. A RDSS may be on premises, in the public cloud or a mixture of both. Previously we identified 10 key high level features of a RDSS and defined an abstract high-level Research Data Reference Architecture (RDRA). Together, these features enable data fabrics, creating repeatable and consistent structures for low friction, highly efficient data movement and near real time data analysis for decision making and scientific workflows. We build on this earlier work in this paper to present a new Research Data Implementation Architecture (RDIA) that meets the RDRA and can guide implementations without locking in any specific product or service. This paper presents and compares six significant RDSSs and shows how they meet both the RDIA and therefore the RDRA. We identify a new structure – the Research Data Storage System Aggregator (RDSS-A) that describes clusters of RDSSs, and survey five such instances. Finally, we provide a real-world validation of the RDIA by documenting the technical specification of a real RDSS. The new work both clarifies a complex landscape, and aids groups building new systems or adopting existing systems.
Jake Carroll, David Abramson 0001, Bronis R. de Supinski
e-Science2
2023 Why We Need a Reference Architecture for Research Data
abstract
There is little doubt that we have entered an era where data underpins modern science and research in general. In support of this, numerous infrastructures have been designed and built, ranging from proprietary on-prem systems through to distributed commercial clouds. Such implementations provide a range of functions during the research lifecycle from provisioning and cataloguing data assets through to storing and presenting data to computing platforms. In this paper we analyse the underlying principles of such systems and develop a high-level Research Data Reference Architecture (RDRA). Specifically, we identify eight key features of a RDRA that can guide the design, construction, and procurement of implementations without mandating any domain, approach, technical solution, or product choice. As a result, it allows implementers to make local and commercial decisions while still meeting the core requirements of a research data management platform. The intended audience is teams charged with implementing infrastructure in research organizations.
David Abramson 0001, Luc Betbeder-Matibet, Stephen Bird, Jake Carroll, Rhys S. Francis, Wojtek Goscinski, Ai-Lin Soo, Garry Swan, Carmel Walsh, Glenn R. Wightwick, J. Max Wilkinson
e-Science1
2023 Moving small files in a networked environment
Chao Jin 0001, David Abramson 0001, Jake Carroll, Zhengchun Liu, Rajkumar Kettimuthu
Future Gener. Comput. Syst.2
2022 Democratising large scale instrument-based science through e-Infrastructure
abstract
Modern scientific instruments are becoming essential for discoveries because they provide unprecedented insight into physical or biological events – often in real time. However, these instruments may generate large amounts of data, and increasingly they require sophisticated e-infrastructure for analysis, storage and archive. The increasing complexity and scale of the data, processing steps and systems has made it difficult for domain scientists to perform their research, narrowing the user base to a select few. In this paper, we present a framework that democratises large-scale instrument-based science, increasing the number of researchers who can engage. We discuss a prototype at the University of Queensland. The system is illustrated through two case studies, one involving light microscopy imaging of the innate immune system, and the other electron microscopy imaging of the SARS-CoV-2 viral proteins.
David Abramson 0001, Deborah S. Barkauskas, Jake Carroll, Nicholas D. Condon, Naphak Modhiran, James Springfield, Daniel Watterson, Chao Jin 0001
e-Science1
2022 On the reliability and the limits of inference of amino acid sequence alignments
abstract
MOTIVATION: Alignments are correspondences between sequences. How reliable are alignments of amino acid sequences of proteins, and what inferences about protein relationships can be drawn? Using techniques not previously applied to these questions, by weighting every possible sequence alignment by its posterior probability we derive a formal mathematical expectation, and develop an efficient algorithm for computation of the distance between alternative alignments allowing quantitative comparisons of sequence-based alignments with corresponding reference structure alignments. RESULTS: By analyzing the sequences and structures of 1 million protein domain pairs, we report the variation of the expected distance between sequence-based and structure-based alignments, as a function of (Markov time of) sequence divergence. Our results clearly demarcate the 'daylight', 'twilight' and 'midnight' zones for interpreting residue-residue correspondences from sequence information alone. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sandun Rajapaksa, Dinithi Sumanaweera, Arthur M. Lesk, Lloyd Allison, Peter J. Stuckey, Maria Garcia de la Banda, David Abramson 0001, Arun Siddharth Konagurthu
Bioinform.7
2021 On identifying statistical redundancy at the level of amino acid subsequences
abstract
This paper presents a framework to characterize and identify local sequences of proteins that are statistically redundant under the measure of Shannon information content while accounting for variations in their occurrences over evolutionary insertions, deletions, and substitutions of amino acids. The identification of such local sequences provides insights for downstream studies on proteins. Here, we have applied our methods to amino acid sequence data sets derived from a database corresponding to 935,552 substructural regions of varying sizes, covering 113,724 proteins from the protein data bank. The results identify, among others, a surjective mapping between 110,598 local sequences (with an average length of 82 amino acids per sequence) and 1,493 topological shapes. The C++ source code and supporting material are available from https://lcb.infotech.monash.edu.au/bibm2021.
Sandun Rajapaksa, Dinithi Sumanaweera, Maria Garcia de la Banda, Peter J. Stuckey, David Abramson 0001, Lloyd Allison, Arthur M. Lesk, Arun Siddharth Konagurthu
BIBM5
2020 Tracking scientific simulation using online time-series modelling
abstract
The increase in compute power and complexity of supercomputing systems requires the decrease in the feature size and the supply voltage of internal components. Such development makes unintended errors such as soft errors, potentially caused by random bit flips, inevitable because of the huge size of the resources (such as CPU cores and memory). In this paper, we discuss a non-parametric statistical modelling technique to implement a soft error detector. By exploring temporal autocorrelation within key variables of a running scientific simulation, we introduce an automatic anomaly detection technique in which runtime data from a time-step based simulation can be converted into a time series, and a time series modelling technique can be used to identify soft errors at runtime. Experiments with LAMMPS, a high-performance molecular dynamics simulator, and with PLUTO, an open-source astrophysical code, reveal that the time-series based detector is subjected to less than 3% of both false-positive rate and false-negative rate while incurring only 6% performance overheads.
Minh Ngoc Dinh, Chien Trung Vo, David Abramson 0001
CCGRID3
2018 Energy efficiency modeling of parallel applications
Mark Endrei, Chao Jin 0001, Minh Ngoc Dinh, David Abramson 0001, Heidi Poxon, Luiz DeRose, Bronis R. de Supinski
SC4
2017 Statistical Compression of Protein Folding Patterns for Inference of Recurrent Substructural Themes
abstract
Computational analyses of the growing corpus of three-dimensional (3D) structures of proteins have revealed a limited set of recurrent substructural themes, termed super-secondary structures. Knowledge of super-secondary structures is important for the study of protein evolution and for the modeling of proteins with unknown structures. Characterizing a comprehensive dictionary of these super-secondary structures has been an unanswered computational challenge in protein structural studies. This paper presents an unsupervised method for learning such a comprehensive dictionary using the statistical framework of lossless compression on a database comprised of concise geometric representations of protein 3D folding patterns. The best dictionary is defined as the one that yields the most compression of the database. Here we describe the inference methodology and the statistical models used to estimate the encoding lengths. An interactive website for this dictionary is available at http://lcb.infotech.monash.edu.au/proteinConcepts/scop100/dictionary.html.
Ramanan Subramanian, Lloyd Allison, Peter J. Stuckey, Maria Garcia de la Banda, David Abramson 0001, Arthur M. Lesk, Arun Siddharth Konagurthu
DCC5
2017 A Metropolitan Area Infrastructure for Data Intensive Science
abstract
The increasing amount of data being collected from simulations, instruments and sensors creates challenges for existing e-Science infrastructure. In particular, it requires new ways of storing, distributing and processing data in order to cope with both the volume and velocity of the data. The University of Queensland has recently designed and deployed MeDiCI, a data fabric that spans the metropolitan area and provides seamless access to data regardless of where it is created, manipulated and archived. MeDiCI is novel in that it exploits temporal and spatial locality to move data on demand in an automated manner. This means that data only needs to reside locally in high speed storage whilst being manipulated, and it can be archived transparently in high capacity, but slower, technologies at other times. MeDiCI is built on commercially available technologies. In this paper, we describe these innovations and present some early results.
David Abramson 0001, Jake Carroll, Chao Jin 0001, Michael Mallon
eScience1
2017 A Computational Pipeline for the IUCN Risk Assessment for Meso-American Reef Ecosystem
abstract
Coral reefs are of global economic and biological significance but are subject to increasing threats. As a result, it is essential to understand the risk of coral reef ecosystem collapse and to develop assessment process for those ecosystems. The International Union for Conservation of Nature (IUCN) Red List of Ecosystem (RLE) is a framework to assess the vulnerability of an ecosystem. Importantly, the assessment processes need to be repeatable as new monitoring data arises. The repeatability will also enhance transparency. In this paper, we discuss the evolution of a computational pipeline for risk assessment of the Meso-American reef ecosystem, a diverse reef ecosystem located in the Caribbean, with the focus on improving the execution time starting from sequential and parallel implementation and finally using Apache Spark. The final form of the pipeline is a scientific workflow to improve its repeatability and reproducibility.
Hoang Anh Nguyen, Lucie Bland, Tristan Roberts, Siddeswara Guru, Minh Ngoc Dinh, David Abramson 0001
eScience6
2017 Caches All the Way Down: Infrastructure for Data Intensive Science
abstract
The rise of big data science has created new demands for modern computer systems. While floating performance has driven computer architecture and system design for the past few decades, there is renewed interest in the speed at which data can be ingested and processed. Early exemplars such as Gordon, the NSF funded system at the San Diego Supercomputing Centre, shifted the focus from pure floating-point performance to memory and IO rates. At the University of Queensland we have continued this trend with the design of FlashLite, a parallel cluster equipped with large amounts of main memory, flash disk, and a distributed shared memory system (ScaleMP's vSMP). This allows applications to place data "close" to the processor, enhancing processing speeds. Further, we have built a geographically distributed multi-tier hierarchical data fabric called MeDiCI, which provides an abstraction of very large data stores across the metropolitan area. MeDiCI leverages industry solutions such as IBM's Spectrum Scale and SGI's DMF platforms. Caching underpins both FlashLite and MeDiCI. In this I will describe the design decisions and illustrate some early application studies that benefit from the approach. I will also highlight some of the challenges that need to be solved for this approach to become mainstream.
David Abramson 0001
HPDC1
2016 FTS 2016 Workshop Keynote Speech
abstract
Provides an abstract of the keynote presentation and a brief professional biography of the presenter. The complete presentation was not made available for publication as part of the conference proceedings.
David Abramson 0001
CLUSTER1
2015 FlashLite: A High Performance Machine for Data Intensive Science
abstract
Data is predicted to transform the 21st century, fuelled by an exponential growth in the amount of data captured, generated and archived. Traditional high performance machines are optimized for numerical computing rather than IO performance or for supporting large memory applications. This paper discusses a new machine, called FlashLite, which addresses these challenges. The paper describes the motivation for the design, and discusses some driving application themes.
David Abramson 0001
ICPADS1
2015 Relative debugging for a highly parallel hybrid computer system
abstract
Relative debugging traces software errors by comparing two executions of a program concurrently - one code being a reference version and the other faulty. Relative debugging is particularly effective when code is migrated from one platform to another, and this is of significant interest for hybrid computer architectures containing CPUs accelerators or coprocessors. In this paper we extend relative debugging to support porting stencil computation on a hybrid computer. We describe a generic data model that allows programmers to examine the global state across different types of applications, including MPI/OpenMP, MPI/OpenACC, and UPC programs. We present case studies using a hybrid version of the `stellarator' particle simulation DELTA5D, on Titan at ORNL, and the UPC version of Shallow Water Equations on Crystal, an internal supercomputer of Cray. These case studies used up to 5,120 GPUs and 32,768 CPU cores to illustrate that the debugger is effective and practical.
Luiz De Rose, Andrew Gontarek, Aaron Vose, Bob Moench, David Abramson 0001, Minh Ngoc Dinh, Chao Jin 0001
SC5
2015 WorkWays: interacting with scientific workflows
abstract
Summary WorkWays is a science gateway that supports human‐in‐the‐loop scientific workflows. Human–workflow interactions are enabled by a dynamic Input Output (IO) model, which allows users to insert data into, or export data out of, a continuously running workflow. WorkWays has been used to solve a number of scientific problems where the user wishes to examine intermediate results in order to interact with the computation as the workflow progresses. This interactive capability not only provides better insights into the computation but also allows users to focus on different input parameter combinations. We have implemented a variety of data types and modes of interaction to account for a wide range of use cases and application domains. This paper demonstrates the applicability of WorkWays on three use cases from different domains. Copyright © 2015 John Wiley & Sons, Ltd.
Hoang Anh Nguyen, David Abramson 0001, Timoleon Kipouros, Andrew L. Janke, Graham J. Galloway
Concurr. Comput. Pract. Exp.2
2015 A data-centric framework for debugging highly parallel applications
abstract
Summary Contemporary parallel debuggers allow users to control more than one processing thread while supporting the same examination and visualisation operations of that of sequential debuggers. This approach restricts the use of parallel debuggers when it comes to large scale scientific applications run across hundreds of thousands compute cores. First, manually observing the runtime data to detect error becomes impractical because the data is too big. Second, performing expensive but useful debugging operations becomes infeasible as the computational codes become more complex, involving larger data structures, and as the machines become larger. This study explores the idea of a data‐centric debugging approach, which could be used to make parallel debuggers more powerful. It discusses the use ofad hocdebug‐time assertions that allow a user to reason about the state of a parallel computation. These assertions support the verification and validation of program state at runtime as a whole rather than focusing on that of only a single process state. Furthermore, the debugger's performance can be improved by exploiting the underlying parallel platform because the available compute cores can execute parallel debugging functions, while a program is idling at a breakpoint. We demonstrate the system with several case studies and evaluate the performance of the tool on a 20 000 cores Cray XE6. Copyright © 2013 John Wiley & Sons, Ltd.
Minh Ngoc Dinh, David Abramson 0001, Chao Jin 0001, Andrew Gontarek, Bob Moench, Luiz De Rose
Softw. Pract. Exp.2
2014 Scalable Relative Debugging
abstract
Detecting and isolating bugs that arise only at high processor counts is a challenging task. Over a number of years, we have implemented a special debugging method, called "relative debugging," that supports debugging applications as they evolve or are ported to larger machines. It allows a user to compare the state of a suspect program against another reference version even as the number of processors is increased. The innovative idea is the comparison of runtime data to reason about the state of the suspect program. While powerful, a naïve implementation of the comparison phase does not scale to large problems running on large machines. In this paper, we propose two different solutions including a hash-based scheme and a direct point-to-point scheme. We demonstrate the implementation, a case study, as well as the performance, of our techniques on 20K cores of a Cray XE6 system.
Minh Ngoc Dinh, David Abramson 0001, Chao Jin 0001
IEEE Trans. Parallel Distributed Syst.2
2013 Topic 11: Multicore and Manycore Programming - (Introduction)
Luiz De Rose, Jan Eitzinger, William Jalby, Alba Cristina Magalhaes Alves de Melo, David Abramson 0001, Alastair F. Donaldson, Tomàs Margalef
Euro-Par5
2013 Statistical Inference of Protein "LEGO Bricks"
abstract
Proteins are biomolecules of life. They fold into a great variety of three-dimensional (3D) shapes. Underlying these folding patterns are many recurrent structural fragments or building blocks (analogous to 'LEGO® bricks'). This paper reports an innovative statistical inference approach to discover a comprehensive dictionary of protein structural building blocks from a large corpus of experimentally determined protein structures. Our approach is built on the Bayesian and information theoretic criterion of minimum message length. To the best of our knowledge, this work is the first systematic and rigorous treatment of a very important data mining problem that arises in the cross-disciplinary area of structural bioinformatics. The quality of the dictionary we find is demonstrated by its explanatory power - any protein within the corpus of known 3D structures can be dissected into successive regions assigned to fragments from this dictionary. This induces a novel one-dimensional representation of three-dimensional protein folding patterns, suitable for application of the rich repertoire of character-string processing algorithms, for rapid identification of folding patterns of newly determined structures. This paper presents the details of the methodology used to infer the dictionary of building blocks, and is supported by illustrative examples to demonstrate its effectiveness and utility.
Arun Siddharth Konagurthu, Lloyd Allison, David Abramson 0001, Peter J. Stuckey, Arthur M. Lesk
ICDM3
2013 Recent advances in e-Science
Daniel S. Katz, David Abramson 0001
Future Gener. Comput. Syst.2
2013 Building an ecoinformatics platform to support climate change adaptation in Victoria
Christopher James Pettit 0001, Steve Williams 0001, Ian D. Bishop, Jean-Philippe Aurambout, A. B. M. Russel, Anthony Michael, Subhash Sharma, David Hunter, Pang Choung Chan, Colin Enticott, Ann Borda, David Abramson 0001
Future Gener. Comput. Syst.12
2013 A local sensitivity analysis method for developing biological models with identifiable parameters: Application to cardiac ionic channel modelling
Anna Sher, Ken Wang, Andrew J. Wathen, Philip Maybank, Gary R. Mirams, David Abramson 0001, Denis Noble, David Gavaghan
Future Gener. Comput. Syst.6
2013 Scheduling parameter sweep workflow in the Grid based on resource competition
Sucha Smanchat, Maria Indrawan, Sea Ling, Colin Enticott, David Abramson 0001
Future Gener. Comput. Syst.5
2013 Approaches to Distributed Execution of Scientific Workflows in Kepler
abstract
The Kepler scientific workflow system enables creation, execution and sharing of workflows across a broad range of scientific and engineering disciplines while also facilitating remote and distributed execution of workflows. In this paper, we present
Marcin Plóciennik, Tomasz Zok, Ilkay Altintas, Jianwu Wang 0001, Daniel Crawl, David Abramson 0001, Frederic Imbeaux, Bernard Guillerminet, Marcos López-Caniego, Isabel Campos Plasencia, Wojciech Pych, Pawel Ciecielag, Bartek Palak, Michal Owsiak, Yann Frauel
Fundam. Informaticae6
2012 A Scalable Parallel Debugging Library with Pluggable Communication Protocols
abstract
Parallel debugging faces challenges in both scalability and efficiency. A number of advanced methods have been invented to improve the efficiency of parallel debugging. As the scale of system increases, these methods highly rely on a scalable communication protocol in order to be utilized in large-scale distributed environments. This paper describes a debugging middleware that provides fundamental debugging functions supporting multiple communication protocols. Its pluggable architecture allows users to select proper communication protocols as plug-ins for debugging on different platforms. It aims to be utilized by various advanced debugging technologies across different computing platforms. The performance of this debugging middleware is examined on a Cray XE Supercomputer with 21,760 CPU cores.
Chao Jin 0001, David Abramson 0001, Minh Ngoc Dinh, Andrew Gontarek, Bob Moench, Luiz De Rose
CCGRID2
2012 Integration of modern data management practice with scientific workflows
abstract
Modern science increasingly involves managing and processing large amounts of distributed data accessed by global teams of researchers. To do this, we need systems that combine data, meta-data and workflows into a single system. This paper discusses such a system, built from a number of existing technologies. We demonstrate the effectiveness on a case study that analyses MRI data.
Neil E. B. Killeen, Jason M. Lohrey, Michael J. Farrell, Wilson Liu, Slavisa Garic, David Abramson 0001, Gary F. Egan
eScience6
2012 WorkWays: Interactive workflow-based science gateways
abstract
Workflow-based science gateways that bring the power of scientific workflows to the Web are becoming increasingly popular. Different IO models enabling interactions between a running workflow and web portal have been explored. However, these are typically not dynamic enough to allow users to insert data into, or export data out of, a continuously running workflow. In this paper, we present a novel IO model, which supports dynamic interaction between a workflow and its portal. We discuss a use case in which web portal are used to control the execution of scientific workflows. This IO model will be part of our workflow-based science gateway named WorkWays.
David Abramson 0001
eScience2
2012 Scalable parallel debugging with statistical assertions
abstract
Traditional debuggers are of limited value for modern scientific codes that manipulate large complex data structures. This paper discusses a novel debug-time assertion, called a "Statistical Assertion", that allows a user to reason about large data structures, and the primitives are parallelised to provide an efficient solution. We present the design and implementation of statistical assertions, and illustrate the debugging technique with a molecular dynamics simulation. We evaluate the performance of the tool on a 12,000 cores Cray XE6.
Minh Ngoc Dinh, David Abramson 0001, Chao Jin 0001, Andrew Gontarek, Bob Moench, Luiz De Rose
PPoPP2
2011 Assertion Based Parallel Debugging
abstract
Programming languages have advanced tremendously over the years, but program debuggers have hardly changed. Sequential debuggers do little more than allow a user to control the flow of a program and examine its state. Parallel ones support the same operations on multiple processes, which are adequate with a small number of processors, but become unwieldy and ineffective on very large machines. Typical scientific codes have enormous multi-dimensional data structures and it is impractical to expect a user to view the data using traditional display techniques. In this paper we discuss the use of debug-time assertions, and show that these can be used to debug parallel programs. The techniques reduce the debugging complexity because they reason about the state of large arrays without requiring the user to know the expected value of every element. Assertions can be expensive to evaluate, but their performance can be improved by running them in parallel. We demonstrate the system with a case study finding errors in a parallel version of the Shallow Water Equations, and evaluate the performance of the tool on a 4,096 cores Cray XE6.
Minh Ngoc Dinh, David Abramson 0001, Donny Kurniawan, Chao Jin 0001, Bob Moench, Luiz De Rose
CCGRID2
2011 Cyberinfrastructure Intership and its Application to e-Science
abstract
Universities are constantly searching for ways that prepare students as effective global professionals. At the same time, cyber infrastructure leverages computing, information, and communication technology to perform research, often in an international context. In this paper we discuss a novel model, called the Cyber infrastructure Internship Program (CIP), which serves both of these goals. Specifically, students apply or develop cyber infrastructure to solve challenging research problems, but they do this via summer internships abroad. CIP has been implemented at three different Universities: The University of California San Diego in the US, Osaka University in Japan and and Monash University in Australia. We discuss details of the schemes and provide some initial evaluations of their success.
David Abramson 0001, Peter W. Arzberger, Gabriele Wienhausen, Jim Galvin, Susumu Date, Fang-Pang Lin, Kai Nan, Shinji Shimojo
eScience1
2011 Keynote: Assertion Based Parallel Debugging
David Abramson 0001
ICA3PP (1)1
2011 ISENGARD: an infrastructure for supporting e-science and grid application development
abstract
Abstract Grid computing facilitates the aggregation and coordination of resources that are distributed across multiple administrative domains for large‐scale and complex e‐Science experiments. Writing, deploying, and testing grid applications over highly heterogeneous and distributed resources are complex and challenging. The process requires grid‐enabled programming tools that can handle the complexity and scale of the infrastructure. However, while a large amount of research has been undertaken into grid middleware, little work has been directed specifically at the area of grid application development tools. This paper presents the design and implementation of ISENGARD, an infrastructure for supporting e‐Science and grid application development. ISENGARD provides services, tools, and APIs that simplify grid software development. Copyright © 2010 John Wiley & Sons, Ltd.
Donny Kurniawan, David Abramson 0001
Concurr. Comput. Pract. Exp.2
2011 Special Theme: Project Management in E-Science: Challenges and Opportunities
Dimitrina Spencer, Ann Zimmerman, David Abramson 0001
Comput. Support. Cooperative Work.3
2011 Parameter Exploration in Science and Engineering Using Many-Task Computing
abstract
Robust scientific methods require the exploration of the parameter space of a system (some of which can be run in parallel on distributed resources), and may involve complete state space exploration, experimental design, or numerical optimization techniques. Many-Task Computing (MTC) provides a framework for performing robust design, because it supports the execution of a large number of otherwise independent processes. Further, scientific workflow engines facilitate the specification and execution of complex software pipelines, such as those found in real science and engineering design problems. However, most existing workflow engines do not support a wide range of experimentation techniques, nor do they support a large number of independent tasks. In this paper, we discuss Nimrod/K—a set of add in components and a new run time machine for a general workflow engine, Kepler. Nimrod/K provides an execution architecture based on the tagged dataflow concepts, developed in 1980s for highly parallel machines. This is embodied in a new Kepler "Director” that supports many-task computing by orchestrating execution of tasks on on clusters, Grids, and Clouds. Further, Nimrod/K provides a set of "Actors” that facilitate the various modes of parameter exploration discussed above. We demonstrate the power of Nimrod/K to solve real problems in cardiac science.
David Abramson 0001, Blair Bethwaite, Colin Enticott, Slavisa Garic, Thomas C. Peachey
IEEE Trans. Parallel Distributed Syst.1
2010 Electrochemical Parameter Optimization Using Scientific Workflows
abstract
Modern chemistry involves a mix of experimentation and computer-supported theory. Historically, these skills have been provided by different groups, and range from traditional “wet” laboratory science to advanced numerical simulation. This paper discusses the application of advanced e-Science software tools to electrochemistry research and involves collaboration between laboratories at Monash and Oxford Universities. In particular, we show how the Nimrod/OK tool can be used to automate the estimation of electrochemical parameters in Fourier transformed voltammetry. Replacing an ad-hoc manual process with e-Science tools both accelerates the process and produces more accurate solutions. In this work much attention was given to the role of the scientist user. During the research, Nimrod/K (on which Nimrod/OK is built) was extended to shield that user from technical details of the grid infrastructure, a new system of file management and grid abstraction has been incorporated.
Colin Enticott, Thomas C. Peachey, David Abramson 0001, Elena Mashkina, Chong-Yong Lee, Alan Bond, Gareth Kennedy, David Gavaghan, Darrell Elton
eScience3
2010 NG-TEPHRA: A Massively Parallel, Nimrod/G-enabled Volcanic Simulation in the Grid and the Cloud
abstract
Volcanoes are a principal factor of hazard across the Pacific Rim, with their focus of interest mostly divided into pyroclastic flows and ash deposition. The latter has significantly more impact due to its widespread geographical reach and prolonged effects in human activities and health. TEPHRA is a volcanic ash dispersion model based on a simple version of the advection-diffusion Suzuki model, which has been revisited and modified for the Iraz\'u volcano in Costa Rica. A full parameter exploration is necessary in this particular case (albeit not sufficient) due to scarce observational data. We present in this paper the model, its assumptions and limitations as well as application lifecycle with resulting ash distribution graphics. The computational experimental settings are described, in particular the use of Nimrod/G with respect to non-homogeneous parameter sweeps and its impact on execution time. We also analyze the implementation of a new parameter discard mechanism common to e-Science experiments where sequential generation of new parameter sets has to be complemented with an early verification in order to avoid allocation of CPU time to non-valid scenarios. Finally four sample 100K-scenario runs are analyzed for both traditional HPC clustering and Cloud computing resources in the Amazon EC2 Cloud.
Santiago Nunez, Blair Bethwaite, José Brenes, Gustavo Barrantes, José Castro, Eduardo Malavassi, David Abramson 0001
eScience7
2010 Realising an eScience Platform to Support Climate Change Adaptation in Victoria
abstract
Our research is focused on developing an ecoinformatics platform to support climate change adaptation in Victoria. A multidisciplinary, cross-organisational approach is taken in developing adaptation strategies to deal with the `diabolical' policy problem of climate change. The platform comprises a number of components including: (i) a metadata discovery tool to support modelling, (ii) a workflow engine for connecting climate change models, (iii) geographical visualisation tools for communicating landscape and farm impacts, (iv) a landscape object library for storing and sharing digital models, and (v) a landscape constructor tool to support participatory decision-making, and (vi) a virtual organisation for collaboration and sharing information. In this paper we will discuss the platform as it has been developed to support collaborative research and to inform stakeholders of the likely impacts of climate change in South West Victoria, Australia. We will discuss some of the drivers for research in developing the ecoinformatics platform and its components. The paper concludes by identifying some future research directions in better connecting researchers and communicating science outcomes associated with climate change impact and adaptation.
Christopher James Pettit 0001, A. B. M. Russel, Anthony Michael, Jean-Philippe Aurambout, Subhash Sharma, Steve Williams 0001, David Hunter, Pang Choung Chan, Ann Borda, Ian D. Bishop, David Abramson 0001
eScience11
2010 A Local Sensitivity Analysis Method for Developing Biological Models with Identifiable Parameters: Application to L-type Calcium Channel Modelling
abstract
Computational cardiac models provide important insights into the underlying mechanisms of heart function. Parameter estimation in these models is an ongoing challenge with many existing models being overparameterised. Sensitivity analysis presents a key tool for exploring the parameter identiflability. While existing methods provide insight into the significance of the parameters, they are unable to identify redundant parameters in an efficient manner. We present a new singular value decomposition based algorithm for determining parameter identifiability in cardiac models. Using this local sensitivity approach, we investigate the Mahajan 2008 rabbit ventricular myocyte L-type calcium current model. We identify non-significant and redundant parameters and improve the leal model by reducing it to a minimum one that is validated to have only identifiable parameters. The newly proposed approach provides a new method for model validation and evaluation of the predictive power of cardiac models.
Anna Sher, Ken Wang, Andrew J. Wathen, Gary R. Mirams, David Abramson 0001, David Gavaghan
eScience5
2010 Firewall Traversal in the Grid Architecture
abstract
Computational grids have been at the forefront of e-science supercomputing for several years now. Many issues have been settled over the years, but some may only appear to be so. One of these issues is firewall traversal. Many solutions have been proposed and developed. We have developed two such solutions ourselves: Remus and Romulus, but like all other solutions, they are limited in application. Others are working on proposed standards or solutions based on existing Internet standards and RFCs. However, we still have production-level grids that instead operate their grid resources on an open firewall policy. Some propose moving grids on top of peer-topeer networks and/or overlay networks, or rebuilding grids on top of clouds instead. Existing grid infrastructures have not rushed to follow either path, however, as the required changes will take considerable effort and cost for currently running systems. This paper investigates the problem and offers a different proposal: a minor revision to the grid architecture. In order to support what we propose, we will look at several proposed solutions and identify their limitations. We also classify them into two distinct approaches, and discuss how each one is not by itself sufficient for all situations. Then we shall show that a slight improvement to the grid protocol architecture provides a multi-pronged architectural solution.
Jefferson Tan, David Abramson 0001, Colin Enticott
HPCC2
2010 Data centric highly parallel debugging
abstract
Debugging parallel programs is an order of magnitude more complex than sequential ones, and yet, most parallel debuggers provide little extra functionality than their sequential counterparts. This problem becomes more serious as computational codes become more complex, involving larger data structures, and as the machines become larger. Peta-scale machines consisting of millions of cores pose a significant challenge for existing techniques. We argue that debugging must become more data-centric, and believe that "assertions" provide a useful model. Assertions allow a user to declare their expectations about the program state as a whole rather than focusing on that of only a single process state. Previously, we have implemented a special type of assertion that supports debugging applications as they evolve or are ported to different platforms. They allow a user to compare the state of one program against another reference version. These 'relative debugging' assertions, whilst powerful, pose significant implementation challenges for large peta-scale machines. In this paper we discuss a hashing technique that provides a scalable solution for very large problems on very large machines. We illustrate the scheme on 65k cores of Kraken, a Cray XT5 at the University of Tennessee.
David Abramson 0001, Minh Ngoc Dinh, Donny Kurniawan, Bob Moench, Luiz De Rose
HPDC1
2009 Virtual Microscopy and Analysis Using Scientific Workflows
abstract
Most commercial microscopes are stand-alone instruments, controlled by dedicated computer systems. These provide limited storage and processing capabilities. Virtual microscopes, on the other hand, link the image capturing hardware and data analysis software into a wide area network of high performance computers, large storage devices and software systems. In this paper we discuss extensions to Grid workflow engines that allow them to execute scientific experiments on virtual microscopes. We demonstrate the utility of such a system in a biomedical case study concerning the imaging of cancer and antibody based therapeutics.
David Abramson 0001, Blair Bethwaite, Minh Ngoc Dinh, Colin Enticott, Stephen Firth, Slavisa Garic, Ian Harper, Martin Lackmann, Tirath Ramdas, A. B. M. Russel, Stefan Schek, Mary Vail
eScience1
2009 Scheduling Multiple Parameter Sweep Workflow Instances on the Grid
abstract
Due to its ability to provide high-performance computing environment, the grid has become an important infrastructure to support eScience. To utilise the grid for parameter sweep experiments, workflow technology combined with tools such as Nimrod/K are used to orchestrate and automate scientific services provided on the grid. As parameter sweeping over a workflow needs to be executed numerous times, it is more efficient to execute multiple instances of the workflow in parallel. However, this parallel execution can be delayed as every workflow instance requires the same set of resources leading to resource competition problem. Although many algorithms exist for scheduling grid workflows, there is little effort in considering multiple workflow instances and resource competition in the scheduling process. In this paper, we proposed a scheduling algorithm for parameter sweep workflow based on resource competition. The proposed algorithm aims to support multiple workflow instances and avoid allocating resources with high resource competition to minimise delay due to the blocking of tasks. The result is evaluated using simulation to compare with an existing scheduling algorithm.
Sucha Smanchat, Maria Indrawan, Sea Ling, Colin Enticott, David Abramson 0001
eScience5
2009 A Virtual Connectivity Layer for Grids
abstract
Computational grids are now mainstream facilities for e-research worldwide. While enterprise grids exist within organizations, national grids have become common, usually consisting of government as well as academic facilities. Such facilities are not uncommonly lenient with blanket policies to allow inbound and outbound grid traffic. This is far from ideal, from a security perspective, but given the dynamic nature of grid use, it is impractical to keep restrictive firewalls and manually keep up with on-demand firewall reconfiguration. Other solutions are necessary, where security is not sacrificed. Apart from first generation solutions that were mostly not sufficiently generic, standardization work is now ongoing, but exclusively aimed at firewall virtualization. We argue for an architectural solution that encompasses firewall virtualization as well as other methods that can be more appropriate in many environments. This paper describes our notion of the missing layer between grid and fabric, which we refer to as the virtual connectivity layer. We have developed two implementations within this layer and discuss how they fit into a complete and well-defined architectural solution.
Jefferson Tan, David Abramson 0001, Colin Enticott
eScience2
2009 Introduction
Jon B. Weissman, Lex Wolters, David Abramson 0001, Marty Humphrey
Euro-Par3
2009 High-throughput protein structure determination using grid computing
abstract
Determining the X-ray crystallographic structures of proteins using the technique of molecular replacement (MR) can be a time and labor-intensive trial-and-error process, involving evaluating tens to hundreds of possible solutions to this complex 3D jigsaw puzzle. For challenging cases indicators of success often do not appear until the later stages of structure refinement, meaning that weeks or even months could be wasted evaluating MR solutions that resist refinement and do not lead to a final structure. In order to improve the chances of success as well as decrease this timeframe, we have developed a novel grid computing approach that performs many MR calculations in parallel, speeding up the process of structure determination from weeks to hours. This high-throughput approach also allows parameter sweeps to be performed in parallel, improving the chances of MR success.
Jason W. Schmidberger, Blair Bethwaite, Colin Enticott, Mark A. Bate, Steve G. Androulakis, Noel Faux, Cyril F. Reboul, Jennifer M. N. Phan, James C. Whisstock, Wojtek Goscinski, Slavisa Garic, David Abramson 0001, Ashley M. Buckle
IPDPS12
2009 Fault-tolerant execution of large parameter sweep applications across multiple VOs with storage constraints
abstract
Abstract Applications that span multiple virtual organizations (VOs) are of great interest to the e‐science community. However, our recent attempts to execute large‐scale parameter sweep applications (PSAs) for real‐world climate studies with the Nimrod/G tool have exposed problems in the areas of fault tolerance, data storage and trust management. In response, we have implemented a task‐splitting approach that facilitates breaking up large PSAs into a sequence of dependent subtasks, improving fault tolerance; provides a garbage collection technique that deletes unnecessary data; and employs a trust delegation technique that facilitates flexible third party data transfers across different VOs. Copyright © 2008 John Wiley & Sons, Ltd.
Shahaan Ayyub, David Abramson 0001, Colin Enticott, Slavisa Garic, Jefferson Tan
Concurr. Comput. Pract. Exp.2
2009 Interoperation of world-wide production e-Science infrastructures
abstract
Abstract Many production Grid and e‐Science infrastructures have begun to offer services to end‐users during the past several years with an increasing number of scientific applications that require access to a wide variety of resources and services in multiple Grids. Therefore, the Grid Interoperation Now—Community Group of the Open Grid Forum—organizes and manages interoperation efforts among those production Grid infrastructures to reach the goal of a world‐wide Grid vision on a technical level in the near future. This contribution highlights fundamental approaches of the group and discusses open standards in the context of production e‐Science infrastructures. Copyright © 2009 John Wiley & Sons, Ltd.
Morris Riedel, Erwin Laure, Thomas Soddemann, Laurence Field, John-Paul Navarro, James Casey, Maarten Litmaath, Jean-Philippe Baud, Birger Koblitz, Charles E. Catlett, Dane Skow, Cindy Zheng, Philip M. Papadopoulos, Mason J. Katz, Neha Sharma 0001, Oxana Smirnova, Balázs Kónya, Peter W. Arzberger, Frank Würthwein, Abhishek Singh Rana, Terrence Martin, M. Wan, Von Welch, Tony Rimovsky, Steven J. Newhouse, Andrea Vanni, Yoshio Tanaka, Yusuke Tanimura, Tsutomu Ikegami, David Abramson 0001, Colin Enticott, Graham Jenkins, Ruth Pordes, Steven Timm, Gidon Moont, Mona Aggarwal, Dave Colling, Olivier van der Aa, Alex Sim, Vijaya Natarajan, Arie Shoshani, Junmin Gu, Gerson Galang, Riccardo Zappi, Luca Magnoni, Vincenzo Ciaschini, Michele Pace, Valerio Venturi, Moreno Marzolla, Paolo Andreetto, Robert Cowles, Shaowen Wang 0001, Yuji Saeki, Hitoshi Sato, Satoshi Matsuoka, Putchong Uthayopas, Somsak Sriprayoonsakul, Oscar Koeroo, Matthew Viljoen, Laura Pearlman, Stephen Pickles, David Wallom, Glenn Moloney, Jerome Lauret, Jim Marsteller, Paul Sheldon, Surya Pathak, Shaun De Witt, Jirí Mencák, Jens Jensen, Matt Hodges, Derek Ross, Sugree Phatanapherom, Gilbert Netzer, Anders Rhod Gregersen, Mike Jones 0002, Péter Kacsuk, Achim Streit, Daniel Mallmann, Felix Wolf 0001, Thomas Lippert, Thierry Delaitre, Eduardo Huedo, Neil Geddes
Concurr. Comput. Pract. Exp.30
2009 REMUS: A Rerouting and Multiplexing System for Grid Connectivity Across Firewalls
Jefferson Tan, David Abramson 0001, Colin Enticott
J. Grid Comput.2
2009 Relative debugging in an integrated development environment
abstract
Abstract Relative Debugging allows a user to compare the internal state of two programs as they run, making it possible to test whether two programs perform the same function given the same input. When implemented with a command line user interface, a relative debugger looks like traditional debugging tools with the addition of commands that describe which structures should be equivalent in the two programs. In this paper, we discuss relative debugging within an integrated development environment, and show that there are significant advantages over a command line form. We describe a pluggable, modular, architecture that works with a variety of different products, including Microsoft's Visual Studio, SUN's NetBeans, and IBM's Eclipse. Copyright © 2009 John Wiley & Sons, Ltd.
David Abramson 0001, Clement Chu, Donny Kurniawan, Aaron Searle
Softw. Pract. Exp.1
2008 Run-time thread sorting to expose data-level parallelism
abstract
We address the problem of data parallel processing for computational quantum chemistry (CQC). CQC is a computationally demanding tool to study the electronic structure of molecules. An important algorithmic component of these computations is the evaluation of Electron Repulsion Integrals (ERIs). A key problem with ERI evaluation is controlflow variation between different ERI evaluations, which can only be resolved at runtime. This causes the computation to be unsuitable for data parallel execution. However, it is observed that although there is variation between ERI evaluations, the variation is limited; in fact there are a limited number of ERI classes present within any given workload. Conceptually, it is possible to classify the ERIs into sizable sets, and execute these sets in a data parallel fashion. Practically, creating these sets is computationally expensive. We describe an architecture to perform this thread sorting, where high throughput is achieved with small associative and multiport memories. The performance of the prototype is evaluated with FPGA synthesis. We go on to envision other uses for thread sorting, in general-purpose manycore architectures.
Tirath Ramdas, Gregory K. Egan, David Abramson 0001, Kim K. Baldridge
ASAP3
2008 Parameter Estimation Using Scientific Workflows
abstract
In recent years there has been a great deal of interest in "scientific workflows". These allow scientists to specify large computational experiments involving a range of different activities, such as data integration, modelling and analysis, and visualization, to name a few. Activities can be composed, often using a graphical programming environment, so that the output of one stage can be passed as input to the next, forming a pipeline of arbitrary complexity. Scientific workflows have been used to great effect in a number of different disciplines including computational chemistry, ecology and bioinformatics.
David Abramson 0001, Colin Enticott, Thomas C. Peachey
eScience1
2008 Grid Interoperability: An Experiment in Bridging Grid Islands
abstract
In the past decade Grid computing has matured considerably. A number of groups have built, operated, and expanded large testbed and production Grids. These Grids have inevitably been designed to meet the needs of a limited set of initial stakeholders, resulting in varying and sometimes ad-hoc specifications. As the use of e-Science becomes more common, this inconsistency is increasingly problematic for the growing set of applications requiring more resources than a single Grid can offer, as spanning these Grid islands is far from trivial. Thus, Grid interoperability is attracting much interest as researchers try to build bridges between separate Grids. Recently we ran a case study that tested interoperation between several Grids, during which we recorded and classified the issues that arose. In this paper we provide empirical evidence supporting existing interoperability efforts, and identify current and potential barriers to Grid interoperability.
Blair Bethwaite, David Abramson 0001, Ashley M. Buckle
eScience2
2008 A Programming Framework for Incremental Data Distribution in Iterative Applications
abstract
Successful HPC over desktop grids and non-dedicated NOWs is challenging, since good performance is difficult to achieve due to dynamic workloads. On iterative data-parallel applications, this is addressed by dynamic data distribution. However, current approaches migrate an application from one distribution to another in one single phase, which can impact performance. In this paper, we present D3-ARC, a programming framework to support adaptive and incremental data distribution, so that data migration takes place over several successive iterations. D3-ARC consists of a runtime system and an API for specifying the distribution of arrays as well as how data redistribution takes place. We demonstrate how D3-ARC can be used to develop an incremental strategy for data distribution in a Poisson solver, utilising a runtime feedback mechanism to determine how much data to migrate during each iteration.
Philip Chan 0002, David Abramson 0001
ISPA2
2008 Nimrod/K: towards massively parallel dynamic grid workflows
abstract
A challenge for Grid computing is the difficulty in developing software that is parallel, distributed and highly dynamic. Whilst there have been many general purpose mechanisms developed over the years, Grid programming still remains a low level, error prone task. Scientific workflow engines can double as programming environments, and allow a user to compose dasiavirtualpsila Grid applications from pre-existing components. Whilst existing workflow engines can specify arbitrary parallel programs, (where components use message passing) they are typically not effective with large and variable parallelism. Here we discuss dynamic dataflow, originally developed for parallel tagged dataflow architectures (TDAs), and show that these can be used for implementing Grid workflows. TDAs spawn parallel threads dynamically without additional programming. We have added TDAs to Kepler, and show that the system can orchestrate workflows that have large amounts of variable parallelism. We demonstrate the system using case studies in chemistry and in cardiac modelling.
David Abramson 0001, Colin Enticott, Ilkay Altintas
SC1
2008 Parallel programming on a high-performance application-runtime
abstract
Abstract High‐performance application development remains challenging, particularly for scientists making the transition to a heterogeneous grid environment. In general areas of computing, virtual environments such as Java and .Net have proved to be successful in fostering application development, allowing users to target and compile to a single environment, rather than a range of platforms, instruction sets and libraries. However, existing runtime environments are focused on business and desktop computing and they do not support the necessary high‐performance computing (HPC) abstractions required by e‐Scientists. Our work is focused on developing an application‐runtime that can support these services natively. The result is a new approach to the development of an application‐runtime for HPC: the Motor system has been developed by integrating a high‐performance communication library directly within a virtual machine. The Motor message passing library is integrated alongside and in cooperation with other runtime libraries and services while retaining a strong message passing performance. As a result, the application developer is provided with a common environment for HPC application development. This environment supports both procedural languages, such as C, and modern object‐oriented languages, such as C#. This paper describes the unique Motor architecture, presents its implementation and demonstrates its performance and use. Copyright © 2008 John Wiley & Sons, Ltd.
Wojtek Goscinski, David Abramson 0001
Concurr. Comput. Pract. Exp.2
2008 Special Issue: Middleware for Grid Computing
abstract
We are pleased to present a special issue of best papers presented at the 4th International Workshop on Middleware for Grid Computing (MGC 2006), which was held in conjunction with the ACM/IFIP/USENIX 7th International Middleware Conference during November 2006 in Melbourne, Australia. The goal of the fourth edition of the MGC workshop was to bring together researches in the field of middleware, addressing large-scale and real-world problems in grid environments, including Grid and Web Services Architectures and Middleware, Object Metadata and Schemas in Grid Middleware, Programming Models and Tools in Grid Middleware, Resource Management and Scheduling, Strategies and Protocols for QoS in Grid Middleware, Information Services in Grid Middleware, Dependability and Fault Tolerance in Grid Middleware, Performance Evaluation and Modeling in Grid Middleware, Agent-based Approaches in Grid Middleware, Utility Computing, Core Grid Infrastructure, Grid Security, Virtualization in Grid Middleware, Wireless Grids Middleware, Data Grid Middleware, Peer to Peer Protocols in Grid Computing, Network Support for Grid Computing, Grid Application Frameworks, Portals, Portlet Containers, Adaptive-Autonomic Middleware, and Messaging in Grid Middleware. The MGC 2006 workshop received 42 papers from researchers all over the world. Each paper was peer reviewed by three program committee members and external reviewers. Based on their recommendations, we accepted 14 highly rated papers as regular papers. All accepted papers were presented at the workshop in Australia. During the review of workshop papers, Program Committee members were requested to nominate top papers for possible consideration for a journal special issue. From highly rated papers, we selected the top six papers for a special issue and invited authors to submit a substantially extended version of their papers for possible consideration to this special issue. These extended papers went through a rigorous peer review process of the journal managed through the CCPE manuscript management system. Each extended paper was reviewed by external referees. Based on these review comments, we selected five papers for inclusion in this special issue. The selected papers broadly cover key research issues in grid computing middleware as summarized below. Zhang and Honeyman 1 consider that sharing data in global scientific collaborations demands a large-scale storage system that is reliable and efficient, yet at the same time convenient to use. To meet these requirements, this paper presents a network file system (NFSv4) that supports mutable replication with strong consistency guarantees. Experimental evaluation shows that the system holds great promise for accessing and sharing data in grid computing, delivering superior performance, even for write-intensive applications and widely dispersed storage servers, while rigorously adhering to conventional file system semantics. Chadwick et al. 2 present access control in Grid Services, describing a novel approach to making access control decisions in grids throughout time and space. While conventional policy decision points (PDPs) such as Sun's eXtensible Access Control Markup Language (XACML) PDP can do neither, and history-based PDPs can only make temporal access control decisions, the work of Chadwick et al. is one of the first attempts to provide a mechanism for distributed and temporal access control decision-making. It does this by storing the ‘retained access control decision information’ (a concept from ISO 10181-3) in a grid-enabled SQL database, which is available to each PDP in the grid. Obligations in each PDP's policy require this information to be updated when access is granted. In this way, each PDP has available to it a history of the previous access control decisions throughout the entire grid. Bittencourt and Madeira 3 present a dynamic adaptive approach to schedule-dependent tasks onto a grid based on the Xavantes grid middleware. With the concept of rounds, which takes turns sending tasks to execution and evaluating the performance of the resources, the proposed algorithm mixes dynamic scheduling and rescheduling with adaptive behaviour when making scheduling decisions. The algorithm uses an accumulated history of performances to resize the rounds on each resource independently. Bringing together rescheduling and adaptation, the paper opens research topics on more dynamic algorithms, which can bring performance prediction closer to dependent task execution on grids with shared resources. Amar et al. 4 argue that, among existing grid middleware approaches, one simple, powerful, and flexible approach consists of using servers available in different administrative domains through the classic-server model using the remote procedure call (RPC) protocol. Clients submit computation requests to a scheduler whose goal is to find a server available on the grid using several performance metrics. The aim of this paper is to give an overview of the DIET middleware developed in the GRAAL team and to describe recent developments around plug-in schedulers, workflow management, and tools. DIET is a hierarchical set of components used for the development of applications based on computational servers on the grid. Chu et al. 5 present KidneyGrid as a service-oriented Grid environment that provides experimental scientists and analysts access to computational simulations and knowledge databases hosted in separate laboratories around the world involved in human and animal kidney research. It provides a software infrastructure (e.g. Gridbus Grid Service Broker and Web-services-based access to remote kidney models) that can be used for collaborative and shared access to grid resources. The innovation is developed within a specialist community of renal scientists, but will be transferable to any other field of research. We would like to thank the authors for contributing papers on their research in middleware for grid computing to this special publication, and thank all the reviewers for providing constructive reviews and in helping to shape this special issue. Finally, we would like to thank Prof. Geoffrey Fox for providing us an opportunity to bring this special issue. Previous issues can be found at 6-12.
Bruno Schulze, David Abramson 0001, Radha Nandkumar, Rajkumar Buyya
Concurr. Comput. Pract. Exp.2
2007 Executing Large Parameter Sweep Applications on a Multi-VO Testbed
abstract
Applications that span multiple virtual organizations (VOs) are of great interest to the eScience community. However, recent attempts to execute large-scale parameter sweep applications (PSAs) with the Nimrod/G tool have exposed problems in the areas of fault tolerance, data storage and trust management. In response, we have implemented a task-splitting approach, which breaks up large PSAs into a sequence of dependent subtasks, improving fault tolerance; provides a garbage collection technique, which deletes unnecessary data; and employs a trust delegation technique that facilitates flexible third party data transfers across different VOs.
Shahaan Ayyub, David Abramson 0001, Colin Enticott, Slavisa Garic, Jefferson Tan
CCGRID2
2007 Active Data: Supporting the Grid Data Life Cycle
abstract
Scientific applications often involve computation intensive workflows and may generate large amount of derived data. In this paper we consider a life cycle, which starts when the data is first generated, and tracks its progress through replication, distribution, deletion and possible re-computation. We describe the design and implementation of an infrastructure, called active data, which combines existing grid middleware to support the scientific data lifecycle in a platform-neutral environment.
Tim Ho, David Abramson 0001
CCGRID2
2007 A Scalable and Efficient Prefix-Based Lookup Mechanism for Large-Scale Grids
abstract
Data sources, storage, computing resources and services are entities on Grids that require mechanisms for publication and lookup. A discovery service relies on efficient lookup to locate these objects from names or attributes. And on large-scale grids, these should be scalable and be able to support range queries and multi- criteria searching. Trie-based approaches like the PHT and DPT, enable sophisticated queries while providing resilience to failures. In this paper, we propose a more efficient variant of the DPT called the Dynamic Prefix Graph (DPG). And we introduce an encoding scheme that allows DPGs to support both string partial matching and numerical range queries. A distributed algorithm for dynamic construction of DPGs is presented. We conduct simulation studies with the DPG under various data sets. And results demonstrate a reduction of up to 30% in logical hops when lookups are performed under DPGs compared to DPTs.
Philip Chan 0002, David Abramson 0001
eScience2
2007 eResearch Solutions for High Throughput Structural Biology
abstract
Structural biology research places significant demands upon computing and informatics infrastructure. Protein production, crystallization and X-ray data collection require solutions to data management, annotation, target tracking and remote experiment monitoring. Structure elucidation is computationally demanding and requires user-friendly interfaces to high-performance computing resources. Here we discuss how these challenges are being met at the Protein Crystallography Unit at Monash University. Specifically, we have developed informatics solutions for each stage in the structural biology pipeline, from DNA cloning through to protein structure determination. This infrastructure will be pivotal for accelerating the process of structural discovery and will be of significant interest to other laboratories worldwide.
Noel Faux, Anthony Beitz, Mark A. Bate, Abdullah A. Amin, Ian Atkinson, Colin Enticott, Khalid Mahmood 0001, Matthew Swift, Andrew E. Treloar, David Abramson 0001, James C. Whisstock, Ashley M. Buckle
eScience10
2007 An Integrated Grid Development Environment in Eclipse
abstract
With the proliferation of grid computing, a large number of computational resources are available for solving complex scientific and engineering problems. Nevertheless, it is non-trivial to write, deploy, and test grid applications over heterogeneous and distributed resources. Further complicating matters, programmers may need to manually manage variations in source code due to resource heterogeneity. This paper presents an implementation of an integrated grid development environment that leverages IBM's Eclipse IDE and our application development framework, Worqbench. It provides novel tools to develop and debug grid software. It regards resources as first-class objects in the IDE and allows tight integration between the test beds and the code development process. We discuss how the environment assists programmers in developing grid applications.
Donny Kurniawan, David Abramson 0001
eScience2
2007 Persistence and communication state transfer in an asynchronous pipe mechanism
abstract
Emergent wide-area distributed systems like computational grids present opportunities for large scientific applications. On these systems, communication mechanisms have to deal with dynamic resource availability and occurrence of network failures. In this paper, we present the design and implementation of an asynchronous and persistent pipe mechanism, called pi-channels. These communication issues are addressed by combining adaptive caching with data streaming for efficient and fault-tolerant communication. We present the underlying distributed algorithm that implements (a) caching of pipe data segments; (b) asynchronous operation; and (c) re-establishment of connections when a peer leaves and rejoins the computation - part of a communication state transfer mechanism. This makes it possible for different segments (from cache and from writer) of the pipe data to be concurrently streamed to the migrated reader, reducing the retrieval time. Finally, we present some performance results showing the benefits of asynchronous operation.
Philip Chan 0002, David Abramson 0001
ICPADS2
2007 GridRod: a dynamic runtime scheduler for grid workflows
abstract
Grid Workflows are emerging as practical programming models for solving large e-scientific problems on the Grid. However, it is typically assumed that the workflow components either read or write data to conventional files, which are copied from one execution stage to another, or they are tightly coupled using IPC libraries such as MPI or distributed streaming. More flexible communication can be achieved by overloading conventional READ and WRITE operations with advanced IO mechanisms such as sockets, streams and pipes, as is done in the GriddLeS environment. Such flexibility allows the pipelining of temporally dependent components, or in contrast, delaying of tightly coupled computations based on the current resource availability and network connectivity. However, it is also harder to schedule the workflow, because the communication mode may not be decided until run time. In this paper, we propose a new scheduling model that leverages such communication flexibility and allows us to generate dynamic runtime schedules. The scheduler in this case, not only allocates components to distributed Grid resources, but also specifies the inter-component communication mechanism (socket, pipe etc.) The current model is implemented as a dynamic workflow scheduling tool called GridRod, which harnesses Nimrod/G's [1] Grid services and GriddLeS [2] web services.
Shahaan Ayyub, David Abramson 0001
ICS2
2007 A WSRF-Compliant Debugger for Grid Applications
abstract
Grid computing allows the utilization of vast computational resources for solving complex scientific and engineering problems. However, development tools for grid applications are not as mature as their traditional counterparts, especially in the area of debugging and testing. Debugging grid applications typically requires a programmer to address non-trivial issues such as heterogeneity, job scheduling, hierarchical resources, and security. This paper presents the design and implementation of a grid service debug architecture that is compliant with the Web Service Resource Framework standard. The debugger provides a library with a set of well-defined debug APIs.
Donny Kurniawan, David Abramson 0001
IPDPS2
2006 Applications Development for the Computational Grid
David Abramson 0001
APWeb1
2006 Deploying Scientific Applications to the PRAGMA Grid Testbed: Strategies and Lessons
abstract
Recent advances in grid infrastructure and middleware development have enabled various types of applications in science and engineering to be deployed on the grid. The characteristics of these applications and the diverse infrastructure and middleware solutions developed, utilized or adapted by PRAGMA member institutes are summarized. The applications include those for climate modeling, computational chemistry, bioinformatics and computational genomics, remote control of instruments, and distributed databases. Many of the applications are deployed to the PRAGMA grid testbed in routine basis experiments. Strategies for deploying applications without modifications, and those taking advantage of new programming models on the grid are explored and valuable lessons learned are reported. Comprehensive end to end solutions from PRAGMA member institutes that provide important grid middleware components and generalized models of integrating applications and instruments on the grid are also described.
David Abramson 0001, Amanda Lynch, Hiroshi Takemiya, Yusuke Tanimura, Susumu Date, Haruki Nakamura, Karpjoo Jeong, Suntae Hwang, Zhonghua Lu, Céline Amoreira, Kim K. Baldridge, Hurng-Chun Lee, Chi-Wei Wang, Horng-Liang Shih, Tomas E. Molina, Wilfred W. Li, Peter W. Arzberger
CCGRID1
2006 Analysis of Jobs in a Multi-Organizational Grid Test-bed
Bu-Sung Lee, Yew-Soon Ong, Cindy Zheng, Peter W. Arzberger, David Abramson 0001
CCGRID7
2006 The PRAGMA Testbed - Building a Multi-Application International Grid
Cindy Zheng, David Abramson 0001, Peter W. Arzberger, Shahaan Ayyub, Colin Enticott, Slavisa Garic, Mason J. Katz, Jae-Hyuck Kwak, Bu-Sung Lee, Philip M. Papadopoulos, Sugree Phatanapherom, Somsak Sriprayoonsakul, Yoshio Tanaka, Yusuke Tanimura, Osamu Tatebe, Putchong Uthayopas
CCGRID2
2006 A Flexible Grid Framework for Automatic Protein-Ligand Docking
abstract
Many important and fundamental questions in biology and biochemistry can be better understood through investigations performed at the protein-ligand or drugreceptor level. A variety of techniques have been used over the years, and it is an area of active research. In this paper we illustrate an approach that leverages a number of different computational chemistry approaches, and combines these with non-linear optimization algorithms and grid based high performance computing platforms. The result is a very flexible, high performance method of evaluating protein-ligand interaction algorithms. We illustrate the approach by evaluating a hybrid molecular modeling and quantum theoretical based algorithm.
David Abramson 0001, Céline Amoreira, Kim K. Baldridge, Laura Berstis, Chris Kondric, Thomas C. Peachey
e-Science1
2006 A Unified Data Grid Replication Framework
abstract
Modern scientific experiments can generate large amounts of data, which may be replicated and distributed across multiple resources to improve application performance and fault tolerance. Whilst a number of different replica management systems exist, particular communities usually adopt a single system. This creates problems when an application program spans more than one community, because it may need to target more than one middleware layer. One solution to this problem is to build a more flexible data access layer above the specific replica middleware. In this paper, we discuss such an architecture, the Grid Replication Framework, which provides applications with an abstract interface to existing replica systems. Further, the framework's flexible plug-in architecture makes it easy to support new middleware as it becomes available.
Tim Ho, David Abramson 0001
e-Science2
2006 Communication over a Secured Heterogeneous Grid with the GriddLeS Runtime Environment
abstract
Scientific workflows are a powerful programming technique for specifying complex computations using a number of otherwise independent components. When used in a Grid environment, it is possible to build powerful "virtual applications" across multiple distributed and heterogeneous resources. Whilst toolkits such as Globus virtualize many system attributes, and thus make it easier to span different organizations, inconsistent security policies and resource heterogeneity can limit the applicability of workflow techniques. In earlier work, we have described a novel run time environment, GriddLeS, that supports flexible communication patterns between workflow components. GriddLeS abstracts IO operations, such that applications are given the illusion of operating on a local file system, whilst in fact they send and receive data between components. In this paper, we describe how GriddLeS assists in resolving some of the issues that arise due to heterogeneity in security policies and system architectures in a Grid environment. We illustrate the solution using a real world scientific workflow for climate modeling, and demonstrate a system that spans multiple conflicting security domains with heterogeneous resources.
Jagan Kommineni, David Abramson 0001, Jefferson Tan
e-Science2
2006 Worqbench: An Integrated Framework for e-Science Application Development
Donny Kurniawan, David Abramson 0001
e-Science2
2006 Motor: A Virtual Machine for High Performance Computing
abstract
High performance application development remains challenging, particularly for scientists making the transition to a grid environment. In general areas of computing, virtual environments such as Java and .Net have proved successful in fostering application development. Unfortunately, these existing virtual environments do not provide the necessary high performance computing abstractions required by e-scientists. In response, we propose and demonstrate a new approach to the development of a high performance virtual infrastructure: Motor is a virtual machine developed by integrating a high performance message passing library directly within a virtual infrastructure. Motor provides high performance application developers with a common runtime, garbage collection and system libraries, including high performance message passing, whilst retaining strong message passing performance
Wojtek Goscinski, David Abramson 0001
HPDC2
2005 Call-Ordering Constraints
abstract
Several kinds of call-ordering problems have been identified, all of which present subtle difficulties in ensuring the correctness of a sequential program. They include object protocols, synchronisation patterns and re-entrance restrictions. This paper presents call-ordering constraints as a unifying solution to these problems. These constraints are new classes of contracts in addition to traditional preconditions, postconditions and invariants. They extend the traditional notion of behavioural subtyping. The paper shows how constraint inheritance can almost ensure behavioural subtyping conformance. The paper also shows how these constraints may be monitored at run time. Call-ordering constraints are included in the BECON contract system, which has been implemented on the Common Language Infrastructure (CLI).
Nam Tran II, David Abramson 0001, Christine Mingins
APSEC2
2005 Flexible IO Services in the Kepler Grid Workflow System
abstract
Existing grid workflow tools assume that individual components either communicate by passing files from one application to another, or are explicitly linked using interprocess communication pipes. When files are used it is usually difficult to overlap reader and write execution. On the other hand, interprocess communication primitives are invasive and require substantial modification to existing code. We have built a library, called GriddLeS, that combines the advantages of both approaches without any changes to the application code. GriddLeS overloads conventional file IO operations, and supports either local, remote or replicated file access, or direct communication over pipes. In this paper we discuss how GriddLeS can be combined with workflow packages and show how flexible and powerful virtual applications can be constructed rapidly. A large atmospheric science case study is discussed
David Abramson 0001, Jagan Kommineni, Ilkay Altintas
e-Science1
2005 Application Deployment over Heterogeneous Grids using Distributed Ant
abstract
The construction of large scale e-Science grid experiments presents a challenge to e-Scientists because of the inherent difficulty of deploying applications over large scale heterogeneous grids. In spite of this, user-oriented application deployment has remained unsupported in grid middleware. This lack of support for application deployment is strongly detrimental to the usability, evolution, uptake and continual development of the grid. This paper presents our motivation, design and implementation of the distributed ant user-oriented application deployment system, including recent extensions to support application deployment over heterogeneous grids. We also present a significant distributed ant deployment case study, demonstrating how a user-oriented application deployment system enables e-Science experiments
Wojtek Goscinski, David Abramson 0001
e-Science2
2005 The GriddLeS Data Replication Service
abstract
The grid provides infrastructure that allows an arbitrary application to be executed on a range of different computational resources. When input files are very large, or when fault tolerance is important, the data may be replicated. Existing grid data replication middleware suffers from two shortcomings. First, it typically requires modification to existing applications. Second, there is limited support on automatic resource selection and a user usually chooses the replica manually to optimize the performance of the system. In this paper we discuss a middleware layer called the GriddLeS replication service (GRS) that sits above existing replication services, solving both of these shortcomings. Two case studies are presented that illustrate the effectiveness of the approach
Tim Ho, David Abramson 0001
e-Science2
2005 Neuroscience instrumentation and distributed analysis of brain activity data: a case for eScience on global Grids
abstract
The distribution of knowledge (by scientists) and data sources (advanced scientific instruments), and the need for large-scale computational resources for analyzing massive scientific data are two major problems commonly observed in scientific disciplines. Two popular scientific disciplines of this nature are brain science and high-energy physics. The analysis of brain-activity data gathered from the MEG (magnetoencephalography) instrument is an important research topic in medical science since it helps doctors in identifying symptoms of diseases. The data needs to be analyzed exhaustively to efficiently diagnose and analyze brain functions and requires access to large-scale computational resources. The potential platform for solving such resource intensive applications is the Grid. This paper presents the design and development of MEG data analysis system by leveraging Grid technologies, primarily Nimrod-G, Gridbus, and Globus. It describes the composition of the neuroscience (brain-activity analysis) application as parameter-sweep application and its on-demand deployment on global Grids for distributed execution. The results of economic-based scheduling of analysis jobs for three different optimizations scenarios on the world-wide Grid testbed resources are presented along with their graphical visualization. Copyright © 2005 John Wiley & Sons, Ltd.
Rajkumar Buyya, Susumu Date, Yuko Mizuno-Matsumoto, Srikumar Venugopal, David Abramson 0001
Concurr. Comput. Pract. Exp.5
2005 An Atmospheric Sciences Workflow and its implementation with Web Services
David Abramson 0001, Jagan Kommineni, John L. McGregor, Jack Katzfey
Future Gener. Comput. Syst.1
2005 Application of grid computing to parameter sweeps and optimizations in molecular modeling
Wibke Sudholt, Kim K. Baldridge, David Abramson 0001, Colin Enticott, Slavisa Garic, Chris Kondric
Future Gener. Comput. Syst.3
2005 The Grid Economy
abstract
This work identifies challenges in managing resources in a Grid computing environment and proposes computational economy as a metaphor for effective management of resources and application scheduling. It identifies distributed resource management challenges and requirements of economy-based Grid systems, and discusses various representative economy-based systems, both historical and emerging, for cooperative and competitive trading of resources such as CPU cycles, storage, and network bandwidth. It presents an extensive, service-oriented Grid architecture driven by Grid economy and an approach for its realization by leveraging various existing Grid technologies. It also presents commodity and auction models for resource allocation. The use of commodity economy model for resource management and application scheduling in both computational and data grids is also presented.
Rajkumar Buyya, David Abramson 0001, Srikumar Venugopal
Proc. IEEE2
2005 Scheduling parameter sweep applications on global Grids: a deadline and budget constrained cost-time optimization algorithm
abstract
Computational Grids and peer-to-peer (P2P) networks enable the sharing, selection, and aggregation of geographically distributed resources for solving large-scale problems in science, engineering, and commerce. The management and composition of resources and services for scheduling applications, however, becomes a complex undertaking. We have proposed a computational economy framework for regulating the supply of and demand for resources and allocating them for applications based on the users' quality-of-service requirements. The framework requires economy-driven deadline- and budget-constrained (DBC) scheduling algorithms for allocating resources to application jobs in such a way that the users' requirements are met. In this paper, we propose a new scheduling algorithm, called the DBC cost–time optimization scheduling algorithm, that aims not only to optimize cost, but also time when possible. The performance of the cost–time optimization scheduling algorithm has been evaluated through extensive simulation and empirical studies for deploying parameter sweep applications on global Grids. Copyright © 2005 John Wiley & Sons, Ltd.
Rajkumar Buyya, M. Manzur Murshed, David Abramson 0001, Srikumar Venugopal
Softw. Pract. Exp.3
2004 A Taxonomy of Call Ordering Problems
abstract
The order of method calls in a program can present subtle problems in ensuring the program's correctness. Some of the problems have been known under different names in the open literature. These include protocols, synchronisation, re-entrance, mandatory calls, and the indirect invariant effect. However, all these problems relate to the temporal ordering of method calls. In essence, the orderings constrain invocations of methods that share program state or otherwise need to cooperate. This paper proposes taxonomy of call ordering problems and their proposed solutions. The taxonomy classifies the problems by showing their common root and a few distinguishing properties. The paper also sketches the key features of a practical unifying solution to these call ordering problems.
Nam Tran II, David Abramson 0001, Christine Mingins
APSEC2
2004 A Flexible IO Scheme for Grid Workflows
abstract
Summary form only given. Computational grids have been proposed as the next generation computing platform for solving large-scale problems in science, engineering, and commerce. There is an enormous amount of interest in applications, called grid workflows in which a number of otherwise independent programs are run in a "pipeline". In practice, there are a number of different mechanisms that can be used to couple the models, ranging from loosely coupled file based IO to tightly coupled message passing. We propose a flexible IO architecture that provides a wide range of mechanisms for building grid workflows without the need for any source code modification and without the need to fix them at design time. Further, the architecture works with legacy applications. We evaluate the performance of our prototype system using a workflow in computational mechanics.
David Abramson 0001, Jagan Kommineni
IPDPS1
2003 An evolutionary programming algorithm for multi-objective optimisation
abstract
This paper describes a new evolutionary programming optimisation algorithm and a method of its application to multi-objective optimization problems. Computational results are presented demonstrating the algorithm's ability to find Pareto-optimal solutions for a real-world problem in radio frequency component design.
Andrew Lewis 0004, David Abramson 0001
IEEE Congress on Evolutionary Computation2
2003 Automating Relative Debugging
abstract
The creation of a new program version based on an existing version is known as software evolution. In 1994, Abramson and Sosic proposed relative debugging to assist users to locate errors in programs developed with software evolutionary techniques. Relative debugging is a paradigm described by D. Abramson et al. (1996) that enables programmers to locate errors by comparing the developed (evolved) program with the original (existing) program as they are concurrently executed. The aim of the proposed research is to further enhance the currently defined paradigm by (partially or completely) automating the process of relatively debugging. We investigate the possibility of automatically identifying the data structures and program points, normally performed by the user, where values should be equivalent during program execution. Minimizing the user's involvement will reduce the cost of enhancing, maintaining and porting software, and has the potential to provide significant productivity gains on current practices in software development.
Aaron Searle, K. John Gough, David Abramson 0001
ASE3
2003 The Virtual Laboratory: a toolset to enable distributed molecular modelling for drug design on the World-Wide Grid
abstract
Abstract Computational Grids are emerging as a new paradigm for sharing and aggregation of geographically distributed resources for solving large‐scale compute and data intensive problems in science, engineering and commerce. However, application development, resource management and scheduling in these environments is a complex undertaking. In this paper, we illustrate the development of a Virtual Laboratory environment by leveraging existing Grid technologies to enable molecular modelling for drug design on geographically distributed resources. It involves screening millions of compounds in the chemical database (CDB) against a protein target to identify those with potential use for drug design. We have used the Nimrod‐G parameter specification language to transform the existing molecular docking application into a parameter sweep application for executing on distributed systems. We have developed new tools for enabling access to ligand records/molecules in the CDB from remote resources. The Nimrod‐G resource broker along with molecule CDB data broker is used for scheduling and on‐demand processing of docking jobs on the World‐Wide Grid (WWG) resources. The results demonstrate the ease of use and power of the Nimrod‐G and virtual laboratory tools for grid computing. Copyright © 2003 John Wiley & Sons, Ltd.
Rajkumar Buyya, Jonathan Giddy, David Abramson 0001
Concurr. Comput. Pract. Exp.4
2003 Debugging scientific applications in the .NET Framework
David Abramson 0001, Greg Watson
Future Gener. Comput. Syst.1
2002 The Virtual Laboratory: A Toolset for Utilising the World-Wide Grid to Design Drug
abstract
Computational Grids [1] enable the sharing, selection and aggregation of distributed resources across multiple organizations for solving large-scale computational and data intensive problems. Molecular modeling for drug
Rajkumar Buyya, Jonathan Giddy, David Abramson 0001
CCGRID4
2002 Economic models for resource management and scheduling in Grid computing
abstract
Abstract The accelerated development in peer‐to‐peer and Grid computing has positioned them as promising next‐generation computing platforms. They enable the creation of virtual enterprises for sharing resources distributed across the world. However, resource management, application development and usage models in these environments is a complex undertaking. This is due to the geographic distribution of resources that are owned by different organizations or peers. The resource owners of each of these resources have different usage or access policies and cost models, and varying loads and availability. In order to address complex resource management issues, we have proposed a computational economy framework for resource allocation and for regulating supply and demand in Grid computing environments. This framework provides mechanisms for optimizing resource provider and consumer objective functions through trading and brokering services. In a real world market, there exist various economic models for setting the price of services based on supply‐and‐demand and their value to the user. They include commodity market, posted price, tender and auction models. In this paper, we discuss the use of these models for interaction between Grid components to decide resource service value, and the necessary infrastructure to realize each model. In addition to usual services offered by Grid computing systems, we need an infrastructure to support interaction protocols, allocation mechanisms, currency, secure banking and enforcement services. We briefly discuss existing technologies that provide some of these services and show their usage in developing the Nimrod‐G grid resource broker. Furthermore, we demonstrate the effectiveness of some of the economic models in resource trading and scheduling using the Nimrod/G resource broker, with deadline and cost constrained scheduling for two different optimization strategies, on the World‐Wide Grid testbed that has resources distributed across five continents. Copyright © 2002 John Wiley & Sons, Ltd.
Rajkumar Buyya, David Abramson 0001, Jonathan Giddy, Heinz Stockinger
Concurr. Comput. Pract. Exp.2
2002 A computational economy for grid computing and its implementation in the Nimrod-G resource broker
David Abramson 0001, Rajkumar Buyya, Jonathan Giddy
Future Gener. Comput. Syst.1
2001 A Case for Economy Grid Architecture for Service-Oriented Grid Computing
abstract
Computational Grids are a promising platform for executing large-scale resource intensive applications. However, resource management and scheduling in the Grid environment is a complex undertaking as resources are (geographically) distributed, heterogeneous in nature, owned by different individuals or organizations with their own policies, have different access and cost models, and have dynamically varying loads and availability. This introduces a number of challenging issues such as site autonomy, heterogeneous interaction, policy extensibility, resource allocation or co-allocation, online control, scalability, transparency, resource brokering, and “computational economy”. A number of Grid systems (such as Globus and Legion) have addressed many of these issues with exception of a computational economy. We argue that a computational economy is required in order to create a real world scalable Grid because it provides a mechanism for regulating the Grid resources demand and supply. It offers incentive for resource owners to be part of the Grid and encourages consumers to optimally utilize resources and balance timeframe and access costs. We propose a ‘computational economy framework ’ that builds on the existing Grid middleware systems and offers an infrastructure for resource management and trading in the Grid environment. We discuss the usage economic models for resource trading in the Nimrod/G resource broker and present deadline and cost-based scheduling experimental results on the Grid.
Rajkumar Buyya, David Abramson 0001, Jonathan Giddy
IPDPS2
2001 An automatic design optimization tool and its application to computational fluid dynamics
abstract
In this paper we describe the Nimrod/O design optimization tool, and its application in computational fluid dynamics. Nimrod/O facilitates the use of an arbitrary computational model to drive an automatic optimization process. This means that the user can parameterise an arbitrary problem, and then ask the tool to compute the parameter values that minimize or maximise a design objective function. The paper describes the Nimrod/O system, and then discusses a case study in the evaluation of an aerofoil problem. The problem involves computing the shape and angle of attack of the aerofoil that maximises the lift to drag ratio. The results show that our general approach is extremely flexible and delivers better results than a program that was developed specifically for the problem. Moreover, it only took us a few hours to set up the tool for the new problem and required no software development.
David Abramson 0001, Andrew Lewis 0004, Thomas C. Peachey, Clive Fletcher
SC1
2000 High Performance Parametric Modeling with Nimrod/G: Killer Application for the Global Grid?
abstract
This paper examines the role of parametric modeling as an application for the global computing grid, and explores some heuristics which make it possible to specific soft real time deadlines for larger computational experiments. We demonstrate the scheme with a case study utilizing the Globus toolkit running on the GUSTO testbed.
David Abramson 0001, Jonathan Giddy, Lew Kotler
IPDPS1
1998 FPGA Based Custom Computing Machines for Irregular Problems
abstract
Over the past few years there has been increased interest in building custom computing machines (CCMs) as a way of achieving very high performance on specific problems. The advent of high density field programmable gate arrays (FPGAs), in combination with new synthesis tools, have made it relatively easy to produce programmable custom machines without building specific hardware. In many cases, the performance achieved by a FPGA based custom computer is attributed to the exploitation of massive concurrency in the underlying application. In this paper we explore the sources of speedup for irregular problems in which is difficult to exploit such parallelism. We highlight 5 main sources of speedup that we have observed, namely the provision of high memory bandwidth, the use of flexible address generation hardware, the use of gather-scatter array operations, the use of lookup tables and the use of multiple tailored arithmetic units. By considering some representative examples of such irregular problems, the paper illustrates that good performance is possible given the current generation of FPGA devices and RISC processors. The paper then explores whether this performance gain will be possible given the next generation of RISC processors and FPGAs. It concludes that the only way to maintain the speedup is to alter the architecture of CCMs in combination with architectural changes to the FPGAs themselves.
David Abramson 0001, Paul Logothetis, Adam Postula, Marcus Randall
HPCA1
1997 Parallel Non-Linear Optimization : Towards The Design Of A Decision Support System For Air Quality Management
abstract
Large numerical simulation codes have been applied to a wide range of scientific and engineering problems. In the environmental arena the ability to predict the results of certain scenarios by computational science has allowed the choice of strategies which maximize desired outcomes (e.g. financial return) whilst minimizing environmental damage. Access to high performance computing resources has focussed attention on the development of environmental decision support systems which can be used by regulatory agencies and industry planners in evaluating different policy options. A common objective is to find a solution which optimizes some pre-defined criteria. In environmental modelling, the type of optimization problems which need to be considered involve non-linear cost functions over both discrete and continuous parameter values.In this paper we address the optimization component of a decision support system, and perform some initial benchmark studies to assess the effectiveness of the overall approach. The algorithm selected for initial study is based on the quasi-Newton BFGS method. Whilst the BFGS algorithm is generally implemented sequentially, because of the focus of the decision support systems described in the paper we are interested in parallelizing the basic algorithm. This is achieved by concurrent evaluation of functions in finite difference approximations to the derivative and a method of interval subdivision in simple bound constrained line searching.In a realistic problem of air quality management, use of the parallel optimization algorithm as part of an optimizing decision support system is shown to have significant performance gains over other methods of solution. In initial tests it uses less than half the evaluations of a computationally demanding numerical simulation previously used simple enumeration techniques require and is four times faster than traditional sequential optimization methods. This case study has successfully demonstrated the application of an optimization system to a core environmental model, and the feasibility of its use to solve real world problems using parallel and distributed supercomputers.
Andrew Lewis 0004, David Abramson 0001, Rod Simpson
SC2
1997 Guard: A Relative Debugger
abstract
A significant amount of software development is evolutionary, involving the modification of already existing programs. To a large extent, the modified programs produce the same results as the original program. This similarity between the original program and the development program is utilized by relative debugging. Relative debugging is a new concept that enables the user to compare the execution of two programs by specifying the expected correspondences between their states. A relative debugger concurrently executes the programs, verifies the correspondences, and reports any differences found. We describe our novel debugger, called Guard, and its relative debugging capabilities. Guard is implemented by using our library of debugging routines, called Dynascope, which provides debugging primitives in heterogeneous networked environments. To demonstrate the capacity of Guard for debugging in heterogeneous environments, we describe an experiment in which the execution of two programs is compared across Internet. The programs are written in different programming languages and executing on different computing platforms. © 1997 by John Wiley & Sons, Ltd.
Rok Sosic, David Abramson 0001
Softw. Pract. Exp.2
1996 A Debugging and Testing Tool for Supporting Software Evolution
David Abramson 0001, Rok Sosic
Autom. Softw. Eng.1
1995 Nimrod: A Tool for Performing Parameterised Simulations Using Distributed Workstations
abstract
This paper discusses Nimrod, a tool for performing parametrised simulations over networks of loosely coupled workstations. Using Nimrod the user interactively generates a parametrised experiment. Nimrod then controls the distribution of jobs to machines and the collection of results. A simple graphical user interface which is built for each application allows the user to view the simulation in terms of their problem domain. The current version of Nimrod is implemented above OSF DCE and runs on DEC Alpha and IBM RS6000 workstations (including a 22 node SP2). Two different case studies are discussed as an illustration of the utility of the system.
David Abramson 0001, Rok Sosic, Jonathan Giddy, B. Hall
HPDC1
1995 Relative Debugging and its Application to the Development of Large Numerical Models
abstract
Because large scientific codes are rarely static objects, developers are often faced with the tedious task of accounting for discrepancies between new and old versions. In this paper, we describe a new technique called relative debugging that addresses this problem by automating the process of comparing a modified code against a correct reference code. We examine the utility of the relative debugging technique by applying a relative debugger called Guard to a range of debugging problems in a large atmospheric circulation model. Our experience confirms the effectiveness of the approach. Using Guard, we are able to validate a new sequential version of the atmospheric model, and to identify the source of a significant discrepancy in a parallel version in a short period of time.
David Abramson 0001, Ian T. Foster, John Michalakes, Rok Sosic
SC1
1992 Addressing Mechanisms for Large Virtual Memories
abstract
Traditionally there has been a clear distinction between computational (short-term) memory and filestore (long-term memory). This distinction has been maintained mainly due to limitations of technology. Recently there has been considerable interest in programming languages and systems which support orthogonal persistence. In such systems arbitrary data structures may persist beyond the life of the program which created them and this distinction is blurred. Systems supporting orthogonal persistence require a persistent store in which to maintain the persistent objects. Such a persistent store can be implemented via an extended virtual memory with addresses large enough to address all objects. Superimposing structure and a protection scheme on these addresses may well result in them being sparsely distributed. An additional incentive for supporting large virtual addresses is an interest in exploiting the potential of very large main memories to achieve supercomputer speed. This paper presents hardware and software mechanisms to implement a paged virtual memory which can be efficiently accessed by large addresses. An implementation of these techniques for a capability-based computer, MONADS-PC, is described.
John Rosenberg, James Leslie Keedy, David Abramson 0001
Comput. J.3
1990 The RMIT Data Flow Computer: A Hybrid Architecture
abstract
This paper examines the two convention models of data-flow machines, static and dynamic. The advantages and disadvantages of each scheme are described. A hybrid architecture is presented which gains the advantages of both. This architecture has been incorporated into a new data-flow machine built at the Royal Melbourne Institute of Technology. Some simulations are presented which illustrate the advantages of the hybrid approach.
David Abramson 0001, Gregory K. Egan
Comput. J.1