VLDB 2026 Research / reviewers in the wild / expert
Jason Maassen
dblp:48/2165
· DBLP profile ↗
42ranked-venue papers
7as first author
3since 2021 · last 2026
0000-0002-8172-4865ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 6 first-author · 1 since 2021Software engineering, systems software and programming languages · 9 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Databases, data management, data science and information retrieval · 4Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RSMM: A focus area maturity model for research software projectsabstractContext: Research software is instrumental in producing research results. It plays a special role in advancing scientific discovery through tasks such as data analysis, simulation, and visualisation. However, despite its importance, the organisations that produce research software face challenges in managing research software projects. Existing frameworks for software engineering often overlook the needs of research software, including research software engineering practices and open science principles. Without clear guidance, organisations that produce research software must develop and invent new techniques for research software engineering, which is a slow and costly process. Objective: This work presents RSMM, a maturity model designed to improve organisational practices in research software project management. Methods: The initial version of RSMM was developed through a systematic literature review. Expert interviews were then conducted to evaluate and refine the model. Finally, multiple case studies were carried out to validate RSMM and demonstrate its applicability in real-world settings. Results: The final version of RSMM (v1.0) comprises 79 best practices, 17 capabilities, and 10 maturity levels, organised into 4 focus areas: ‘Software Project Management’ , ‘Research Software Management’ , ‘Community Engagement’ , and ‘Software Adoptability’ . We provide a comprehensive analysis of RSMM v1.0 and demonstrate its practical applicability through two illustrative case studies. Conclusion: The RSMM is designed to help organisations in evaluating and improving their research software project management by assessing a project’s current maturity level and providing best practices across four focus areas to guide its progression. Deekshitha, Rena Bakhshi, Jason Maassen, Carlos Martinez-Ortiz, Rob van Nieuwpoort, Antti Knutas, Slinger Jansen |
Inf. Softw. Technol. | 3 |
| 2023 | FAIRSECO: An Extensible Framework for Impact Measurement of Research SoftwareabstractThe growing usage of research software in the research community has highlighted the need to recognize and acknowledge the contributions made not only by researchers but also by Research Software Engineers. However, the existing methods for crediting research software and Research Software Engineers have proven to be insufficient. In response, we have developed FAIRSECO, an extensible open source framework with the objective of assessing the impact of research software in research through the evaluation of various factors. The FAIRSECO framework addresses two critical information needs: firstly, it provides potential users of research software with metrics related to software quality and FAIRness. Secondly, the framework provides information for those who wish to measure the success of a project by offering impact data. By exploring the quality and impact of research software, our aim is to ensure that Research Software Engineers receive the recognition they deserve for their valuable contributions. Deekshitha, Siamak Farshidi, Jason Maassen, Rena Bakhshi, Rob van Nieuwpoort, Slinger Jansen |
e-Science | 3 |
| 2022 | Lightning: Scaling the GPU Programming Model Beyond a Single GPUabstractThe GPU programming model is primarily aimed at the development of applications that run one GPU. However, this limits the scalability of GPU code to the capabilities of a single GPU in terms of compute power and memory capacity. To scale GPU applications further, a great engineering effort is typically required: work and data must be divided over multiple GPUs by hand, possibly in multiple nodes, and data must be manually spilled from GPU memory to higher-level memories. We present Lightning: a framework that follows the common GPU programming paradigm but enables scaling to large problems with ease. Lightning supports multi-GPU execution of GPU kernels, even across multiple nodes, and seamlessly spills data to higher-level memories (main memory and disk). Existing CUDA kernels can easily be adapted for use in Lightning, with data access annotations on these kernels allowing Lightning to infer their data requirements and the dependencies between subsequent kernel launches. Lightning efficiently distributes the work/data across GPUs and maximizes efficiency by overlapping scheduling, data movement, and kernel execution when possible. We present the design and implementation of Lightning, as well as experimental results on up to 32 GPUs for eight benchmarks and one real-world application. Evaluation shows excellent performance and scalability, such as a speedup of 57.2 x over the CPU using Lighting with 16 GPUs over 4 nodes and 80 GB of data, far beyond the memory capacity of one GPU. Stijn Heldens, Pieter Hijma, Ben van Werkhoven, Jason Maassen, Rob van Nieuwpoort |
IPDPS | 4 |
| 2020 | Rocket: efficient and scalable all-pairs computations on heterogeneous platformsabstractAll-pairs compute problems apply a user-defined function to each combination of two items of a given data set. Although these problems present an abundance of parallelism, data reuse must be exploited to achieve good performance. Several researchers considered this problem, either resorting to partial replication with static work distribution or dynamic scheduling with full replication. In contrast, we present a solution that relies on hierarchical multi-level software-based caches to maximize data reuse at each level in the distributed memory hierarchy, combined with a divide-and-conquer approach to exploit data locality, hierarchical work-stealing to dynamically balance the workload, and asynchronous processing to maximize resource utilization. We evaluate our solution using three real-world applications, from digital forensics, localization microscopy, and bioinformatics, on different platforms, from desktop machine to a supercomputer. Results shows excellent efficiency and scalability when scaling to 96 GPUs, even obtaining super-linear speedups due to a distributed cache. Stijn Heldens, Pieter Hijma, Ben van Werkhoven, Jason Maassen, Henri E. Bal, Rob van Nieuwpoort |
SC | 4 |
| 2019 | Reference Exascale ArchitectureabstractWhile political commitments for building exascale systems have been made, turning these systems into platforms for a wide range of exascale applications faces several technical, organisational and skills-related challenges. The key technical challenges are related to the availability of data. While the first exascale machines are likely to be built within a single site, the input data is in many cases impossible to store within a single site. Alongside handling of extreme-large amount of data, the exascale system has to process data from different sources, support accelerated computing, handle high volume of requests per day, minimize the size of data flows, and be extensible in terms of continuously increasing data as well as increase in parallel requests being sent. These technical challenges are addressed by the general reference exascale architecture. It is divided into three main blocks: virtualization layer, distributed virtual file system, and manager of computing resources. Its main property is modularity which is achieved by containerization at two levels: 1) application containers - containerization of scientific workflows, 2) micro-infrastructure - containerization of extreme-large data service-oriented infrastructure. The paper also presents an instantiation of the reference architecture - the architecture of the PROCESS project (PROviding Computing solutions for ExaScale ChallengeS) and discuss its relation to the reference exascale architecture. The PROCESS architecture has been used as an exascale platform within various exascale pilot applications. This work will present the requirements and the derived architecture as well as the 5 use cases pilots that it made possible. Martin Bobák 0001, Balázs Somosköi, Mara Graziani, Matti Heikkurinen, Maximilian Höb, Jan Schmidt, Ladislav Hluchý, Adam Belloum, Reginald Cushing, Jan Meizner, Piotr Nowakowski, Viet D. Tran, Ondrej Habala, Jason Maassen |
eScience | 14 |
| 2019 | Unlocking the LOFAR LTAabstractIn this paper we discuss our efforts in "unlocking" the Long Term Archive (LTA) of the LOFAR radio telescope. This is a large (> 43 PB) archive that expands with about 7 PB per year by the ingestion of new observations. It consists of coarsely calibrated "visibilities", i.e. correlations between signals from LOFAR stations. Currently, only a small fraction of the LOFAR LTA consists of sky maps, which are needed for most astronomical research. Unfortunately, creating such sky maps can be challenging, due to the data sizes of the observations and the complexity and compute requirements of the software involved. We try to fix this by enabling a simple one-click-reduction of LOFAR observations into sky maps for any user of this archive. This work was performed as part of the PROCESS 1 project, which aims to provide generalizable open source solutions for user friendly exascale data processing. Hanno Spreeuw, Souley Madougou, Ronald van Haren, Berend Weel, Adam Belloum, Jason Maassen |
eScience | 6 |
| 2018 | A Portable and Scalable Workflow for Detecting Structural Variants in Whole-Genome Sequencing DataabstractCancer affects millions of people worldwide. With the advent of novel DNA sequencing technologies,whole genome sequencing (WGS) is becoming an integral part of cancer diagnostics that can potentially enable tailored treatments of individual patients. To alleviate these problems,a user-friendly, portable and extendable SV calling workflow, sv-callers developed, that includes four state-of-the-art tools to detect SVs in cancer genomes using on-premises HPC systems. The workflow's parallel execution environment enables to scale from a single computer to high-performance compute clusters with minimal effort. The workflow supports SV analysis in either germline or somatic mode, and requires a list of (paired) WGS samples including a reference genome as input. Users may change the workflow parameters and/or software versions using the YAML configuration files. We performed SV analyses on single and paired (tumor/normal) WGS samples, and report on the results obtained using different academic HPC systems. Arnold Kuzniar, Jason Maassen, Stefan Verhoeven, Luca Santuari, Carl Shneider, Wigard Kloosterman, Jeroen de Ridder |
eScience | 2 |
| 2018 | Painting the Picture of Software Impact with the Research Software DirectoryabstractIn this lightning talk we will describe the Research Software Directory; a content management system that is tailored to research software with the goal of enabling a qualitative assessment of software impact and improving software findability. Jurriaan H. Spaaks, Tom Klaver, Stefan Verhoeven, Jason Maassen, Tom Bakker, Atze van der Ploeg, Ben van Werkhoven, Willem Robert van Hage, Rob van Nieuwpoort |
eScience | 4 |
| 2017 | On the complexities of utilizing large-scale lightpath-connected distributed cyberinfrastructureabstractSummary In Autumn 2013, we—an international team of climate scientists, computer scientists, eScience researchers, and e‐Infrastructure specialists—participated in the enlighten your research global competition, organized to showcase advanced lightpath technologies in support of state‐of‐the‐art research questions. As one of the winning entries, our enlighten your research global team embarked on a very ambitious project to run an extremely high resolution climate model on a collection of supercomputers distributed over two continents and connected using an advanced 10 G lightpath networking infrastructure. Although good progress was made, we were not able to perform all desired experiments due to a varying combination of technical problems, configuration issues, policy limitations and lack of (budget for) human resources to solve these issues. In this paper, we describe our goals, the technical and non‐technical barriers, we encountered and provide recommendations on how these barriers can be removed so future project of this kind may succeed. Copyright © 2016 John Wiley & Sons, Ltd. Jason Maassen, Ben van Werkhoven, Maarten A. J. van Meersbergen, Henri E. Bal, Michael Kliphuis, Sandra E. Brunnabend, Henk A. Dijkstra, Gerben van Malenstein, Migiel de Vos, Sylvia Kuijpers, Sander Boele, Jules Wolfrat, Nick Hill, David Wallom, Christian Grimm, Dieter Kranzlmüller, Dinesh Ganpathi, Shantenu Jha, Yaakoub El Khamra, Frank O. Bryan, Benjamin Kirtman, Frank J. Seinstra |
Concurr. Comput. Pract. Exp. | 1 |
| 2015 | A Round Table for Multi-disciplinary Research on Geospatial and Climate DataabstractEarth observation sciences produce large sets of data which are inherently rich in spatial and geo-spatial information. Together with live data collected from monitoring systems and large collections of semantically rich objects they provide new opportunities for advanced eScience research on climatology, urban planing and smart cities. Such combination of heterogeneous data sets forms a new source of knowledge. Efficient knowledge extraction from them is an eScience challenge. It requires efficient bulk data injection from both static and streaming data sources, dynamic adaptation of the physical and logical schema, efficient methods to correlate spatial and temporal data, and flexibility to (re-)formulate the research question at any time. In this work, we present a data management layer over a column-oriented relational data management system that provides efficient analysis of spatiotemporal data. It provides fast data ingestion through different data loaders, tabular and array based storage, and a dynamic step-wise exploration. Romulo Goncalves, Milena Ivanova, Foteini Alvanaki, Jason Maassen, Kostis Kyzirakos, Oscar Martinez-Rubi, Hannes Mühleisen |
e-Science | 4 |
| 2014 | Performance Models for CPU-GPU Data TransfersabstractMany GPU applications perform data transfers to and from GPU memory at regular intervals. For example because the data does not fit into GPU memory or because of internode communication at the end of each time step. Overlapping GPU computation with CPU-GPU communication can reduce the costs of moving data. Several different techniques exist for transferring data to and from GPU memory and for overlapping those transfers with GPU computation. It is currently not known when to apply which method. Implementing and benchmarking each method is often a large programming effort and not feasible. To solve these issues and to provide insight in the performance of GPU applications, we propose an analytical performance model that includes PCIe transfers and overlapping computation and communication. Our evaluation shows that the performance models are capable of correctly classifying the relative performance of the different implementations. Ben van Werkhoven, Jason Maassen, Frank J. Seinstra, Henri E. Bal |
CCGRID | 2 |
| 2014 | Optimizing convolution operations on GPUs using adaptive tiling
Ben van Werkhoven, Jason Maassen, Henri E. Bal, Frank J. Seinstra |
Future Gener. Comput. Syst. | 2 |
| 2013 | Scalable RDF data compression with MapReduceabstractSUMMARY The Semantic Web contains many billions of statements, which are released using the resource description framework (RDF) data model. To better handle these large amounts of data, high performance RDF applications must apply a compression technique. Unfortunately, because of the large input size, even this compression is challenging. In this paper, we propose a set of distributed MapReduce algorithms to efficiently compress and decompress a large amount of RDF data. Our approach uses a dictionary encoding technique that maintains the structure of the data. We highlight the problems of distributed data compression and describe the solutions that we propose. We have implemented a prototype using the Hadoop framework, and evaluate its performance. We show that our approach is able to efficiently compress a large amount of data and scales linearly on both input size and number of nodes. Copyright © 2012 John Wiley & Sons, Ltd. Jacopo Urbani, Jason Maassen, Niels Drost, Frank J. Seinstra, Henri E. Bal |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | User transparent data and task parallel multimedia computing with Pyxis-DT
Timo van Kessel, Ben van Werkhoven, Niels Drost, Jason Maassen, Henri E. Bal, Frank J. Seinstra |
Future Gener. Comput. Syst. | 4 |
| 2012 | User Transparent Data and Task Parallel Multimedia Computing with Pyxis-DTabstractThe research area of Multimedia Content Analysis (MMCA) considers all aspects of the automated extraction of knowledge from multimedia archives and data streams. To satisfy the increasing computational demands of emerging MMCA problems, there is an urgent need to apply High Performance Computing (HPC) techniques. However, as most MMCA researchers are not also HPC experts, in the field there is a demand~for~programming models and tools that are both efficient and easy~to~use. Today several user transparent library-based parallelization tools exist that aim to satisfy both these requirements. Such tools generally use a data parallel approach in which data structures (e.g. video frames) are scattered among the available nodes in a compute cluster. However, for certain MMCA applications a data parallel approach induces intensive communication, which significantly decreases performance. In these situations, we can benefit from applying alternative approaches. This paper presents Pyxis-DT: a user transparent parallel programming model for MMCA applications that employs both data and task parallelism. Hybrid parallel execution is obtained by run-time construction and execution of a task graph consisting of strictly defined building block operations. Each of these building block operations can be executed in data parallel fashion. Results show that for realistic MMCA applications the concurrent use of data and task parallelism can significantly improve performance compared to using either approach in isolation. Timo van Kessel, Niels Drost, Jason Maassen, Henri E. Bal, Frank J. Seinstra |
CCGRID | 3 |
| 2012 | WebPIE: A Web-scale Parallel Inference Engine using MapReduce
Jacopo Urbani, Spyros Kotoulas, Jason Maassen, Frank van Harmelen, Henri E. Bal |
J. Web Semant. | 3 |
| 2012 | Reply to comment on "WebPIE: A Web-scale parallel inference engine using MapReduce"
Jacopo Urbani, Spyros Kotoulas, Jason Maassen, Frank van Harmelen, Henri E. Bal |
J. Web Semant. | 3 |
| 2012 | Corrigendum to "WebPIE: A Web-scale Parallel Inference Engine using MapReduce" [Web Semant. Sci. Serv. Agents World Wide Web 10 (2012) 59-75]
Jacopo Urbani, Spyros Kotoulas, Jason Maassen, Frank van Harmelen, Henri E. Bal |
J. Web Semant. | 3 |
| 2011 | JEL: unified resource tracking for parallel and distributed applicationsabstractAbstract When parallel applications are run in large‐scale distributed environments, such as grids, peer‐to‐peer (P2P) systems, and clouds, the set of resources used can change dynamically as machines crash, reservations end, and new resources become available. It is vital for applications to respond to these changes. Therefore, it is necessary to keep track of the available resources—a problem which is known to be notoriously difficult. In this article we argue that resource tracking must be provided as the standard functionality in the lower parts of the software stack. We propose a general solution to resource tracking: the Join–Elect–Leave (JEL) model. JEL provides unified resource tracking for parallel and distributed applications across environments. JEL is a simple yet powerful model based on notifying when resources have Joined or Left the computation. We demonstrate that JEL is suitable for resource tracking in a wide variety of programming models, ranging from the fixed resource sets traditionally used in MPI‐1 to flexible grid‐oriented programming models. We compare several JEL implementations, and show these to perform and scale well in several real‐world scenarios involving grids, clouds and P2P systems applied concurrently, and wide‐area systems with failing resources. Using JEL, we have won the first prize in a number of international distributed computing competitions. Copyright © 2010 John Wiley & Sons, Ltd. Niels Drost, Rob van Nieuwpoort, Jason Maassen, Frank J. Seinstra, Henri E. Bal |
Concurr. Comput. Pract. Exp. | 3 |
| 2011 | Zorilla: a peer-to-peer middleware for real-world distributed systemsabstractAbstract The inherent complex nature of current distributed computing architectures hinders the widespread adoption of these systems for mainstream use. In general, users have access to a highly heterogeneous set of compute resources, which may include clusters, grids, desktop grids, clouds, and other compute platforms. This heterogeneity is especially problematic when running parallel and distributed applications. Software is needed which easily combines as many resources as possible into one coherent computing platform. In this paper, we introduce Zorilla: peer‐to‐peer (P2P) middleware that creates a single distributed environment from any available set of compute resources. Zorilla imposes minimal requirements on the resource used, is platform independent, and does not rely on central components. In addition to providing functionality on bare resources, Zorilla can exploit locally available middleware. Zorilla explicitly supports distributed and parallel applications, and allows resources from multiple sites to cooperate in a single computation. Zorilla makes extensive use of both virtualization and P2P techniques. We will demonstrate how virtualization and P2P combine into a simple design, while enhancing functionality and ease of use. Together, these techniques bring our goal a step closer: transparent, easy use of resources, even on very heterogeneous distributed systems. Copyright © 2011 John Wiley & Sons, Ltd. Niels Drost, Rob van Nieuwpoort, Jason Maassen, Frank J. Seinstra, Henri E. Bal |
Concurr. Comput. Pract. Exp. | 3 |
| 2010 | OWL Reasoning with WebPIE: Calculating the Closure of 100 Billion Triples
Jacopo Urbani, Spyros Kotoulas, Jason Maassen, Frank van Harmelen, Henri E. Bal |
ESWC (1) | 3 |
| 2010 | Massive Semantic Web data compression with MapReduceabstractThe Semantic Web consists of many billions of statements made of terms that are either URIs or literals. Since these terms usually consist of long sequences of characters, an effective compression technique must be used to reduce the data size and increase the application performance. One of the best known techniques for data compression is dictionary encoding. In this paper we propose a MapReduce algorithm that efficiently compresses and decompresses a large amount of Semantic Web data. We have implemented a prototype using the Hadoop framework and we report an evaluation of the performance. The evaluation shows that our approach is able to efficiently compress a large amount of data and that it scales linearly regarding the input size and number of nodes. Jacopo Urbani, Jason Maassen, Henri E. Bal |
HPDC | 2 |
| 2009 | Introduction
Domenico Talia, Jason Maassen, Fabrice Huet, Shantenu Jha |
Euro-Par | 2 |
| 2009 | Ibis: Real-world problem solving using real-world gridsabstractIbis is an open source software framework that drastically simplifies the process of programming and deploying large-scale parallel and distributed grid applications. Ibis supports a range of programming models that yield efficient implementations, even on distributed sets of heterogeneous resources. Also, Ibis is specifically designed to run in hostile grid environments that are inherently dynamic and faulty, and that suffer from connectivity problems. Recently, Ibis has been put to the test in two competitions organized by the IEEE Technical Committee on Scalable Computing, as part of the CCGrid 2008 and Cluster/Grid 2008 international conferences. Each of the competitions' categories focused either on the aspect of scalability, efficiency, or fault-tolerance. Our Ibis-based applications have won the first prize in all of these categories. In this paper we give an overview of Ibis, and - to exemplify its power and flexibility - we discuss our contributions to the competitions, and present an overview of our lessons learned. Henri E. Bal, Niels Drost, Roelof Kemp, Jason Maassen, Rob van Nieuwpoort, C. van Reeuwijk, Frank J. Seinstra |
IPDPS | 4 |
| 2009 | Assessing the impact of future reconfigurable optical networks on application performanceabstractThe introduction of optical private networks (lightpaths) has significantly improved the capacity of long distance network links, making it feasible to run large parallel applications in a distributed fashion on multiple sites of a computational grid. Besides offering bandwidths of 10 Gbit/s or more, lightpaths also allow network connections to be dynamically reconfigured. This paper describes our experiences with running data-intensive applications on a grid that offers a (manually) reconfigurable optical wide-area network. We show that the flexibility offered by such a network is useful for applications and that it is often possible to estimate the necessary network configuration in advance. Jason Maassen, Kees Verstoep, Henri E. Bal, Paola Grosso, Cees T. A. M. de Laat |
IPDPS | 1 |
| 2009 | eyeDentify: Multimedia Cyber Foraging from a SmartphoneabstractThe recent introduction of smartphones has resulted in an explosion of innovative mobile applications. The computational requirements of many of these applications, however, can not be met by the smartphone itself. The compute power of the smartphone can be enhanced by distributing the application over other compute resources. Existing solutions comprise of a light weight client running on the smartphone and a heavy weight compute server running on, for example, a cloud. This places the user in a dependent position, however, because the user only controls the client application. In this paper, we follow a different model, called cyber foraging, that gives users full control over all parts of the application. We have implemented the model using the Ibis middleware. We evaluate the model using an innovative application in the domain of multimedia computing, and show that cyber foraging increases the application's responsiveness and accuracy whilst decreasing its energy usage. Roelof Kemp, Nicholas Palmer, Thilo Kielmann, Frank J. Seinstra, Niels Drost, Jason Maassen, Henri E. Bal |
ISM | 6 |
| 2009 | Dynamic photonic lightpaths in the StarPlane network
Paola Grosso, Damien Marchal, Jason Maassen, Eric Bernier, Cees T. A. M. de Laat |
Future Gener. Comput. Syst. | 3 |
| 2008 | Experiences with Fine-Grained Distributed Supercomputing on a 10G TestbedabstractThis paper shows how lightpath-based networks can allow challenging, fine-grained parallel supercomputing applications to be run on a grid, using parallel retrograde analysis on DAS-3 as a case study. Detailed performance analysis shows that several problems arise that are not present on tightly-coupled systems like clusters. In particular, flow control, asynchronous communication, and host- level communication overheads become new obstacles. By optimizing these aspects, however, a 10 G grid can obtain high performance for this type of communication-intensive application. The class of large-scale distributed applications suitable for running on a grid is therefore larger than previously thought realistic. Kees Verstoep, Jason Maassen, Henri E. Bal, John W. Romein |
CCGRID | 2 |
| 2008 | Resource tracking in parallel and distributed applicationsabstractwww.cs.vu.nl/ibis In this paper, we introduce the Join-Elect-Leave (JEL) model, a simple yet powerful model for tracking the resources participating in an application. This model is based on the concept of signaling, i.e., notifying the application when resources have Joined or Left the computation. In addition, the model includes Elections, which can be used to select resources with a special role. JEL supports several consistency models and is suitable for resource coordination of a wide variety of applications, ranging from the traditional fixed resource sets used in MPI, to flexible grid-oriented programming models. Categories and Subject Descriptors: C.2.1 [Computer-Communication Networks]: [Distributed Niels Drost, Rob van Nieuwpoort, Jason Maassen, Henri E. Bal |
HPDC | 3 |
| 2007 | Smartsockets: solving the connectivity problems in grid computingabstractTightly coupled parallel applications are increasingly run in Grid environments. Unfortunately, on many Grid sites the ability of machines to create or accept network connections is severely limited by ?rewalls, network address translation (NAT)or non-routed networks. Multi homing further complicates connection setup and machine identi?cation. Although ad-hoc solutions exist for some of these problems, it is usually up to the application's user to discover the cause of the connectivity problems and ?nd a solution. In this paper we describe SmartSockets1 a communication library that lifts this burden by automatically discovering the connectivity problems and solving them with as little support from the user as possible. Jason Maassen, Henri E. Bal |
HPDC | 1 |
| 2007 | Self-adaptive applications on the gridabstractGrids are inherently heterogeneous and dynamic. One important problemin grid computing is resource selection, that is, finding anappropriate resource set for the application. Another problem is adaptation to the changing characteristics of the grid environment. Existing solutions to these two problems require that a performance model for an application is known. However, constructing such models is a complex task. In this paper, we investigate an approach that does not require performance models. We start an application on any set of resources. During the application run, we periodically collect the statistics about the application run and deduce application requirements from these statistics. Then, we adjustthe resource set to better fit the application needs. This approach allows us to avoid performance bottlenecks, such as overloaded WAN links or very slow processors, and therefore can yield significant performance improvements. We evaluate our approach in a number of scenarios typical for the Grid. Gosia Wrzesinska, Jason Maassen, Henri E. Bal |
PPoPP | 2 |
| 2006 | Satin++: Divide-and-Share on the GridabstractDivide-and-conquer is a popular and effective paradigm for writing grid-enabled applications. I t has been shown to perform well i n environments wtth high network latencies and dynamically changing numbers of processors. However, an important disadvantage of the divide-and-conquer paradigm is its limited applicability due to the lack of a shared data abstraction. W e propose a divide-and-share model: the divide-andconquer model extended with shared objects. Shared objects implement a relaxed consistency model called guard consistency. W e have implemented Satin++: a framework for writing divide-and-share applications. With Satin++ we implemented a number of applications including VLSI routing, N-body simulation and a S A T solver. W e evaluate the performance of our model on a cluster supercomputer and on the heterogeneous, wide-area Grid15000 testbed and demonstrate that our applications can achieve high eficiencies on the Grid. Gosia Wrzesinska, Jason Maassen, Kees Verstoep, Henri E. Bal |
e-Science | 2 |
| 2006 | Programming scientific and distributed workflow with Triana servicesabstractAbstract In this paper, we discuss a real‐world application scenario that uses three distinct types of workflow within the Triana problem‐solving environment: serial scientific workflow for the data processing of gravitational wave signals; job submission workflows that execute Triana services on a testbed; and monitoring workflows that examine and modify the behaviour of the executing application. We briefly describe the Triana distribution mechanisms and the underlying architectures that we can support. Our middleware independent abstraction layer, called the Grid Application Prototype (GAP), enables us to advertise, discover and communicate with Web and peer‐to‐peer (P2P) services. We show how gravitational wave search algorithms have been implemented to distribute both the search computation and data across the European GridLab testbed, using a combination of Web services, Globus interaction and P2P infrastructures. Copyright © 2005 John Wiley & Sons, Ltd. David Churches, Gabor Gombás, Andrew Harrison 0001, Jason Maassen, Craig Robinson, Matthew S. Shields, Ian J. Taylor, Ian Wang |
Concurr. Comput. Pract. Exp. | 4 |
| 2006 | Middleware adaptation with the Delphoi serviceabstractAbstract Grid middleware needs to adapt to changing resources for a large variety of operations. Currently, however, there is only low‐level information available about Grid resources, coming from various but functionally isolated monitoring and information systems. In this paper, we present the Delphoi service. It provides a unified interface to the necessary information, and also matches the abstraction level needed by middleware services to adapt their behavior. We describe Delphoi's architecture, the information it provides, and we evaluate the quality of its performance information. Delphoi has been developed as part of the EC‐funded GridLab project and is currently being deployed on the project's testbed for adding adaptivity to GridLab's middleware services. Copyright © 2006 John Wiley & Sons, Ltd. Jason Maassen, Rob van Nieuwpoort, Thilo Kielmann, Kees Verstoep, Mathijs den Burger |
Concurr. Comput. Pract. Exp. | 1 |
| 2005 | Ibis: a flexible and efficient Java-based Grid programming environmentabstractAbstract In computational Grids, performance‐hungry applications need to simultaneously tap the computational power of multiple, dynamically available sites. The crux of designing Grid programming environments stems exactly from the dynamic availability of compute cycles: Grid programming environments (a) need to beportableto run on as many sites as possible, (b) they need to beflexibleto cope with different network protocols and dynamically changing groups of compute nodes, while (c) they need to provideefficient(local) communication that enables high‐performance computing in the first place. Existing programming environments are either portable (Java), or flexible (Jini, Java Remote Method Invocation or (RMI)), or they are highly efficient (Message Passing Interface). No system combines all three properties that are necessary for Grid computing. In this paper, we present Ibis, a new programming environment that combines Java's ‘run everywhere’ portability both with flexible treatment of dynamically available networks and processor pools, and with highly efficient, object‐based communication. Ibis can transfer Java objects very efficiently by combining streaming object serialization with a zero‐copy protocol. Using RMI as a simple test case, we show that Ibis outperforms existing RMI implementations, achieving up to nine times higher throughputs with trees of objects. Copyright © 2005 John Wiley & Sons, Ltd. Rob van Nieuwpoort, Jason Maassen, Gosia Wrzesinska, Rutger F. H. Hofman, Ceriel J. H. Jacobs, Thilo Kielmann, Henri E. Bal |
Concurr. Pract. Exp. | 2 |
| 2004 | An simple and efficient fault tolerance mechanism for divide-and-conquer systemsabstractSummary form only given. We study if fault tolerance can be made simpler and more efficient by exploiting the structure of the application. More specifically, we study divide-and-conquer parallelism, which is a popular and effective paradigm for writing parallel Grid applications. We have designed a novel fault tolerance mechanism for divide-and-conquer applications that reduces the amount of redundant computation by storing results of the discarded in a global (replicated) table. These results can later be reused, thereby minimizing the amount of work lost as a result of a crash. The execution time overhead of our mechanism is close to zero. Our mechanism can handle crashes of multiple processors or entire clusters at the same time.. It can also handle crashes of the root node that initially started the parallel computation. We have incorporated our fault tolerance mechanism in Satin, which is a Java-based divide-and-conquer system. Satin is implemented on top of the Ibis communication library. The core of Ibis is implemented in pure Java, without using any native libraries. The Satin runtime system and our fault tolerance extension also are written entirely in Java. The resulting system therefore is highly portable allowing the software to run unmodified on a heterogeneous Grid. We evaluated the performance of our fault tolerance scheme on a cluster of the Distributed ASCI Supercomputer 2 (DAS-2). In the first part of our tests, we show that the execution time overhead of our mechanism is close to zero. The results of the second part of our tests show that our algorithm salvages most of the work done by alive processors. Finally, we carried out tests on the European GridLab testbed. We ran one of our applications on a set of six heterogeneous parallel machines (four different operating systems, four different architectures) located in four different European countries. After manually killing one of the sites, the program recovered and finished normally. Gosia Wrzesinska, Rob van Nieuwpoort, Jason Maassen, Henri E. Bal |
CCGRID | 3 |
| 2003 | CCJ: object-based message passing and collective communication in JavaabstractAbstract CCJ is a communication library that adds MPI‐like message passing and collective operations to Java. Rather than trying to adhere to the precise MPI syntax, CCJ aims at a clean integration of communication into Java's object‐oriented framework. For example, CCJ uses thread groups to support Java's multithreading model and it allows any data structure (not just arrays) to be communicated. CCJ is implemented entirely in Java, on top of RMI, so it can be used with any Java virtual machine. The paper discusses three parallel Java applications that use collective communication. It compares the performance (on top of a Myrinet cluster) of CCJ, RMI and mpiJava versions of these applications and also compares their code complexity. A detailed performance comparison between CCJ and mpiJava is given using the Java Grande Forum MPJ benchmark suite. The results show that neither CCJ's object‐oriented design nor its implementation on top of RMI impose a performance penalty on applications compared to their mpiJava counterparts. The source of CCJ is available from our Web site http://www.cs.vu.nl/manta. Copyright © 2003 John Wiley & Sons, Ltd. Arnold Nelisse, Jason Maassen, Thilo Kielmann, Henri E. Bal |
Concurr. Comput. Pract. Exp. | 2 |
| 2002 | Programming environments for high-performance Grid computing: the Albatross project
Thilo Kielmann, Henri E. Bal, Jason Maassen, Rob van Nieuwpoort, Lionel Eyraud-Dubois, Rutger F. H. Hofman, Kees Verstoep |
Future Gener. Comput. Syst. | 3 |
| 2001 | Parallel application experience with replicated method invocationabstractAbstract We describe and evaluate a new approach to object replication in Java, aimed at improving the performance of parallel programs. Our programming model allows the programmer to define groups of objects that can be replicated and updated as a whole, using reliable, totally‐ordered broadcast to send update methods to all machines containing a copy. The model has been implemented in the Manta highperformance Java system. We evaluate system performance both with microbenchmarks and with a set of five parallel applications. For the applications, we also evaluate ease of programming, compared to RMI implementations. We present performance results for a Myrinet‐based workstation cluster as well as for a wide‐area distributed system consisting of four such clusters. The microbenchmarks show that updating a replicated object on 64 machines only takes about three times the RMI latency in Manta. Applications using Manta's object replication mechanism perform at least as fast as manually optimized versions based on RMI, while keeping the application code as simple as with naive versions that use shared objects without taking locality into account. Using a replication mechanism in Manta's runtime system enables several unmodified applications to run efficiently even on the wide‐area system. Copyright © 2001 John Wiley & Sons, Ltd. Jason Maassen, Thilo Kielmann, Henri E. Bal |
Concurr. Comput. Pract. Exp. | 1 |
| 2001 | Efficient Java RMI for parallel programmingabstractJava offers interesting opportunities for parallel computing. In particular, Java Remote Method Invocation (RMI) provides a flexible kind of remote procedure call (RPC) that supports polymorphism. Sun's RMI implementation achieves this kind of flexibility at the cost of a major runtime overhead. The goal of this article is to show that RMI can be implemented efficiently, while still supporting polymorphism and allowing interoperability with Java Virtual Machines (JVMs). We study a new approach for implementing RMI, using a compiler-based Java system called Manta. Manta uses a native (static) compiler instead of a just-in-time compiler. To implement RMI efficiently, Manta exploits compile-time type information for generating specialized serializers. Also, it uses an efficient RMI protocol and fast low-level communication protocols.A difficult problem with this approach is how to support polymorphism and interoperability. One of the consequences of polymorphism is that an RMI implementation must be able to download remote classes into an application during runtime. Manta solves this problem by using a dynamic bytecode compiler, which is capable of compiling and linking bytecode into a running application. To allow interoperability with JVMs, Manta also implements the Sun RMI protocol (i.e., the standard RMI protocol), in addition to its own protocol.We evaluate the performance of Manta using benchmarks and applications that run on a 32-node Myrinet cluster. The time for a null-RMI (without parameters or a return value) of Manta is 35 times lower than for the Sun JDK 1.2, and only slightly higher than for a C-based RPC protocol. This high performance is accomplished by pushing almost all of the runtime overhead of RMI to compile time. We study the performance differences between the Manta and the Sun RMI protocols in detail. The poor performance of the Sun RMI protocol is in part due to an inefficient implementation of the protocol. To allow a fair comparison, we compiled the applications and the Sun RMI protocol with the native Manta compiler. The results show that Manta's null-RMI latency is still eight times lower than for the compiled Sun RMI protocol and that Manta's efficient RMI protocol results in 1.8 to 3.4 times higher speedups for four out of six applications. Jason Maassen, Rob van Nieuwpoort, Ronald Veldema, Henri E. Bal, Thilo Kielmann, Ceriel J. H. Jacobs, Rutger F. H. Hofman |
ACM Trans. Program. Lang. Syst. | 1 |
| 2000 | Wide-area parallel programming using the remote method invocation modelabstractJava's support for parallel and distributed processing makes the language attractive for metacomputing applications, such as parallel applications that run on geographically distributed (wide-area) systems. To obtain actual experience with a Java-centric approach to metacomputing, we have built and used a high-performance wide-area Java system, called Manta. Manta implements the Java Remote Method Invocation (RMI) model using different communication protocols (active messages and TCP/IP) for different networks. The paper shows how wide-area parallel applications can be expressed and optimized using Java RMI. Also, it presents performance results of several applications on a wide-area system consisting of four Myrinet-based clusters connected by ATM WANs. We finally discuss alternative programming models, namely object replication, JavaSpaces, and MPI for Java. Copyright © 2000 John Wiley & Sons, Ltd. Rob van Nieuwpoort, Jason Maassen, Henri E. Bal, Thilo Kielmann, Ronald Veldema |
Concurr. Pract. Exp. | 2 |
| 1999 | An Efficient Implementation of Java's Remote Method InvocationabstractJava offers interesting opportunities for parallel computing. In particular, Java Remote Method Invocation provides an unusually flexible kind of Remote Procedure Call. Unlike RPC, RMI supports polymorphism, which requires the system to be able to download remote classes into a running application. Sun's RMI implementation achieves this kind of flexibility by passing around object type information and processing it at run time, which causes a major run time overhead. Using Sun's JDK 1.1.4 on a Pentium Pro/Myri.net cluster, for example, the latency for a null RMI (without parameters or a return value) is 1228 μsec, which is about a factor of 40 higher than that of a user-level RPC. In this paper, we study an alternative approach for implementing RMI, based on native compilation. This approach allows for better optimization, eliminates the need for processing of type information at run time, and makes a light weight communication protocol possible. We have built a Java system based on a native compiler, which supports both compile time and run time generation of marshallers. We find that almost all of the run time overhead of RMI can be pushed to compile time. With this approach, the latency of a null RMI is reduced to 34 μsec, while still supporting polymorphic RMIs (and allowing interoperability with other JVMs). Jason Maassen, Rob van Nieuwpoort, Ronald Veldema, Henri E. Bal, Aske Plaat |
PPoPP | 1 |