VLDB 2026 Research / reviewers in the wild / expert
Marco Danelutto
dblp:35/3852
· DBLP profile ↗
88ranked-venue papers
20as first author
23since 2021 · last 2025
0000-0002-7433-376XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 50 · 8 first-author · 9 since 2021Software engineering, systems software and programming languages · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SPARE: Self-adaptive Platform for Allocating Resources in Emergencies for Urgent Edge ComputingabstractThis paper presents SPARE, a novel serverless platform that supports self-adaptive resource allocation and reconfiguration, thereby increasing the availability of computing resources for time-critical tasks in urgent events. In emergency scenarios, SPARE reallocates resources by forwarding serverless function invocations to the nearest edge nodes having sufficient capacity. Additionally, the platform employs the use of unikernels and lightweight virtualization through Firecracker, which helps to reduce cold start times and improve function responsiveness. The experimental results demonstrate that SPARE is capable of releasing up to one-third of edge nodes within a serverless edge platform, while only experiencing a mild increase in latency, thus maintaining service continuity. Valerio Besozzi, Marco Danelutto, Patrizio Dazzi, Emanuele Carlini 0001, Matteo Mordacchini |
PDP | 2 |
| 2025 | HPC master design: experience from PisaabstractHPC is a widely used term, often referred to the applications, architectures and programming models and tools targeting highly parallel machines such as those of the top500.org lists. Recent advances in computing hardware resources require application of HPC techniques when using much smaller machines. Indeed, proper parallel programming tools and applications are needed also to exploit parallel hardware resources in personal computers (laptops, desktops, servers). This paper outlines key challenges in designing master’s degree programs in HPC and shares lessons learned from various experiences in developing and implementing such programs in Italy and Europe. Marco Danelutto |
PDP | 1 |
| 2024 | Structuring the Continuum
Marco Danelutto, Patrizio Dazzi, Massimo Torquati |
AINA (5) | 1 |
| 2024 | Power Aware Scheduling of Tasks on FPGAs in Data CentersabstractA variety of computing platforms like Field Pro-grammable Gate Array (FPGA), Graphics Processing Unit (GPU) and multicore Central Processing Unit (CPU) in data centers are suitable for the acceleration of data-intensive workloads. FPGA platforms in data centers are significantly gaining popularity for high-performance computations due to their high speed, re-configurable nature and cost-effectiveness. Heterogeneous, highly parallel computational architectures in data centers and high-speed communication technologies like$5\mathrm{G}$are becoming in-creasingly suitable for real-time applications. However, flexibility, cost-effectiveness, high computational capabilities and energy efficiency remain challenging issues in FPGA based data centers. This paper introduces a power-aware scheduling methodology to accommodate execution of periodic hardware tasks within the available FPGAs of a data center at their potentially maximum speed. The proposed methodology guarantees the execution of available tasks using the maximum number of parallel compu-tation units (CUs) possible to implement in the FPGAs with minimum power consumption. The proposed scheduling method-ology is implemented in a data center with multiple Alveo- U50 Xilinx-AMD FPGAs and Vitis 2023 tool. The evidence from the implementation shows the proposed scheduling methodology is efficient compared to existing solutions. Rourab Paul, Marco Danelutto |
PDP | 2 |
| 2024 | General-purpose data stream processing on heterogeneous architectures with WindFlowabstractMany emerging applications analyze data streams by running graphs of communicating tasks called operators. To develop and deploy such applications, Stream Processing Systems (SPSs) like Apache Storm and Flink have been made available to researchers and practitioners. They exhibit imperative or declarative programming interfaces to develop operators running arbitrary algorithms working on structured or unstructured data streams. In this context, the interest in leveraging hardware acceleration with GPUs has become more pronounced in high-throughput use cases. Unfortunately, GPU acceleration has been studied for relational operators working on structured streams only, while non-relational operators have often been overlooked. This paper presents WindFlow, a library supporting the seamless GPU offloading of general partitioned-stateful operators, extending the range of operators that benefit from hardware acceleration. Its design provides high throughput still exposing a high-level API to users compared with the raw utilization of GPUs in Apache Flink. Gabriele Mencagli, Massimo Torquati, Dalvan Griebler, Alessandra Fais, Marco Danelutto |
J. Parallel Distributed Comput. | 5 |
| 2024 | Boosting general-purpose stream processing with reconfigurable hardwareabstractAbstract Reconfigurable devices such as field-programmable gate arrays (FPGAs) offer flexible solutions to workload acceleration with high energy efficiency. Despite such a potential advantage, they often reveal hard to program by application programmers. High-level synthesis languages have been developed to provide higher-level abstractions, allowing the developers to define the FPGA behavior using an imperative programming approach based on C/C++ languages. However, such approaches still leave the developer with the responsibility to harness the low-level optimizations required to develop efficient FPGA programs. Along this line, this paper introduces , a framework helping programmers to develop FPGA-accelerated data stream processing (DSP) applications. The approach provides a high-level Python API to develop the data-flow graph of operators, which is automatically translated into an efficient Vitis source code targeting Xilinx devices. The execution of the bitstreams implementing two benchmark applications showcases the efficiency of using FPGAs for DSP workloads. In general, provides, with a reasonable time-to-solution, higher performance compared with state-of-the-art DSP frameworks. Alberto Ottimo, Gabriele Mencagli, Marco Danelutto |
J. Supercomput. | 3 |
| 2024 | Enhancing self-adaptation for efficient decision-making at run-time in streaming applications on multicoresabstractAbstract Parallel computing is very important to accelerate the performance of computing applications. Moreover, parallel applications are expected to continue executing in more dynamic environments and react to changing conditions. In this context, applying self-adaptation is a potential solution to achieve a higher level of autonomic abstractions and runtime responsiveness. In our research, we aim to explore and assess the possible abstractions attainable through the transparent management of parallel executions by self-adaptation. Our primary objectives are to expand the adaptation space to better reflect real-world applications and assess the potential for self-adaptation to enhance efficiency. We provide the following scientific contributions: (I) A conceptual framework to improve the designing of self-adaptation; (II) A new decision-making strategy for applications with multiple parallel stages; (III) A comprehensive evaluation of the proposed decision-making strategy compared to the state-of-the-art. The results demonstrate that the proposed conceptual framework can help design and implement self-adaptive strategies that are more modular and reusable. The proposed decision-making strategy provides significant gains in accuracy compared to the state-of-the-art, increasing the parallel applications’ performance and efficiency. Adriano Vogel, Marco Danelutto, Massimo Torquati, Dalvan Griebler, Luiz Gustavo Fernandes |
J. Supercomput. | 2 |
| 2023 | A Proposal for a Continuum-aware Programming Model: From Workflows to Services Autonomously Interacting in the Compute ContinuumabstractThis paper proposes a continuum-aware programming model enabling the execution of application workflows across the compute continuum: cloud, fog and edge resources. It simplifies the management of heterogeneous nodes while alleviating the burden of programmers and unleashing innovation. This model optimizes the continuum through advanced development experiences by transforming workflows into autonomous service collaborations. It reduces complexity in positioning/interconnecting services across the continuum. A meta-model introduces high-level workflow descriptions as service networks with defined contracts and quality of service, thus enabling the deployment/management of workflows as first-class entities. It also provides automation based on policies, monitoring and heuristics. Tailored mechanisms orchestrate/manage services across the continuum, optimizing performance, cost, data protection and sustainability while managing risks. This model facilitates incremental development with visibility of design impacts and seamless evolution of applications and infrastructures. In this work, we explore this new computing paradigm showing how it can trigger the development of a new generation of tools to support the compute continuum progress. Marco Aldinucci, Robert Birke, Antonio Brogi, Emanuele Carlini 0001, Massimo Coppola, Marco Danelutto, Patrizio Dazzi, Luca Ferrucci, Stefano Forti 0002, Hanna Kavalionak, Gabriele Mencagli, Matteo Mordacchini, Marcelo Pasin, Federica Paganelli, Massimo Torquati |
COMPSAC | 6 |
| 2023 | FastFlow targeting FPGAsabstractWriting good code for FPGA is a challenge “per se”, but also running already existing and optimized FPGA kernels often requires writing specific “host side” code and some target hardware knowledge to achieve good performances. In this work, we describe a FastFlow extension supporting seamless off loading of tasks to FPGA, once an FPGA kernel is available. In particular, we show how kernels implemented in Vitis and running on XILINX Alveo FPGA boards may be integrated to implement “normal” parallel stages (pipeline stages, map/farm workers) in a structured parallel FastFlow computation. Experimental results are shown, demonstrating the feasibility of the approach. Marco Danelutto, Gabriele Mencagli, Alberto Ottimo, Francesco Iannone, Paolo Palazzari |
PDP | 1 |
| 2023 | Message from the Organizing Committee Chairs: PDP 2023abstractOn behalf of the Organizing Committee, we welcome you to the 31st Euromicro International Conference on Parallel, Distributed, and Network-Based Processing (PDP2023), organized by the Department of Science and Technology of the University of Naples “Parthenope”. The conference was hosted in Naples in the prestigious Villa Doria d'Angri from the 1st to the 3rd of March 2023. Raffaele Montella, Angelo Ciaramella, Marco Lapegna, Marco Danelutto, Dora Blanco Heras |
PDP | 4 |
| 2023 | Message from the General Chairs: PDP 2023abstractWelcome to the 31st Euromicro International Conference on Parallel, Distributed, and Network-Based Processing (PDP2023). Raffaele Montella, Angelo Ciaramella, Marco Lapegna, Marco Danelutto, Dora Blanco Heras |
PDP | 4 |
| 2023 | FSP: a Framework for Data Stream Processing Applications targeting FPGAsabstractFPGA architectures are becoming popular because of their high performance-to-energy ratio. Nonetheless, their effective exploitation is often counterbalanced by a high programming effort, since most of the modern hardware description languages provide only low-level programming abstractions. This paper proposes FSP, a framework to productively support the development of Data Stream Processing applications on CPU+FPGA System-on-Chip devices (SoCs). By exploiting a code generation approach starting from a high-level DSL in Python, FSP generates an efficient OpenCL skeleton implementation of the parallel pipeline on FPGA and the library to be used by host programs to transfer inputs and collect results to/from the FPGA program. The experimental results showcase the effectiveness of FSP on an SoC equipped with an Intel Arria 10 FPGA by running two streaming benchmark applications. Alberto Ottimo, Gabriele Mencagli, Marco Danelutto |
PDP | 3 |
| 2023 | Revisiting self-adaptation for efficient decision-making at run-time in parallel executionsabstractSelf-adaptation is a potential alternative to provide a higher level of autonomic abstractions and run-time responsiveness in parallel executions. However, the recurrent problem is that self-adaptation is still limited in flexibility and efficiency. For instance, there is a lack of mechanisms to apply adaptation actions and efficient decision-making strategies to decide which configurations should be conveniently enforced at run-time. In this work, we are interested in providing and evaluating potential abstractions achievable with self-adaptation transparently managing parallel executions. Therefore, we provide a new mechanism to support self-adaptation in applications with multiple parallel stages executed in multi-cores. Moreover, we reproduce, reimplement, and evaluate an existing decision-making strategy in our scenario. The observations from the results show that the proposed mechanism for self-adaptation can provide new parallelism abstractions and autonomous responsiveness at run-time. On the other hand, there is a need for more accurate decision-making strategies to enable efficient executions of applications in resource-constrained scenarios like multi-cores. Adriano Vogel, Marco Danelutto, Dalvan Griebler, Luiz Gustavo Fernandes |
PDP | 2 |
| 2023 | NAS Parallel Benchmarks with CUDA and beyondabstractAbstract NAS Parallel Benchmarks (NPB) is a standard benchmark suite used in the evaluation of parallel hardware and software. Several research efforts from academia have made these benchmarks available with different parallel programming models beyond the original versions with OpenMP and MPI. This work joins these research efforts by providing a new CUDA implementation for NPB. Our contribution covers different aspects beyond the implementation. First, we define design principles based on the best programming practices for GPUs and apply them to each benchmark using CUDA. Second, we provide ease of use parametrization support for configuring the number of threads per block in our version. Third, we conduct a broad study on the impact of the number of threads per block in the benchmarks. Fourth, we propose and evaluate five strategies for helping to find a better number of threads per block configuration. The results have revealed relevant performance improvement solely by changing the number of threads per block, showing performance improvements from 8% up to 717% among the benchmarks. Fifth, we conduct a comparative analysis with the literature, evaluating performance, memory consumption, code refactoring required, and parallelism implementations. The performance results have shown up to 267% improvements over the best benchmarks versions available. We also observe the best and worst design choices, concerning code size and the performance trade‐off. Lastly, we highlight the challenges of implementing parallel CFD applications for GPUs and how the computations impact the GPU's behavior. Gabriell Alves de Araujo, Dalvan Griebler, Dinei A. Rockenbach, Marco Danelutto, Luiz Gustavo Fernandes |
Softw. Pract. Exp. | 4 |
| 2022 | Towards Parallel Data Stream Processing on System-on-Chip CPU+GPU DevicesabstractData Stream Processing is a pervasive computing paradigm with a wide spectrum of applications. Traditional streaming systems exploit the processing capabilities provided by homogeneous Clusters and Clouds. Due to the transition to streaming systems suitable for IoT/Edge environments, there has been the urgent need of new streaming frameworks and tools tailored for embedded platforms, often available as System-onChips composed of a small multicore CPU and an integrated onchip GPU. Exploiting this hybrid hardware requires special care in the runtime system design. In this paper, we discuss the support provided by the WindFlow library, showing its design principles and its effectiveness on the NVIDIA Jetson Nano board. Gabriele Mencagli, Dalvan Griebler, Marco Danelutto |
PDP | 3 |
| 2022 | Self-adaptation on parallel stream processing: A systematic reviewabstractSummary A recurrent challenge in real‐world applications is autonomous management of the executions at run‐time. In this vein, stream processing is a class of applications that compute data flowing in the form of streams (e.g., video feeds, images, and data analytics), where parallel computing can help accelerate the executions. On the one hand, stream processing applications are becoming more complex, dynamic, and long‐running. On the other hand, it is unfeasible for humans to monitor and manually change the executions continuously. Hence, self‐adaptation can reduce costs and human efforts by providing a higher‐level abstraction with an autonomic/seamless management of executions. In this work, we aim at providing a literature review regarding self‐adaptation applied to the parallel stream processing domain. We present a comprehensive revision using a systematic literature review method. Moreover, we propose a taxonomy to categorize and classify the existing self‐adaptive approaches. Finally, applying the taxonomy made it possible to characterize the state‐of‐the‐art, identify trends, and discuss open research challenges and future opportunities. Adriano Vogel, Dalvan Griebler, Marco Danelutto, Luiz Gustavo Fernandes |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | The Italian research on HPC key technologies across EuroHPCabstractHigh-Performance Computing (HPC) is one of the strategic priorities for research and innovation worldwide due to its relevance for industrial and scientific applications. We envision HPC as composed of three pillars: infrastructures, applications, and key technologies and tools. While infrastructures are by construction centralized in large-scale HPC centers, and applications are generally within the purview of domain-specific organizations, key technologies fall in an intermediate case where coordination is needed, but design and development are often decentralized. A large group of Italian researchers has started a dedicated laboratory within the National Interuniversity Consortium for Informatics (CINI) to address this challenge. The laboratory, albeit young, has managed to succeed in its first attempts to propose a coordinated approach to HPC research within the EuroHPC Joint Undertaking, participating in the calls 2019--20 to five successful proposals for an aggregate total cost of 95M€. In this paper, we outline the working group's scope and goals and provide an overview of the five funded projects, which become fully operational in March 2021, and cover a selection of key technologies provided by the working group partners, highlighting their usage development within the projects. Marco Aldinucci, Giovanni Agosta, Antonio Andreini, Claudio A. Ardagna, Andrea Bartolini, Alessandro Cilardo, Biagio Cosenza, Marco Danelutto, Roberto Esposito, William Fornaciari, Roberto Giorgi, Davide Lengani, Raffaele Montella, Mauro Olivieri, Sergio Saponara, Daniele Simoni, Massimo Torquati |
CF | 8 |
| 2021 | TEXTAROSSA: Towards EXtreme scale Technologies and Accelerators for euROhpc hw/Sw Supercomputing Applications for exascaleabstractTo achieve high performance and high energy efficiency on near-future exascale computing systems, three key technology gaps needs to be bridged. These gaps include: energy efficiency and thermal control; extreme computation efficiency via HW acceleration and new arithmetics; methods and tools for seamless integration of reconfigurable accelerators in heterogeneous HPC multi-node platforms. TEXTAROSSA aims at tackling this gap through a co-design approach to heterogeneous HPC solutions, supported by the integration and extension of HW and SW IPs, programming models and tools derived from European research. Giovanni Agosta, Daniele Cattaneo 0002, William Fornaciari, Andrea Galimberti, Giuseppe Massari, Federico Reghenzani, Federico Terraneo, Davide Zoni, Carlo Brandolese, Massimo Celino, Francesco Iannone, Paolo Palazzari, Giuseppe Zummo, Massimo Bernaschi, Pasqua D'Ambra, Sergio Saponara, Marco Danelutto, Massimo Torquati, Marco Aldinucci, Yasir Arfat, Barbara Cantalupo, Iacopo Colonnelli, Roberto Esposito, Alberto Riccardo Martinelli, Gianluca Mittone, Olivier Beaumont, Bérenger Bramas, Lionel Eyraud-Dubois, Brice Goglin, Abdou Guermouche, Raymond Namyst, Samuel Thibault, Antonio Filgueras, Miquel Vidal, Carlos Álvarez 0001, Xavier Martorell, Ariel Oleksiak, Michal Kulczewski, Alessandro Lonardo, Piero Vicini, Francesca Lo Cicero, Francesco Simula, Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Pier Stanislao Paolucci, Matteo Turisini, Francesco Giacomini, Tommaso Boccali, Simone Montangero, Roberto Ammendola |
DSD | 17 |
| 2021 | Assessing Coding Metrics for Parallel Programming of Stream Processing Programs on Multi-coresabstractFrom the popularization of multi-core architectures, several parallel APIs have emerged, helping to abstract the programming complexity and increasing productivity in application development. Unfortunately, only a few research efforts in this direction managed to show the usability pay-back of the programming abstraction created, because it is not easy and poses many challenges for conducting empirical software engineering. We believe that coding metrics commonly used in software engineering code measurements can give useful indicators on the programming effort of parallel applications and APIs. These metrics were designed for general purposes without considering the evaluation of applications from a specific domain. In this study, we aim to evaluate the feasibility of seven coding metrics to be used in the parallel programming domain. To do so, five stream processing applications implemented with different parallel APIs for multi-cores were considered. Our experiments have shown COCOMO II is a suitable model for evaluating the productivity of different parallel APIs targeting multi-cores on stream processing applications while other metrics are restricted to the code size. Gabriella Andrade, Dalvan Griebler, Rodrigo Pereira dos Santos, Marco Danelutto, Luiz Gustavo Fernandes |
SEAA | 4 |
| 2021 | Towards On-the-fly Self-Adaptation of Stream Parallel PatternsabstractStream processing applications compute streams of data and provide insightful results in a timely manner, where parallel computing is necessary for accelerating the application executions. Considering that these applications are becoming increasingly dynamic and long-running, a potential solution is to apply dynamic runtime changes. However, it is challenging for humans to continuously monitor and manually self-optimize the executions. In this paper, we propose self-adaptiveness of the parallel patterns used, enabling flexible on-the-fly adaptations. The proposed solution is evaluated with an existing programming framework and running experiments with a synthetic and a real-world application. The results show that the proposed solution is able to dynamically self-adapt to the most suitable parallel pattern configuration and achieve performance competitive with the best static cases. The feasibility of the proposed solution encourages future optimizations and other applicabilities. Adriano Vogel, Gabriele Mencagli, Dalvan Griebler, Marco Danelutto, Luiz Gustavo Fernandes |
PDP | 4 |
| 2021 | Latency-aware adaptive micro-batching techniques for streamed data compression on graphics processing unitsabstractSummary Stream processing is a parallel paradigm used in many application domains. With the advance of graphics processing units (GPUs), their usage in stream processing applications has increased as well. The efficient utilization of GPU accelerators in streaming scenarios requires to batch input elements in microbatches, whose computation is offloaded on the GPU leveraging data parallelism within the same batch of data. Since data elements are continuously received based on the input speed, the bigger the microbatch size the higher the latency to completely buffer it and to start the processing on the device. Unfortunately, stream processing applications often have strict latency requirements that need to find the best size of the microbatches and to adapt it dynamically based on the workload conditions as well as according to the characteristics of the underlying device and network. In this work, we aim at implementing latency‐aware adaptive microbatching techniques and algorithms for streaming compression applications targeting GPUs. The evaluation is conducted using the Lempel‐Ziv‐Storer‐Szymanski compression application considering different input workloads. As a general result of our work, we noticed that algorithms with elastic adaptation factors respond better for stable workloads, while algorithms with narrower targets respond better for highly unbalanced workloads. Charles Michael Stein, Dinei A. Rockenbach, Dalvan Griebler, Massimo Torquati, Gabriele Mencagli, Marco Danelutto, Luiz Gustavo Fernandes |
Concurr. Comput. Pract. Exp. | 6 |
| 2021 | The NAS Parallel Benchmarks for evaluating C++ parallel programming frameworks on shared-memory architectures
Junior Loff, Dalvan Griebler, Gabriele Mencagli, Gabriell Alves de Araujo, Massimo Torquati, Marco Danelutto, Luiz Gustavo Fernandes |
Future Gener. Comput. Syst. | 6 |
| 2021 | WindFlow: High-Speed Continuous Stream Processing With Parallel Building BlocksabstractNowadays, we are witnessing the diffusion of Stream Processing Systems (SPSs) able to analyze data streams in near realtime. Traditional SPSs likeStormandFlinktarget distributed clusters and adopt thecontinuous streaming model, where inputs are processed as soon as they are available while outputs are continuously emitted. Recently, there has been a great focus on SPSs for scale-up machines. Some of them (e.g.,BriskStream) still use the continuous model to achieve low latency. Others optimize throughput with batching approaches that are, however, often inadequate to minimize latency for live-streaming applications. Our contribution is to show a novel software engineering approach to design the runtime system of SPSs targeting multicores, with the aim of providing a uniform solution able to optimize throughput and latency. The approach has a formal nature based on the assembly of components calledbuilding blocks, whose composition allows optimizations to be easily expressed in a compositional manner. We use this methodology to build a new SPS calledWindFlow. Our evaluation showcases the benefits ofWindFlow: it provides lower latency than SPSs for continuous streaming, and can be configured to optimize throughput, to perform similarly and even better than batch-based scale-up SPSs. Gabriele Mencagli, Massimo Torquati, Andrea Cardaci, Alessandra Fais, Luca Rinaldi, Marco Danelutto |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2020 | Efficient NAS Parallel Benchmark Kernels with CUDAabstractNAS Parallel Benchmarks (NPB) are one of the standard benchmark suites used to evaluate parallel hardware and software. There are many research efforts trying to provide different parallel versions apart from the original OpenMP and MPI. Concerning GPU accelerators, there are only the OpenCL and OpenACC available as consolidated versions. Our goal is to provide an efficient parallel implementation of the five NPB kernels with CUDA. Our contribution covers different aspects. First, best parallel programming practices were followed to implement NPB kernels using CUDA. Second, the support of larger workloads (class B and C) allow to stress and investigate the memory of robust GPUs. Third, we show that it is possible to make NPB efficient and suitable for GPUs although the benchmarks were designed for CPUs in the past. We succeed in achieving double performance with respect to the state-of-the-art in some cases as well as implementing efficient memory usage. Fourth, we discuss new experiments comparing performance and memory usage against OpenACC and OpenCL state-of-the-art versions using a relative new GPU architecture. The experimental results also revealed that our version is the best one for all the NPB kernels compared to OpenACC and OpenCL. The greatest differences were observed for the FT and EP kernels. Gabriell Alves de Araujo, Dalvan Griebler, Marco Danelutto, Luiz Gustavo Fernandes |
PDP | 3 |
| 2020 | Challenging the abstraction penalty in parallel patterns libraries
José Daniel García, David del Rio Astorga, Marco Aldinucci, Fabio Tordini, Marco Danelutto, Gabriele Mencagli, Massimo Torquati |
J. Supercomput. | 5 |
| 2020 | Simplifying and implementing service level objectives for stream parallelism
Dalvan Griebler, Adriano Vogel, Daniele De Sensi, Marco Danelutto, Luiz Gustavo Fernandes |
J. Supercomput. | 4 |
| 2019 | Accelerating Actor-Based Applications with Parallel PatternsabstractParallel programmers mandate high-level parallel programming tools allowing to reduce the effort of the efficient parallelization of their applications. Parallel programming leveraging parallel patterns has recently received renovated attention thanks to their clear functional and parallel semantics. In this work, we propose a synergy between the well-known Actors-based programming model and the pattern-based parallelization methodology. We present our preliminary results in that direction, discussing and assessing the implementation of the Map parallel pattern by using an Actor-based software accelerator abstraction that seamlessly integrates within the C++ Actor Framework (CAF). The results obtained on the Intel Xeon Phi KNL platform demonstrate good performance figures achieved with negligible programming efforts. Luca Rinaldi, Massimo Torquati, Gabriele Mencagli, Marco Danelutto, Tullio Menga |
PDP | 4 |
| 2019 | Stream Parallelism on the LZSS Data Compression Application for Multi-Cores with GPUsabstractGPUs have been used to accelerate different data parallel applications. The challenge consists in using GPUs to accelerate stream processing applications. Our goal is to investigate and evaluate whether stream parallel applications may benefit from parallel execution on both CPU and GPU cores. In this paper, we introduce new parallel algorithms for the Lempel-Ziv-Storer-Szymanski (LZSS) data compression application. We implemented the algorithms targeting both CPUs and GPUs. GPUs have been used with CUDA and OpenCL to exploit inner algorithm data parallelism. Outer stream parallelism has been exploited using CPU cores through SPar. The parallel implementation of LZSS achieved 135 fold speedup using a multi-core CPU and two GPUs. We also observed speedups in applications where we were not expecting to get it using the same combine data-stream parallel exploitation techniques. Charles Michael Stein, Dalvan Griebler, Marco Danelutto, Luiz Gustavo Fernandes |
PDP | 3 |
| 2019 | Power-aware pipelining with automatic concurrency controlabstractSummary Continuous streaming computations are usually composed of different modules, exchanging data through shared message queues. The selection of the algorithm used to access such queues (ie, theconcurrency control) is a critical aspect both for performance and power consumption. In this paper, we describe the design of automatic concurrency control algorithm for implementing power‐efficient communications on shared‐memory multicores. The algorithm automatically switches between nonblocking and blocking concurrency protocols, getting the best from the two worlds, ie, obtaining the same throughput offered by the nonblocking implementation and the same power efficiency of the blocking concurrency protocol. We demonstrate the effectiveness of our approach using two micro‐benchmarks and two real streaming applications. Massimo Torquati, Daniele De Sensi, Gabriele Mencagli, Marco Aldinucci, Marco Danelutto |
Concurr. Comput. Pract. Exp. | 5 |
| 2019 | Supporting structured parallel program design, development and tuning in FastFlow
Leonardo Gazzarri, Marco Danelutto |
J. Supercomput. | 2 |
| 2019 | Stream parallelism with ordered data constraints on multi-core systems
Dalvan Griebler, Renato B. Hoffmann, Marco Danelutto, Luiz Gustavo Fernandes |
J. Supercomput. | 3 |
| 2019 | On dynamic memory allocation in sliding-window parallel patterns for streaming analytics
Massimo Torquati, Gabriele Mencagli, Maurizio Drocco, Marco Aldinucci, Tiziano De Matteis, Marco Danelutto |
J. Supercomput. | 6 |
| 2018 | Increasing Efficiency in Parallel Programming TeachingabstractThe ability to teach parallel programming principles and techniques is becoming fundamental to prepare a new generation of programmers able to master the pervasive parallelism made available by hardware vendors. Classical parallel programming courses leverage either low-level programming frameworks (e.g. those based on Pthreads) or higher level frameworks such as OpenMP or MPI. We discuss our teaching experience within the Master in "Computer Science and networking" where parallel programming is taught leveraging structured parallel programming principles and frameworks. The paper summarizes the results achieved in eight years of experience and shows how the adoption of a structured parallel programming approach improves the efficiency of the teaching process. Marco Danelutto, Massimo Torquati |
PDP | 1 |
| 2018 | Efficient NAS Benchmark Kernels with C++ Parallel ProgrammingabstractBenchmarking is a way to study the performance of new architectures and parallel programming frameworks. Well-established benchmark suites such as the NAS Parallel Benchmarks (NPB) comprise legacy codes that still lack portability to C++ language. As a consequence, a set of high-level and easy-to-use C++ parallel programming frameworks cannot be tested in NPB. Our goal is to describe a C++ porting of the NPB kernels and to analyze the performance achieved by different parallel implementations written using the Intel TBB, OpenMP and FastFlow frameworks for Multi-Cores. The experiments show an efficient code porting from Fortran to C++ and an efficient parallelization on average. Dalvan Griebler, Junior Loff, Gabriele Mencagli, Marco Danelutto, Luiz Gustavo Fernandes |
PDP | 4 |
| 2018 | Elastic-PPQ: A two-level autonomic system for spatial preference query processing over dynamic data streams
Gabriele Mencagli, Massimo Torquati, Marco Danelutto |
Future Gener. Comput. Syst. | 3 |
| 2018 | Simplifying self-adaptive and power-aware computing with Nornir
Daniele De Sensi, Tiziano De Matteis, Marco Danelutto |
Future Gener. Comput. Syst. | 3 |
| 2018 | A parallel pattern for iterative stencil + reduce
Marco Aldinucci, Marco Danelutto, Maurizio Drocco, Peter Kilpatrick, Claudia Misale, Guilherme Peretti Pezzi, Massimo Torquati |
J. Supercomput. | 2 |
| 2018 | Data stream processing via code annotations
Marco Danelutto, Tiziano De Matteis, Gabriele Mencagli, Massimo Torquati |
J. Supercomput. | 1 |
| 2017 | Evaluating Concurrency Throttling and Thread Packing on SMT MulticoresabstractPower-aware computing is gaining an increasing attention both in academic and industrial settings. The problem of guaranteeing a given QoS requirement (either in terms of performance or power consumption) can be faced by selecting and dynamically adapting the amount of physical and logical resources used by the application. In this study, we considered standard multicore platforms by taking as a reference approaches for power-aware computing two well-known dynamic reconfiguration techniques: Concurrency Throttling and Thread Packing. Furthermore, we also studied the impact of using simultaneous multithreading (e.g., Intel's HyperThreading) in both techniques. In this work, leveraging on the applications of the PARSEC benchmark suite, we evaluate these techniques by considering performance-power trade-offs, resource efficiency, predictability and required programming effort. The results show that, according to the comparison criteria, these techniques complement each other. Marco Danelutto, Tiziano De Matteis, Daniele De Sensi, Massimo Torquati |
PDP | 1 |
| 2017 | Enabling semantics to improve detection of data races and misuses of lock-free data structuresabstractSummary The rapid progress of multi/many‐core architectures has caused data‐intensive parallel applications not yet fully optimized to deliver the best performance. In the advent of concurrent programming, frameworks offering structured patterns have alleviated developers' burden adapting such applications to multithreaded architectures. While some of these patterns are implemented using synchronization primitives, others avoid them by means of lock‐free data mechanisms. However, lock‐free programming is not straightforward, ensuring an appropriate use of their interfaces can be challenging, since different memory models plus instruction reordering at compiler/processor levels can interfere in the occurrence of data races. The benefits of race detectors are formidable in this sense; however, they may emit false positives if are unaware of the underlying lock‐free structure semantics. To mitigate this issue, this paper extends ThreadSanitizer, a race detection tool, with the semantics of 2 lock‐free data structures: the single‐producer/single‐consumer and the multiple‐producer/multiple‐consumer queues. With it, we are able to drop false positives and detect potential semantic violations. The experimental evaluation, using different queue implementations on a set ofμbenchmarks and real applications, demonstrates that it is possible to reduce, on average, 60% the number of data race warnings and detect wrong uses of these structures. Manuel F. Dolz, David del Rio Astorga, Javier Fernández 0001, Massimo Torquati, José Daniel García, Félix García Carballeira, Marco Danelutto |
Concurr. Comput. Pract. Exp. | 7 |
| 2017 | Bringing Parallel Patterns Out of the Corner: The P3 ARSEC Benchmark SuiteabstractHigh-level parallel programming is an active research topic aimed at promoting parallel programming methodologies that provide the programmer with high-level abstractions to develop complex parallel software with reduced time to solution. Pattern-based parallel programming is based on a set of composable and customizable parallel patterns used as basic building blocks in parallel applications. In recent years, a considerable effort has been made in empowering this programming model with features able to overcome shortcomings of early approaches concerning flexibility and performance. In this article, we demonstrate that the approach is flexible and efficient enough by applying it on 12 out of 13 PARSEC applications. Our analysis, conducted on three different multicore architectures, demonstrates that pattern-based parallel programming has reached a good level of maturity, providing comparable results in terms of performance with respect to both other parallel programming methodologies based on pragma-based annotations (i.e., Open mp and O mp S s ) and native implementations (i.e., P threads ). Regarding the programming effort, we also demonstrate a considerable reduction in lines of code and code churn compared to P threads and comparable results with respect to other existing implementations. Daniele De Sensi, Tiziano De Matteis, Massimo Torquati, Gabriele Mencagli, Marco Danelutto |
ACM Trans. Archit. Code Optim. | 5 |
| 2017 | Parallel Continuous Preference Queries over Out-of-Order and Bursty Data StreamsabstractTechniques to handle traffic bursts and out-of-order arrivals are of paramount importance to provide real-time sensor data analytics in domains like traffic surveillance, transportation management, healthcare and security applications. In these systems the amount of raw data coming from sensors must be analyzed by continuous queries that extract value-added information used to make informed decisions in real-time. To perform this task with timing constraints, parallelism must be exploited in the query execution in order to enable the real-time processing on parallel architectures. In this paper we focus on continuous preference queries, a representative class of continuous queries for decision making, and we propose a parallel query model targeting the efficient processing over out-of-order and bursty data streams. We study how to integrate punctuation mechanisms in order to enable out-of-order processing. Then, we present advanced scheduling strategies targeting scenarios with different burstiness levels, parameterized using the index of dispersion quantity. Extensive experiments have been performed using synthetic datasets and real-world data streams obtained from an existing real-time locating system. The experimental evaluation demonstrates the efficiency of our parallel solution and its effectiveness in handling the out-of-orderness degrees and burstiness levels of real-world applications. Gabriele Mencagli, Massimo Torquati, Marco Danelutto, Tiziano De Matteis |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | Introducing Parallelism by Using REPARA C++11 AttributesabstractPatterns provide a mechanism to express parallelism at a high level of abstraction and to make easier the transformation of existing legacy applications to target parallel frameworks. That also opens a path for writing new parallel applications. In this paper we introduce the REPARA approach for expressing parallel patterns and transforming the source code to parallelism frameworks. We take advantage of C++11 attributes as a mechanism to introduce annotations and enrich semantic information on valid source code. We also present a methodology for performing transformation of source code that allows to target multiple parallel programming models. Another contribution is a rule based mechanism to transform annotated code to those specific programming models. The REPARA approach requires programmer intervention only to perform initial code annotation while providing speedups that are comparable to those obtained by manual parallelization. Marco Danelutto, José Daniel García, Luis Miguel Sánchez, Rafael Sotomayor, Massimo Torquati |
PDP | 1 |
| 2016 | RPL: A Domain-Specific Language for Designing and Implementing Parallel C++ ApplicationsabstractParallelising sequential applications is usually a very hard job, due to many different ways in which an application can be parallelised and a large number of programming models (each with its own advantages and disadvantages) that can be used. In this paper, we describe a method to semi-automatically generate and evaluate different parallelisations of the same application, allowing programmers to find the best parallelisation without significant manual reengineering of the code. We describe a novel, high-level domain-specific language, Refactoring Pattern Language (RPL), that is used to represent the parallel structure of an application and to capture its extra-functional properties (such as service time). We then describe a set of RPL rewrite rules that can be used to generate alternative, but semantically equivalent, parallel structures (parallelisations) of the same application. We also describe the RPL Shell that can be used to evaluate these parallelisations, in terms of the desired extra-functional properties. Finally, we describe a set of C++ refactorings, targeting OpenMP, Intel TBB and FastFlow parallel programming models, that semi-automatically apply the desired parallelisation to the application's source code, therefore giving a parallel version of the code. We demonstrate how the RPL and the refactoring rules can be used to derive efficient parallelisations of two realistic C++ use cases (Image Convolution and Ant Colony Optimisation). Vladimir Janjic, Christopher Brown 0002, Kenneth MacKenzie, Kevin Hammond, Marco Danelutto, Marco Aldinucci, José Daniel García |
PDP | 5 |
| 2016 | A Reconfiguration Algorithm for Power-Aware Parallel ApplicationsabstractIn current computing systems, many applications require guarantees on their maximum power consumption to not exceed the available power budget. On the other hand, for some applications, it could be possible to decrease their performance, yet maintain an acceptable level, in order to reduce their power consumption. To provide such guarantees, a possible solution consists in changing the number of cores assigned to the application, their clock frequency, and the placement of application threads over the cores. However, power consumption and performance have different trends depending on the application considered and on its input. Finding a configuration of resources satisfying user requirements is, in the general case, a challenging task. In this article, we propose Nornir, an algorithm to automatically derive, without relying on historical data about previous executions, performance and power consumption models of an application in different configurations. By using these models, we are able to select a close-to-optimal configuration for the given user requirement, either performance or power consumption. The configuration of the application will be changed on-the-fly throughout the execution to adapt to workload fluctuations, external interferences, and/or application’s phase changes. We validate the algorithm by simulating it over the applications of the Parsecbenchmark suit. Then, we implement our algorithm and we analyse its accuracy and overhead over some of these applications on a real execution environment. Eventually, we compare the quality of our proposal with that of the optimal algorithm and of some state-of-the-art solutions. Daniele De Sensi, Massimo Torquati, Marco Danelutto |
ACM Trans. Archit. Code Optim. | 3 |
| 2015 | Energy Driven Adaptivity in Stream Parallel ComputationsabstractDetermining the right amount of resources needed for a given computation is a critical problem. In many cases, computing systems are configured to use an amount of resources to manage high load peaks even though this cause energy waste when the resources are not fully utilised. To avoid this problem, adaptive approaches are used to dynamically increase/decrease computational resources depending on the real needs. A different approach based on Dynamic Voltage and Frequency Scaling (DVFS) is emerging as a possible alternative solution to reduce energy consumption of idle CPUs by lowering their frequencies. In this work, we propose to tackle the problem in stream parallel computations by using both the classic adaptivity concepts and the possibility provided by modern CPUs to dynamically change their frequency. We validate our approach showing a real network application that performs Deep Packet Inspection over network traffic. We are able to manage bandwidth changing over time, guaranteeing minimal packet loss during reconfiguration and minimal energy consumption. Marco Danelutto, Daniele De Sensi, Massimo Torquati |
PDP | 1 |
| 2015 | A Green Perspective on Structured Parallel ProgrammingabstractStructured parallel programming, and in particular programming models using the algorithmic skeleton or parallel design pattern concepts, are increasingly considered to be the only viable means of supporting effective development of scalable and efficient parallel programs. Structured parallel programming models have been assessed in a number of works in the context of performance. In this paper we consider how the use of structured parallel programming models allows knowledge of the parallel patterns present to be harnessed to address both performance and energy consumption. We consider different features of structured parallel programming that may be leveraged to impact the performance/energy trade-off and we discuss a preliminary set of experiments validating our claims. Marco Danelutto, Massimo Torquati, Peter Kilpatrick |
PDP | 1 |
| 2014 | Loop Parallelism: A New Skeleton Perspective on Data Parallel PatternsabstractTraditionally, skeleton based parallel programming frameworks support data parallelism by providing the programmer with a comprehensive set of data parallel skeletons, based on different variants of map and reduce patterns. On the other side, more conventional parallel programming frameworks provide application programmers with the possibility to introduce parallelism in the execution of loops with a relatively small programming effort. In this work, we discuss a "ParallelFor" skeleton provided within the FastFlow framework and aimed at filling the usability and expressivity gap between the classical data parallel skeleton approach and the loop parallelisation facilities offered by frameworks such as OpenMP and Intel TBB. By exploiting the low run-time overhead of the FastFlow parallel skeletons and the new facilities offered by the C++11 standard, our ParallelFor skeleton succeeds to obtain comparable or better performance than both OpenMP and TBB on the Intel Phi many-core and Intel Nehalem multi-core for a set of benchmarks considered, yet requiring a comparable programming effort. Marco Danelutto, Massimo Torquati |
PDP | 1 |
| 2014 | Parallel patterns for heterogeneous CPU/GPU architectures: Structured parallelism from cluster to cloud
Sonia Campa, Marco Danelutto, Mehdi Goli 0001, Horacio González-Vélez, Alina Madalina Popescu, Massimo Torquati |
Future Gener. Comput. Syst. | 2 |
| 2013 | Towards The Deployment Of Fastflow On Distributed Virtual ArchitecturesabstractIn this paper we investigate the deployment of FastFlow applications on multi-core virtual platforms. The overhead introduced by the virtual environment has been measured using a well-known application benchmark both in the sequential and in the FastFlow parallel setting. The overhead introduced for both the sequential and the parallel executions of CPU and memory-intensive applications is in the range of 2-30%, while the execution speedup is almost preserved. Additionally, we have ported the FastFlow benchmark to a cloud-based distributed environment in which a task-intensive application has been tested and the performance compared with the corresponding run on a smaller cluster of multi-core machines without virtualisation.
From a parallel programming perspective, we have demonstrated how a unique programming framework based on the structured parallel programming paradigm can cope with very different kind of target architectures without any (or minimal) code intervention. Sonia Campa, Marco Danelutto, Massimo Torquati, Horacio González-Vélez, Alina Madalina Popescu |
ECMS | 2 |
| 2013 | Parallel Patterns for General Purpose Many-CoreabstractEfficient programming of general purpose many-core accelerators poses several challenging problems. The high number of cores available, the peculiarity of the interconnection network, and the complex memory hierarchy organization, all contribute to make efficient programming of such devices difficult. We propose to use parallel design patterns, implemented using algorithmic skeletons, to abstract and hide most of the difficulties related to the efficient programming of many-core accelerators. In particular, we discuss the porting of the FastFlow framework on the Tilera TilePro64 architecture and the results obtained running synthetic benchmarks as well as true application kernels. These results demonstrate the efficiency achieved while using patterns on the TilePro64 both to program stand-alone skeleton-based parallel applications and to accelerate existing sequential code. Daniele Buono, Marco Danelutto, Silvia Lametti, Massimo Torquati |
PDP | 2 |
| 2013 | A RISC Building Block Set for Structured Parallel ProgrammingabstractWe propose a set of building blocks (RISC-pb2l) suitable to build high-level structured parallel programming frameworks. The set is designed following a RISC approach. RISC-pb2l is architecture independent but the implementation of the different blocks may be specialized to make the best usage of the target architecture peculiarities. A number of optimizations may be designed transforming basic building blocks compositions into more efficient compositions, such that parallel application efficiency may be derived by construction rather than by debugging. Marco Danelutto, Massimo Torquati |
PDP | 1 |
| 2012 | An Efficient Unbounded Lock-Free Queue for Multi-core Systems
Marco Aldinucci, Marco Danelutto, Peter Kilpatrick, Massimiliano Meneghin, Massimo Torquati |
Euro-Par | 2 |
| 2012 | Topic 9: Parallel and Distributed Programming
Sergei Gorlatch, Rizos Sakellariou, Marco Danelutto, Thilo Kielmann |
Euro-Par | 3 |
| 2012 | Parallel Patterns + Macro Data Flow for Multi-core ProgrammingabstractData flow techniques have been around since the early '70s when they were used in compilers for sequential languages. Shortly after their introduction they were also considered as a possible model for parallel computing, although the impact here was limited. Recently, however, data flow has been identified as a candidate for efficient implementation of various programming models on multi-core architectures. In most cases, however, the burden of determining data flow ``macro'' instructions is left to the programmer, while the compiler/run time system manages only the efficient scheduling of these instructions. We discuss a structured parallel programming approach supporting automatic compilation of programs to macro data flow and we show experimental results demonstrating the feasibility of the approach and the efficiency of the resulting ``object'' code on different classes of state-of-the-art multi-core architectures. The experimental results use different base mechanisms to implement the macro data flow run time support, from plain pthreads with condition variables to more modern and effective lock- and fence-free parallel frameworks. Experimental results comparing efficiency of the proposed approach with those achieved using other, more classical, parallel frameworks are also presented. Marco Aldinucci, L. Anardu, Marco Danelutto, Massimo Torquati, Peter Kilpatrick |
PDP | 3 |
| 2011 | Accelerating Code on Multi-cores with FastFlow
Marco Aldinucci, Marco Danelutto, Peter Kilpatrick, Massimiliano Meneghin, Massimo Torquati |
Euro-Par (2) | 2 |
| 2010 | High Performance Architectures and Compilers
Pedro C. Diniz, Marco Danelutto, Denis Barthou, Marc Gonzales, Michael Hübner 0001 |
Euro-Par (1) | 2 |
| 2010 | Perspectives on grid computing
Uwe Schwiegelshohn, Rosa M. Badia, Marian Bubak, Marco Danelutto, Schahram Dustdar, Fabrizio Gagliardi, Alfred Geiger, Ladislav Hluchý, Dieter Kranzlmüller, Erwin Laure, Thierry Priol, Alexander Reinefeld, Michael M. Resch, Andreas Reuter 0001, Otto Rienhoff, Thomas Rüter, Peter M. A. Sloot, Domenico Talia, Klaus Ullmann, Ramin Yahyapour |
Future Gener. Comput. Syst. | 4 |
| 2009 | Stkm on Sca: A Unified Framework with Components, Workflows and Algorithmic Skeletons
Marco Aldinucci, Hinde-Lilia Bouziane, Marco Danelutto, Christian Pérez |
Euro-Par | 3 |
| 2009 | Autonomic management of non-functional concerns in distributed & parallel application programmingabstractAn approach to the management of non-functional concerns in massively parallel and/or distributed architectures that marries parallel programming patterns with autonomic computing is presented. The necessity and suitability of the adoption of autonomic techniques are evidenced. Issues arising in the implementation of autonomic managers taking care of multiple concerns and of coordination among hierarchies of such autonomic managers are discussed. Experimental results are presented that demonstrate the feasibility of the approach. Marco Aldinucci, Marco Danelutto, Peter Kilpatrick |
IPDPS | 2 |
| 2009 | Towards Hierarchical Management of Autonomic Components: A Case StudyabstractWe address the issue of autonomic management in hierarchical component-based distributed systems. The long term aim is to provide a modeling framework for autonomic management in which QoS goals can be defined, plans for system adaptation described and proofs of achievement of goals by (sequences of) adaptations furnished. Here we present an early step on this path. We restrict our focus to skeleton-based systems in order to exploit their well-defined structure. The autonomic cycle is described using the Orc system orchestration language while the plans are presented as structural modifications together with associated costs and benefits. A case study is presented to illustrate the interaction of managers to maintain QoS goals for throughput under varying conditions of resource availability. Marco Aldinucci, Marco Danelutto, Peter Kilpatrick |
PDP | 2 |
| 2008 | Topic 6: Grid and Cluster Computing
Marco Danelutto, Juan Touriño, Mark Baker, Rajkumar Buyya, Paraskevi Fragopoulou, Christian Pérez, Erich Schikuta |
Euro-Par | 1 |
| 2008 | Behavioural Skeletons in GCM: Autonomic Management of Grid ComponentsabstractAutonomic management can be used to improve the QoS provided by parallel/distributed applications. We discuss behavioural skeletons introduced in earlier work: rather than relying on programmer ability to design "from scratch" efficient autonomic policies, we encapsulate general autonomic controller features into algorithmic skeletons. Then we leave to the programmer the duty of specifying the parameters needed to specialise the skeletons to the needs of the particular application at hand. This results in the programmer having the ability to fast prototype and tune distributed/parallel applications with non-trivial autonomic management capabilities. We discuss how behavioural skeletons have been implemented in the framework of GCM (the grid component model developed within the CoreGRID NoE and currently being implemented within the GridCOMP STREP project). We present results evaluating the overhead introduced by autonomic management activities as well as the overall behaviour of the skeletons. We also present results achieved with a long running application subject to autonomic management and dynamically adapting to changing features of the target architecture. Overall the results demonstrate both the feasibility of implementing autonomic control via behavioural skeletons and the effectiveness of our sample behavioural skeletons in managing the "functional replication" pattern(s). Marco Aldinucci, Sonia Campa, Marco Danelutto, Marco Vanneschi, Peter Kilpatrick, Patrizio Dazzi, Domenico Laforenza, Nicola Tonellotto |
PDP | 3 |
| 2008 | Securing skeletal systems with limited performance penalty: The muskel
Marco Aldinucci, Marco Danelutto |
J. Syst. Archit. | 2 |
| 2007 | Management in Distributed Systems: A Semi-formal Approach
Marco Aldinucci, Marco Danelutto, Peter Kilpatrick |
Euro-Par | 2 |
| 2007 | The cost of security in skeletal systemsabstractSkeletal systems exploit algorithmical skeletons technology to provide the user very high level, efficient parallel programming environments. They have been recently demonstrated to be suitable for highly distributed architectures, such as workstation clusters, networks and grids. However, when using skeletal system for grid programming care must be taken to secure data and code transfers across non-dedicated, non-secure network links. In this work we take into account the cost of security introduction in muskel, a Java based skeletal system exploiting macro data flow implementation technology. We consider the adoption of mechanisms that allow securing all the communications taking place between remote, unreliable nodes and we evaluate the cost of such mechanisms. In particular, we consider the implications on the computational grains needed to scale secure and insecure skeletal computations. Marco Aldinucci, Marco Danelutto |
PDP | 2 |
| 2007 | A Performance Model for Stream-based ComputationsabstractComponent-based grid applications have complex deployment models. Performance-sensitive decisions should be taken by automatic tools. Such tools must match developer knowledge on component performance with QoS requirements on the applications, in order to find deployment plans that satisfy a service level agreement (SLA). This paper presents a steady state performance model that can be employed to reason about and automatically map stream-based component programs Marco Danelutto, Marco Vanneschi, Corrado Zoccolo, Domenico Laforenza, Nicola Tonellotto |
PDP | 1 |
| 2007 | Skeleton-based parallel programming: Functional and parallel semantics in a single shot
Marco Aldinucci, Marco Danelutto |
Comput. Lang. Syst. Struct. | 2 |
| 2006 | Autonomic QoS in ASSIST Grid-Aware ComponentsabstractCurrent grid-aware applications are developed on existing software infrastructures, such as Globus, by developers who are experts on grid software implementation. Although many useful applications have been produced this way, this approach may hardly support the additional complexity to quality of service (QoS) control in real application. We describe the ASSIST programming environment, the prototype of parallel programming environment currently under development at our group, as a suitable basis to capture all the desired features for QoS control for the grid. Grid applications, built as compositions of ASSIST components, are supported by an innovative grid abstract machine, which includes essential abstractions of standard middleware services and a hierarchical application manager, which may be considered as an early prototype of autonomic manager. Marco Aldinucci, Marco Danelutto, Marco Vanneschi |
PDP | 2 |
| 2006 | An Alternative Implementation Schema for ASSIST parmodabstractASSIST is a structured parallel programming environment targeting networks/clusters of workstations and grids. It introduced the parmod parallel construct, supporting a variety of parallelism exploitation patterns, including classical ones. The original implementation of parmod relies on static assignment of parallel activities to the processing elements at hand. In this work, we discuss an alternative implementation of the parmod construct that implements completely dynamic assignment of parallel activities to the processing elements. We show that the new implementation introduces very limited overhead in case of regular computations, whereas it performs much better than the original one in case of irregular applications. The whole implementation of parmod is available as a C++/MPI library. Marco Danelutto, C. Migliore, C. Pantaleo |
PDP | 1 |
| 2006 | Algorithmic skeletons meeting grids
Marco Danelutto, Marco Aldinucci |
Parallel Comput. | 1 |
| 2005 | Topic 9 - Parallel Programming: Models, Methods and Languages
Marco Danelutto, Denis Caromel, Duane Szafron, Fernando M. A. Silva |
Euro-Par | 1 |
| 2004 | A Framework for Orthogonal Data and Control Parallelism Exploitation
Sonia Campa, Marco Danelutto |
ICCSA (2) | 2 |
| 2003 | ASSIST Demo: A High Level, High Performance Portable, Structured Parallel Programming Environment at Work
Marco Aldinucci, Sonia Campa, Pierpaolo Ciullo, Massimo Coppola, Marco Danelutto, Paolo Pesciullesi, Roberto Ravazzolo, Massimo Torquati, Marco Vanneschi, Corrado Zoccolo |
Euro-Par | 5 |
| 2003 | Topic Introduction
José C. Cunha, Marco Danelutto, Peter H. Welch |
Euro-Par | 2 |
| 2003 | An advanced environment supporting structured parallel programming in Java
Marco Aldinucci, Marco Danelutto, P. Teti |
Future Gener. Comput. Syst. | 2 |
| 2003 | HPC the easy way: new technologies for high performance application development and deployment
Marco Danelutto |
J. Syst. Archit. | 1 |
| 2002 | Advanced environments for parallel and distributed computing
Pasqua D'Ambra, Marco Danelutto, Daniela di Serafino |
Parallel Comput. | 2 |
| 2002 | Advanced environments for parallel and distributed applications: a view of current status
Pasqua D'Ambra, Marco Danelutto, Daniela di Serafino, Marco Lapegna |
Parallel Comput. | 2 |
| 2000 | SKElib : Parallel Programming with Skeletons in C
Marco Danelutto, Massimiliano Stigliani |
Euro-Par | 1 |
| 1999 | SkIE: A heterogeneous environment for HPC applications
Bruno Bacci, Marco Danelutto, Susanna Pelagatti, Marco Vanneschi |
Parallel Comput. | 2 |
| 1998 | Optimizing Data-Parallel Programs Using the BSP Cost Model
David B. Skillicorn, Marco Danelutto, Susanna Pelagatti, Andrea Zavanella |
Euro-Par | 2 |
| 1997 | Skeletons for Data Parallelism in p3l
Marco Danelutto, Fabrizio Pasqualetti, Susanna Pelagatti |
Euro-Par | 1 |
| 1995 | P3 L: A structured high-level parallel language, and its structured supportabstractAbstract The paper presents a parallel programming methodology that ensures easy programming, efficiency and portability of programs to different machines belonging to the class of the general‐purpose, distributed‐memory, MIMD architectures. The methodology is based on the definition of a new, high‐level, explicitly parallel language, called P3L, and of a set of static tools that automatically adapt the program features for each target architecture. P3L does not require programmers to specify process activations, the actual parallelism degree, scheduling, or interprocess communications, i.e. all those features that need to be adjusted to harness each specific target machine. Parallelism is, on the other hand, expressed in a structured and qualitative way, by hierarchical composition of a restricted set of language constructs, corresponding to those forms of parallelism that are frequently encountered in parallel applications, and that can be efficiently implemented. The efficient portability of P3L applications is guaranteed by the compiler along with the novel structure of the support. The compiler automatically adapts the program features for each specific architecture, using the costs (in terms of performance) of the low‐level mechanisms exported by the architecture itself. In our methodology, these costs, along with other features of the architecture, are viewed through an abstract machine, whose interface is used by the compiler to produce the final object code. Bruno Bacci, Marco Danelutto, Salvatore Orlando 0001, Susanna Pelagatti, Marco Vanneschi |
Concurr. Pract. Exp. | 2 |
| 1992 | A methodology for the development and the support of massively parallel programs
Marco Danelutto, Roberto Di Meglio, Salvatore Orlando 0001, Susanna Pelagatti, Marco Vanneschi |
Future Gener. Comput. Syst. | 1 |
| 1992 | Implementation of a synchronous communication in a loosely coupled system: A correctness proof
Andrea Masini, Marco Danelutto |
Future Gener. Comput. Syst. | 2 |
| 1991 | Pisa parallel processing project on general-purpose highly-parallel computersabstractA methodology is presented which is aimed at the development of efficient and portable software for general-purpose highly-parallel computers. The methodology has two major components: a programming language that allows the programmer to express the parallelism of an application at a high level and an abstract model of parallel computers that allows programs written for it to be mapped efficiently to different multiple instruction-multiple data (MIMD) parallel computers.> Fabrizio Baiardi, Marco Danelutto, Roberto Di Meglio, Mehdi Jazayeri, Michael Mackey, Susanna Pelagatti, Fabrizio Petrini, Timothy S. Sullivan, Marco Vanneschi |
COMPSAC | 2 |
| 1990 | Design and Distributed Implementation of the Parallel Logic Language Shared PrologabstractThe parallel logic language Shared Prolog embeds Prolog as its sequential component. A program is Shared Prolog is composed of a set of logic agents, i.e. Prolog programs, that communicate associatively via a shared workspace called blackboard. Vincenzo Ambriola, Paolo Ciancarini, Marco Danelutto |
PPoPP | 3 |