VLDB 2026 Research / reviewers in the wild / expert
Massimo Torquati
dblp:97/4161
· DBLP profile ↗
64ranked-venue papers
3as first author
21since 2021 · last 2026
0000-0001-6323-3459ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 38 · 2 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Streaming I/O for scientific workflow engine accelerationabstractScientific workflows are increasingly characterized by complex task dependencies and large-scale data exchanges, which place significant pressure on the input/output (I/O) systems of traditional Workflow Engines (WFEs). These challenges are particularly evident in data-intensive and real-time processing contexts, where conventional disk-based I/O mechanisms often become performance bottlenecks. This paper presents an approach to enhancing the DAGonStar scientific workflow engine by integrating CAPIO, a middleware designed to support memory-based streaming I/O. The integration combines DAGonStar’s orchestration capabilities with CAPIO’s efficient data handling to better support workflows operating on continuous or large-scale datasets. We describe the architectural modifications introduced to enable this collaboration and provide an analysis of the resulting system. The proposed solution aims to improve the responsiveness and flexibility of scientific workflows by streamlining data transfers and simplifying task coordination. This work contributes to the evolution of workflow systems toward more efficient and scalable models for scientific computing. • Integration of memory-based streaming I/O into scientific workflow engines. • Automated generation of synchronization rules through workflow dependency analysis of DAGonStar. • Enhanced pipeline’s tasks execution efficiency via system call interception through the usage of CAPIO. • Benchmark evaluation showing up to 33% reduction in execution time with DAGonCAPIO. • Support for both local batch and SLURM-based distributed executions. Simone Perrotta, Ciro Giuseppe De Vita, Gennaro Mellone, Marco Edoardo Santimaria, Massimo Torquati, Francisco Javier García Blas, Raffaele Montella |
Future Gener. Comput. Syst. | 5 |
| 2026 | Dynamic transparent streaming in file-based workflows with CAPIO
Marco Edoardo Santimaria, Iacopo Colonnelli, Barbara Cantalupo, Massimo Torquati, Doriana Medic, Nicola Tuccari, Eva Sciacca, Marco Aldinucci |
Future Gener. Comput. Syst. | 4 |
| 2025 | Towards RISC-V-based HPC: The Italian Pathfinding Activities in the DARE-SGA1 ProjectabstractThe European Union’s efforts towards technological sovereignty in High-Performance Computing are driving research and development of RISC-V-based supercomputers. The DARE SGA1 project, in particular, aims to develop chips designed and owned by Europeans. This paper introduces the Italian contribution to DARE SGA1 regarding pathfinding activities toward future RISC-V-based accelerator designs, reliability improvements, system software, and AI and Quantum Chemistry applications. Giovanni Agosta, Marco Aldinucci, Andrea Bartolini, Laura Bellentani, Andrea Biagioni, Daniele Cesarini, Carlotta Chiarini, Iacopo Colonnelli, Pietro Delugas, Lev Denisov, Ottorino Frezza, Marco Grangetto, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Andrea Maslov, Mauro Olivieri, Pierpaolo Perticaroli, Luca Pontisso, Cristian Rossi, Davide Rossi 0001, Sergio Saponara, Antonio Sciarappa, Francesco Simula, Matteo Sonza Reorda, Massimo Torquati, Piero Vicini |
DSD | 26 |
| 2025 | Directed Acyclic Graph on Cross-Application Programmable I/O: Adding streaming flavour to scientific workflowsabstractThis paper introduces DAGonCAPIO, a workflow framework that integrates the DAGonStar engine with the CAPIO middleware to enable I/O streaming in scientific workflows. DAGonStar utilizes a Directed Acyclic Graph (DAG) to orchestrate tasks. At the same time, CAPIO enables downstream tasks to process data as soon as partial outputs become available, without requiring modifications to application code. This integration reduces delays from traditional file-based communication. The system uses the workflow:// schema to define data dependencies and generate CAPIO coordination scripts. DAGonStar was modified to support early scratch directory naming and decoupled task execution. Experiments with a WRF-based weather forecasting workflow on a 256-core cluster demonstrate that DAGonCAPIO reduces time-to-first-result by up to 4171 seconds, achieving a nearly 10x speedup. Simone Perrotta, Marco Edoardo Santimaria, Ciro Giuseppe De Vita, Massimo Torquati, Diana Di Luccio, Pasquale Corvino, Antonella Pirozzi, Raffaele Montella |
eScience | 4 |
| 2025 | FastFlow-Python: Parallel Building Blocks in Python Through FastFlow IntegrationabstractWe present FastFlow-Python, a framework that brings parallelism to Python for stream-processing applications. FastFlow-Python enables developers to build high-throughput, low-latency data-flow networks by instantiating high-level, ready-to-use parallel building blocks. Built on the C++ FastFlow library, it leverages Python bindings via the C/Python API to efficiently manage parallel execution using both subinterpreters and multiprocessing, all abstracted by the framework. We demonstrate the performance benefits of FastFlow-Python through a comparative analysis with a pure Python stream-processing implementation, highlighting its effectiveness in overcoming the limitations imposed by the Global Interpreter Lock (GIL). Experimental results show almost linear scalability when increasing the number of workers. Matteo Della Bartola, Jacopo Massa, Patrizio Dazzi, Massimo Torquati |
IC2E | 4 |
| 2025 | Scalable compute continuumabstractThe Compute Continuum paradigm addresses the challenges of heterogeneous and dynamic computing resources, facilitating distributed application execution while enhancing data locality, performance, availability, adaptability, and energy efficiency. By integrating IoT, edge, and cloud resources into a cohesive continuum, applications can operate closer to data sources and end users. This approach supports refined adaptation strategies tailored to specific infrastructure components, enabling reduced latency, optimized bandwidth use, and improved privacy. To fully realize the Compute Continuum’s potential, autonomous and proactive management is essential, leveraging interdisciplinary methods from optimization theory, control theory, machine learning, and artificial intelligence. This special issue highlights advancements in three key areas: resource characterization and scheduling, middleware for application deployment and reconfiguration, and applications in the Compute Continuum. These contributions highlight innovative solutions for resource optimization, dynamic management, and real-world implementations, showcasing the potential of the Compute Continuum to revolutionize distributed computing across diverse domains. Valeria Cardellini, Patrizio Dazzi, Gabriele Mencagli, Matteo Nardelli 0001, Massimo Torquati |
Future Gener. Comput. Syst. | 5 |
| 2024 | Structuring the Continuum
Marco Danelutto, Patrizio Dazzi, Massimo Torquati |
AINA (5) | 3 |
| 2024 | General-purpose data stream processing on heterogeneous architectures with WindFlowabstractMany emerging applications analyze data streams by running graphs of communicating tasks called operators. To develop and deploy such applications, Stream Processing Systems (SPSs) like Apache Storm and Flink have been made available to researchers and practitioners. They exhibit imperative or declarative programming interfaces to develop operators running arbitrary algorithms working on structured or unstructured data streams. In this context, the interest in leveraging hardware acceleration with GPUs has become more pronounced in high-throughput use cases. Unfortunately, GPU acceleration has been studied for relational operators working on structured streams only, while non-relational operators have often been overlooked. This paper presents WindFlow, a library supporting the seamless GPU offloading of general partitioned-stateful operators, extending the range of operators that benefit from hardware acceleration. Its design provides high throughput still exposing a high-level API to users compared with the raw utilization of GPUs in Apache Flink. Gabriele Mencagli, Massimo Torquati, Dalvan Griebler, Alessandra Fais, Marco Danelutto |
J. Parallel Distributed Comput. | 2 |
| 2024 | Analyzing FOSS license usage in publicly available software at scale via the SWH-analytics frameworkabstractAbstract The Software Heritage (SWH) dataset represents an invaluable source of open-source code as it aims to collect, preserve, and share all publicly available software in source code form ever produced by humankind. Although designed to archive deduplicated small files thanks to the use of a Merkle tree as the underlying data structure, querying the SWH dataset presents challenges due to the nature of these structures, which organize content based on hash values rather than any locality principle. The magnitude of the repository, coupled with the resource-intensive nature of the download process, highlights the need for specialized infrastructure and computational resources to effectively handle and study the extensive dataset housed within SWH. Currently, there is a lack of infrastructures specifically tailored for running analytics on the SWH dataset, leaving users to handle these issues manually. To address these challenges, we implemented the SWH-Analytics (SWHA) framework, a development environment that transparently runs custom analytic applications on publicly available software data preserved over time by SWH. Specifically, this work shows how SWHA can be effectively exploited to study usage patterns of free and open-source software licenses, highlighting the need to improve license literacy among developers. Alessia Antelmi, Massimo Torquati, Giacomo Corridori, Daniele Gregori, Francesco Polzella, Gianmarco Spinatelli, Marco Aldinucci |
J. Supercomput. | 2 |
| 2024 | Enhancing self-adaptation for efficient decision-making at run-time in streaming applications on multicoresabstractAbstract Parallel computing is very important to accelerate the performance of computing applications. Moreover, parallel applications are expected to continue executing in more dynamic environments and react to changing conditions. In this context, applying self-adaptation is a potential solution to achieve a higher level of autonomic abstractions and runtime responsiveness. In our research, we aim to explore and assess the possible abstractions attainable through the transparent management of parallel executions by self-adaptation. Our primary objectives are to expand the adaptation space to better reflect real-world applications and assess the potential for self-adaptation to enhance efficiency. We provide the following scientific contributions: (I) A conceptual framework to improve the designing of self-adaptation; (II) A new decision-making strategy for applications with multiple parallel stages; (III) A comprehensive evaluation of the proposed decision-making strategy compared to the state-of-the-art. The results demonstrate that the proposed conceptual framework can help design and implement self-adaptive strategies that are more modular and reusable. The proposed decision-making strategy provides significant gains in accuracy compared to the state-of-the-art, increasing the parallel applications’ performance and efficiency. Adriano Vogel, Marco Danelutto, Massimo Torquati, Dalvan Griebler, Luiz Gustavo Fernandes |
J. Supercomput. | 3 |
| 2023 | Adaptive multi-tier intelligent data manager for ExascaleabstractThe main objective of the ADMIRE project1 is the creation of an active I/O stack that dynamically adjusts computation and storage requirements through intelligent global coordination, the elasticity of computation and I/O, and the scheduling of storage resources along all levels of the storage hierarchy, while offering quality-of-service (QoS), energy efficiency, and resilience for accessing extremely large data sets in very heterogeneous computing and storage environments. We have developed a framework prototype that is able to dynamically adjust computation and storage requirements through intelligent global coordination, separated control, and data paths, the malleability of computation and I/O, the scheduling of storage resources along all levels of the storage hierarchy, and scalable monitoring techniques. The leading idea in ADMIRE is to co-design applications with ad-hoc storage systems that can be deployed with the application and adapt their computing and I/O behaviour on runtime, using malleability techniques, to increase the performance of applications and the throughput of the applications. Jesús Carretero 0001, Francisco Javier García Blas, Marco Aldinucci, Jean-Baptiste Besnard, Jean-Thomas Acquaviva, André Brinkmann, Marc-Andre Vef, Emmanuel Jeannot, Alberto Miranda, Ramon Nou, Morris Riedel, Massimo Torquati, Felix Wolf 0001 |
CF | 12 |
| 2023 | Accelerating Stream Processing Queries with Congestion-aware Scheduling and Real-time Linux ThreadsabstractStream Processing Engines (SPEs) have been used by companies and industries to develop queries able to extract insights from data streams. The Edge/IoT context poses additional challenges, since streaming queries need to run closer to data producers to save latency, i.e., on resource-constrained devices. Lachesis is a middleware helping Linux to schedule more efficiently threads of the SPE, which revealed useful especially for devices with limited CPU resources. Lachesis does not require any architectural change to the SPE implementation. It collects metrics from the SPE, and computes high-level priorities that are converted into hints to the Operating System to affect its actual scheduling of threads. This paper extends the initial contribution of Lachesis in two main directions: i) we optimize the policy assigning to threads a priority proportional to their actual load by accurately studying the implementation of Storm and Flink, two popular SPEs; ii) instead of restricting the OS scheduling to traditional SCHED_OTHER threads as done previously by Lachesis, we leverage the real-time capability of the modern Linux kernel. Our experimental evaluation shows that both enhancements provide important benefits compared with the previous version of Lachesis: we get +9.75% (average) throughput (+19% peak) with --27% latency on average (--40% peak). Fausto Frasca, Vincenzo Gulisano, Gabriele Mencagli, Dimitris Palyvos-Giannas, Massimo Torquati |
CF | 5 |
| 2023 | Experimenting with Emerging RISC-V Systems for Decentralised Machine LearningabstractDecentralised Machine Learning (DML) enables collaborative machine learning without centralised input data. Federated Learning (FL) and Edge Inference are examples of DML. While tools for DML (especially FL) are starting to flourish, many are not flexible and portable enough to experiment with novel processors (e.g., RISC-V), non-fully connected network topologies, and asynchronous collaboration schemes. We overcome these limitations via a domain-specific language allowing us to map DML schemes to an underlying middleware, i.e. the FastFlow parallel programming library. We experiment with it by generating different working DML schemes on x86-64 and ARM platforms and an emerging RISC-V one. We characterise the performance and energy efficiency of the presented schemes and systems. As a byproduct, we introduce a RISC-V porting of the PyTorch framework, the first publicly available to our knowledge. Gianluca Mittone, Nicolò Tonci, Robert Birke, Iacopo Colonnelli, Doriana Medic, Andrea Bartolini, Roberto Esposito, Emanuele Parisi, Francesco Beneventi, Mirko Polato, Massimo Torquati, Luca Benini, Marco Aldinucci |
CF | 11 |
| 2023 | A Proposal for a Continuum-aware Programming Model: From Workflows to Services Autonomously Interacting in the Compute ContinuumabstractThis paper proposes a continuum-aware programming model enabling the execution of application workflows across the compute continuum: cloud, fog and edge resources. It simplifies the management of heterogeneous nodes while alleviating the burden of programmers and unleashing innovation. This model optimizes the continuum through advanced development experiences by transforming workflows into autonomous service collaborations. It reduces complexity in positioning/interconnecting services across the continuum. A meta-model introduces high-level workflow descriptions as service networks with defined contracts and quality of service, thus enabling the deployment/management of workflows as first-class entities. It also provides automation based on policies, monitoring and heuristics. Tailored mechanisms orchestrate/manage services across the continuum, optimizing performance, cost, data protection and sustainability while managing risks. This model facilitates incremental development with visibility of design impacts and seamless evolution of applications and infrastructures. In this work, we explore this new computing paradigm showing how it can trigger the development of a new generation of tools to support the compute continuum progress. Marco Aldinucci, Robert Birke, Antonio Brogi, Emanuele Carlini 0001, Massimo Coppola, Marco Danelutto, Patrizio Dazzi, Luca Ferrucci, Stefano Forti 0002, Hanna Kavalionak, Gabriele Mencagli, Matteo Mordacchini, Marcelo Pasin, Federica Paganelli, Massimo Torquati |
COMPSAC | 15 |
| 2023 | CAPIO: a Middleware for Transparent I/O Streaming in Data- Intensive WorkflowsabstractWith the increasing amount of digital data available for analysis and simulation, the class of I/O-intensive HPC workflows is fated to quickly expand, further exacerbating the performance gap between computing, memory, and storage technologies. This paper introduces CAPIO (Cross-Application Programmable I/O), a middleware capable of injecting I/O streaming capabilities into file-based workflows, improving the computation- I/O overlap without the need to change the application code. The contribution is twofold: 1) at design time, a new I/O coordination language allows users to annotate workflow data dependencies with synchronization semantics; 2) at run time, a user-space middleware automatically and transparently to the user turns a workflow batch execution into a streaming execution according to the semantics expressed in the configuration file. CAPIO has been tested on synthetic benchmarks simulating typical workflow I/O patterns and two real-world workflows. Experiments show that CAPIO reduces the execution time by 10% to 66% for data-intensive workflows that use the file system as a communication medium. Alberto Riccardo Martinelli, Massimo Torquati, Marco Aldinucci, Iacopo Colonnelli, Barbara Cantalupo |
HiPC | 2 |
| 2021 | The Italian research on HPC key technologies across EuroHPCabstractHigh-Performance Computing (HPC) is one of the strategic priorities for research and innovation worldwide due to its relevance for industrial and scientific applications. We envision HPC as composed of three pillars: infrastructures, applications, and key technologies and tools. While infrastructures are by construction centralized in large-scale HPC centers, and applications are generally within the purview of domain-specific organizations, key technologies fall in an intermediate case where coordination is needed, but design and development are often decentralized. A large group of Italian researchers has started a dedicated laboratory within the National Interuniversity Consortium for Informatics (CINI) to address this challenge. The laboratory, albeit young, has managed to succeed in its first attempts to propose a coordinated approach to HPC research within the EuroHPC Joint Undertaking, participating in the calls 2019--20 to five successful proposals for an aggregate total cost of 95M€. In this paper, we outline the working group's scope and goals and provide an overview of the five funded projects, which become fully operational in March 2021, and cover a selection of key technologies provided by the working group partners, highlighting their usage development within the projects. Marco Aldinucci, Giovanni Agosta, Antonio Andreini, Claudio A. Ardagna, Andrea Bartolini, Alessandro Cilardo, Biagio Cosenza, Marco Danelutto, Roberto Esposito, William Fornaciari, Roberto Giorgi, Davide Lengani, Raffaele Montella, Mauro Olivieri, Sergio Saponara, Daniele Simoni, Massimo Torquati |
CF | 17 |
| 2021 | TEXTAROSSA: Towards EXtreme scale Technologies and Accelerators for euROhpc hw/Sw Supercomputing Applications for exascaleabstractTo achieve high performance and high energy efficiency on near-future exascale computing systems, three key technology gaps needs to be bridged. These gaps include: energy efficiency and thermal control; extreme computation efficiency via HW acceleration and new arithmetics; methods and tools for seamless integration of reconfigurable accelerators in heterogeneous HPC multi-node platforms. TEXTAROSSA aims at tackling this gap through a co-design approach to heterogeneous HPC solutions, supported by the integration and extension of HW and SW IPs, programming models and tools derived from European research. Giovanni Agosta, Daniele Cattaneo 0002, William Fornaciari, Andrea Galimberti, Giuseppe Massari, Federico Reghenzani, Federico Terraneo, Davide Zoni, Carlo Brandolese, Massimo Celino, Francesco Iannone, Paolo Palazzari, Giuseppe Zummo, Massimo Bernaschi, Pasqua D'Ambra, Sergio Saponara, Marco Danelutto, Massimo Torquati, Marco Aldinucci, Yasir Arfat, Barbara Cantalupo, Iacopo Colonnelli, Roberto Esposito, Alberto Riccardo Martinelli, Gianluca Mittone, Olivier Beaumont, Bérenger Bramas, Lionel Eyraud-Dubois, Brice Goglin, Abdou Guermouche, Raymond Namyst, Samuel Thibault, Antonio Filgueras, Miquel Vidal, Carlos Álvarez 0001, Xavier Martorell, Ariel Oleksiak, Michal Kulczewski, Alessandro Lonardo, Piero Vicini, Francesca Lo Cicero, Francesco Simula, Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Pier Stanislao Paolucci, Matteo Turisini, Francesco Giacomini, Tommaso Boccali, Simone Montangero, Roberto Ammendola |
DSD | 18 |
| 2021 | The 4th International Workshop on Autonomic Solutions for Parallel and Distributed Data Stream Processing (Auto-DaSP 2021)abstractThe organizers of the 4th International Workshop on Autonomic Solutions for Parallel and Distributed Data Stream Processing (Auto-DaSP 2021) are delighted to welcome you to the workshop proceedings as part of the ICPE 2021 conference companion. Valeria Cardellini, Gabriele Mencagli, Massimo Torquati |
ICPE | 3 |
| 2021 | Latency-aware adaptive micro-batching techniques for streamed data compression on graphics processing unitsabstractSummary Stream processing is a parallel paradigm used in many application domains. With the advance of graphics processing units (GPUs), their usage in stream processing applications has increased as well. The efficient utilization of GPU accelerators in streaming scenarios requires to batch input elements in microbatches, whose computation is offloaded on the GPU leveraging data parallelism within the same batch of data. Since data elements are continuously received based on the input speed, the bigger the microbatch size the higher the latency to completely buffer it and to start the processing on the device. Unfortunately, stream processing applications often have strict latency requirements that need to find the best size of the microbatches and to adapt it dynamically based on the workload conditions as well as according to the characteristics of the underlying device and network. In this work, we aim at implementing latency‐aware adaptive microbatching techniques and algorithms for streaming compression applications targeting GPUs. The evaluation is conducted using the Lempel‐Ziv‐Storer‐Szymanski compression application considering different input workloads. As a general result of our work, we noticed that algorithms with elastic adaptation factors respond better for stable workloads, while algorithms with narrower targets respond better for highly unbalanced workloads. Charles Michael Stein, Dinei A. Rockenbach, Dalvan Griebler, Massimo Torquati, Gabriele Mencagli, Marco Danelutto, Luiz Gustavo Fernandes |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | The NAS Parallel Benchmarks for evaluating C++ parallel programming frameworks on shared-memory architectures
Junior Loff, Dalvan Griebler, Gabriele Mencagli, Gabriell Alves de Araujo, Massimo Torquati, Marco Danelutto, Luiz Gustavo Fernandes |
Future Gener. Comput. Syst. | 5 |
| 2021 | WindFlow: High-Speed Continuous Stream Processing With Parallel Building BlocksabstractNowadays, we are witnessing the diffusion of Stream Processing Systems (SPSs) able to analyze data streams in near realtime. Traditional SPSs likeStormandFlinktarget distributed clusters and adopt thecontinuous streaming model, where inputs are processed as soon as they are available while outputs are continuously emitted. Recently, there has been a great focus on SPSs for scale-up machines. Some of them (e.g.,BriskStream) still use the continuous model to achieve low latency. Others optimize throughput with batching approaches that are, however, often inadequate to minimize latency for live-streaming applications. Our contribution is to show a novel software engineering approach to design the runtime system of SPSs targeting multicores, with the aim of providing a uniform solution able to optimize throughput and latency. The approach has a formal nature based on the assembly of components calledbuilding blocks, whose composition allows optimizations to be easily expressed in a compositional manner. We use this methodology to build a new SPS calledWindFlow. Our evaluation showcases the benefits ofWindFlow: it provides lower latency than SPSs for continuous streaming, and can be configured to optimize throughput, to perform similarly and even better than batch-based scale-up SPSs. Gabriele Mencagli, Massimo Torquati, Andrea Cardaci, Alessandra Fais, Luca Rinaldi, Marco Danelutto |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | Data stream processing in HPC systems: New frameworks and architectures for high-frequency streaming
Marco Aldinucci, Valeria Cardellini, Gabriele Mencagli, Massimo Torquati |
Parallel Comput. | 4 |
| 2020 | Challenging the abstraction penalty in parallel patterns libraries
José Daniel García, David del Rio Astorga, Marco Aldinucci, Fabio Tordini, Marco Danelutto, Gabriele Mencagli, Massimo Torquati |
J. Supercomput. | 7 |
| 2019 | Accelerating Actor-Based Applications with Parallel PatternsabstractParallel programmers mandate high-level parallel programming tools allowing to reduce the effort of the efficient parallelization of their applications. Parallel programming leveraging parallel patterns has recently received renovated attention thanks to their clear functional and parallel semantics. In this work, we propose a synergy between the well-known Actors-based programming model and the pattern-based parallelization methodology. We present our preliminary results in that direction, discussing and assessing the implementation of the Map parallel pattern by using an Actor-based software accelerator abstraction that seamlessly integrates within the C++ Actor Framework (CAF). The results obtained on the Intel Xeon Phi KNL platform demonstrate good performance figures achieved with negligible programming efforts. Luca Rinaldi, Massimo Torquati, Gabriele Mencagli, Marco Danelutto, Tullio Menga |
PDP | 2 |
| 2019 | Power-aware pipelining with automatic concurrency controlabstractSummary Continuous streaming computations are usually composed of different modules, exchanging data through shared message queues. The selection of the algorithm used to access such queues (ie, theconcurrency control) is a critical aspect both for performance and power consumption. In this paper, we describe the design of automatic concurrency control algorithm for implementing power‐efficient communications on shared‐memory multicores. The algorithm automatically switches between nonblocking and blocking concurrency protocols, getting the best from the two worlds, ie, obtaining the same throughput offered by the nonblocking implementation and the same power efficiency of the blocking concurrency protocol. We demonstrate the effectiveness of our approach using two micro‐benchmarks and two real streaming applications. Massimo Torquati, Daniele De Sensi, Gabriele Mencagli, Marco Aldinucci, Marco Danelutto |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | New Landscapes of the Data Stream Processing in the era of Fog Computing
Valeria Cardellini, Gabriele Mencagli, Domenico Talia, Massimo Torquati |
Future Gener. Comput. Syst. | 4 |
| 2019 | On dynamic memory allocation in sliding-window parallel patterns for streaming analytics
Massimo Torquati, Gabriele Mencagli, Maurizio Drocco, Marco Aldinucci, Tiziano De Matteis, Marco Danelutto |
J. Supercomput. | 1 |
| 2018 | Increasing Efficiency in Parallel Programming TeachingabstractThe ability to teach parallel programming principles and techniques is becoming fundamental to prepare a new generation of programmers able to master the pervasive parallelism made available by hardware vendors. Classical parallel programming courses leverage either low-level programming frameworks (e.g. those based on Pthreads) or higher level frameworks such as OpenMP or MPI. We discuss our teaching experience within the Master in "Computer Science and networking" where parallel programming is taught leveraging structured parallel programming principles and frameworks. The paper summarizes the results achieved in eight years of experience and shows how the adoption of a structured parallel programming approach improves the efficiency of the teaching process. Marco Danelutto, Massimo Torquati |
PDP | 2 |
| 2018 | Reducing Message Latency and CPU Utilization in the CAF Actor FrameworkabstractIn this work, we consider the C++ Actor Framework (CAF), a recent proposal that revamped the interest in building concurrent and distributed applications using the actor programming model in C++. CAF has been optimized for high-throughput computing, whereas message latency between actors is greatly influenced by the message data rate: at low and moderate rates the latency is higher than at high data rates. To this end, we propose a modification of the polling strategies in the work-stealing CAF scheduler, which can reduce message latency at low and moderate data rates up to two orders of magnitude without compromising the overall throughput and message latency at maximum pressure. The technique proposed uses a lightweight event notification protocol that is general enough to be used used to optimize the runtime of other frameworks experiencing similar issues. Massimo Torquati, Tullio Menga, Tiziano De Matteis, Daniele De Sensi, Gabriele Mencagli |
PDP | 1 |
| 2018 | Elastic-PPQ: A two-level autonomic system for spatial preference query processing over dynamic data streams
Gabriele Mencagli, Massimo Torquati, Marco Danelutto |
Future Gener. Comput. Syst. | 2 |
| 2018 | Harnessing sliding-window execution semantics for parallel stream processing
Gabriele Mencagli, Massimo Torquati, Fabio Lucattini, Salvatore Cuomo, Marco Aldinucci |
J. Parallel Distributed Comput. | 2 |
| 2018 | A parallel pattern for iterative stencil + reduce
Marco Aldinucci, Marco Danelutto, Maurizio Drocco, Peter Kilpatrick, Claudia Misale, Guilherme Peretti Pezzi, Massimo Torquati |
J. Supercomput. | 7 |
| 2018 | Data stream processing via code annotations
Marco Danelutto, Tiziano De Matteis, Gabriele Mencagli, Massimo Torquati |
J. Supercomput. | 4 |
| 2017 | Evaluating Concurrency Throttling and Thread Packing on SMT MulticoresabstractPower-aware computing is gaining an increasing attention both in academic and industrial settings. The problem of guaranteeing a given QoS requirement (either in terms of performance or power consumption) can be faced by selecting and dynamically adapting the amount of physical and logical resources used by the application. In this study, we considered standard multicore platforms by taking as a reference approaches for power-aware computing two well-known dynamic reconfiguration techniques: Concurrency Throttling and Thread Packing. Furthermore, we also studied the impact of using simultaneous multithreading (e.g., Intel's HyperThreading) in both techniques. In this work, leveraging on the applications of the PARSEC benchmark suite, we evaluate these techniques by considering performance-power trade-offs, resource efficiency, predictability and required programming effort. The results show that, according to the comparison criteria, these techniques complement each other. Marco Danelutto, Tiziano De Matteis, Daniele De Sensi, Massimo Torquati |
PDP | 4 |
| 2017 | Enabling semantics to improve detection of data races and misuses of lock-free data structuresabstractSummary The rapid progress of multi/many‐core architectures has caused data‐intensive parallel applications not yet fully optimized to deliver the best performance. In the advent of concurrent programming, frameworks offering structured patterns have alleviated developers' burden adapting such applications to multithreaded architectures. While some of these patterns are implemented using synchronization primitives, others avoid them by means of lock‐free data mechanisms. However, lock‐free programming is not straightforward, ensuring an appropriate use of their interfaces can be challenging, since different memory models plus instruction reordering at compiler/processor levels can interfere in the occurrence of data races. The benefits of race detectors are formidable in this sense; however, they may emit false positives if are unaware of the underlying lock‐free structure semantics. To mitigate this issue, this paper extends ThreadSanitizer, a race detection tool, with the semantics of 2 lock‐free data structures: the single‐producer/single‐consumer and the multiple‐producer/multiple‐consumer queues. With it, we are able to drop false positives and detect potential semantic violations. The experimental evaluation, using different queue implementations on a set ofμbenchmarks and real applications, demonstrates that it is possible to reduce, on average, 60% the number of data race warnings and detect wrong uses of these structures. Manuel F. Dolz, David del Rio Astorga, Javier Fernández 0001, Massimo Torquati, José Daniel García, Félix García Carballeira, Marco Danelutto |
Concurr. Comput. Pract. Exp. | 4 |
| 2017 | Bringing Parallel Patterns Out of the Corner: The P3 ARSEC Benchmark SuiteabstractHigh-level parallel programming is an active research topic aimed at promoting parallel programming methodologies that provide the programmer with high-level abstractions to develop complex parallel software with reduced time to solution. Pattern-based parallel programming is based on a set of composable and customizable parallel patterns used as basic building blocks in parallel applications. In recent years, a considerable effort has been made in empowering this programming model with features able to overcome shortcomings of early approaches concerning flexibility and performance. In this article, we demonstrate that the approach is flexible and efficient enough by applying it on 12 out of 13 PARSEC applications. Our analysis, conducted on three different multicore architectures, demonstrates that pattern-based parallel programming has reached a good level of maturity, providing comparable results in terms of performance with respect to both other parallel programming methodologies based on pragma-based annotations (i.e., Open mp and O mp S s ) and native implementations (i.e., P threads ). Regarding the programming effort, we also demonstrate a considerable reduction in lines of code and code churn compared to P threads and comparable results with respect to other existing implementations. Daniele De Sensi, Tiziano De Matteis, Massimo Torquati, Gabriele Mencagli, Marco Danelutto |
ACM Trans. Archit. Code Optim. | 3 |
| 2017 | Parallel Continuous Preference Queries over Out-of-Order and Bursty Data StreamsabstractTechniques to handle traffic bursts and out-of-order arrivals are of paramount importance to provide real-time sensor data analytics in domains like traffic surveillance, transportation management, healthcare and security applications. In these systems the amount of raw data coming from sensors must be analyzed by continuous queries that extract value-added information used to make informed decisions in real-time. To perform this task with timing constraints, parallelism must be exploited in the query execution in order to enable the real-time processing on parallel architectures. In this paper we focus on continuous preference queries, a representative class of continuous queries for decision making, and we propose a parallel query model targeting the efficient processing over out-of-order and bursty data streams. We study how to integrate punctuation mechanisms in order to enable out-of-order processing. Then, we present advanced scheduling strategies targeting scenarios with different burstiness levels, parameterized using the index of dispersion quantity. Extensive experiments have been performed using synthetic datasets and real-world data streams obtained from an existing real-time locating system. The experimental evaluation demonstrates the efficiency of our parallel solution and its effectiveness in handling the out-of-orderness degrees and burstiness levels of real-world applications. Gabriele Mencagli, Massimo Torquati, Marco Danelutto, Tiziano De Matteis |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | Introducing Parallelism by Using REPARA C++11 AttributesabstractPatterns provide a mechanism to express parallelism at a high level of abstraction and to make easier the transformation of existing legacy applications to target parallel frameworks. That also opens a path for writing new parallel applications. In this paper we introduce the REPARA approach for expressing parallel patterns and transforming the source code to parallelism frameworks. We take advantage of C++11 attributes as a mechanism to introduce annotations and enrich semantic information on valid source code. We also present a methodology for performing transformation of source code that allows to target multiple parallel programming models. Another contribution is a rule based mechanism to transform annotated code to those specific programming models. The REPARA approach requires programmer intervention only to perform initial code annotation while providing speedups that are comparable to those obtained by manual parallelization. Marco Danelutto, José Daniel García, Luis Miguel Sánchez, Rafael Sotomayor, Massimo Torquati |
PDP | 5 |
| 2016 | PWHATSHAP: efficient haplotyping for future generation sequencingabstractBACKGROUND: Haplotype phasing is an important problem in the analysis of genomics information. Given a set of DNA fragments of an individual, it consists of determining which one of the possible alleles (alternative forms of a gene) each fragment comes from. Haplotype information is relevant to gene regulation, epigenetics, genome-wide association studies, evolutionary and population studies, and the study of mutations. Haplotyping is currently addressed as an optimisation problem aiming at solutions that minimise, for instance, error correction costs, where costs are a measure of the confidence in the accuracy of the information acquired from DNA sequencing. Solutions have typically an exponential computational complexity. WHATSHAP is a recent optimal approach which moves computational complexity from DNA fragment length to fragment overlap, i.e., coverage, and is hence of particular interest when considering sequencing technology's current trends that are producing longer fragments. RESULTS: Given the potential relevance of efficient haplotyping in several analysis pipelines, we have designed and engineered PWHATSHAP, a parallel, high-performance version of WHATSHAP. PWHATSHAP is embedded in a toolkit developed in Python and supports genomics datasets in standard file formats. Building on WHATSHAP, PWHATSHAP exhibits the same complexity exploring a number of possible solutions which is exponential in the coverage of the dataset. The parallel implementation on multi-core architectures allows for a relevant reduction of the execution time for haplotyping, while the provided results enjoy the same high accuracy as that provided by WHATSHAP, which increases with coverage. CONCLUSIONS: Due to its structure and management of the large datasets, the parallelisation of WHATSHAP posed demanding technical challenges, which have been addressed exploiting a high-level parallel programming framework. The result, PWHATSHAP, is a freely available toolkit that improves the efficiency of the analysis of genomics information. Andrea Bracciali, Marco Aldinucci, Murray Patterson, Tobias Marschall, Nadia Pisanti, Ivan Merelli, Massimo Torquati |
BMC Bioinform. | 7 |
| 2016 | A Reconfiguration Algorithm for Power-Aware Parallel ApplicationsabstractIn current computing systems, many applications require guarantees on their maximum power consumption to not exceed the available power budget. On the other hand, for some applications, it could be possible to decrease their performance, yet maintain an acceptable level, in order to reduce their power consumption. To provide such guarantees, a possible solution consists in changing the number of cores assigned to the application, their clock frequency, and the placement of application threads over the cores. However, power consumption and performance have different trends depending on the application considered and on its input. Finding a configuration of resources satisfying user requirements is, in the general case, a challenging task. In this article, we propose Nornir, an algorithm to automatically derive, without relying on historical data about previous executions, performance and power consumption models of an application in different configurations. By using these models, we are able to select a close-to-optimal configuration for the given user requirement, either performance or power consumption. The configuration of the application will be changed on-the-fly throughout the execution to adapt to workload fluctuations, external interferences, and/or application’s phase changes. We validate the algorithm by simulating it over the applications of the Parsecbenchmark suit. Then, we implement our algorithm and we analyse its accuracy and overhead over some of these applications on a real execution environment. Eventually, we compare the quality of our proposal with that of the optimal algorithm and of some state-of-the-art solutions. Daniele De Sensi, Massimo Torquati, Marco Danelutto |
ACM Trans. Archit. Code Optim. | 2 |
| 2015 | Energy Driven Adaptivity in Stream Parallel ComputationsabstractDetermining the right amount of resources needed for a given computation is a critical problem. In many cases, computing systems are configured to use an amount of resources to manage high load peaks even though this cause energy waste when the resources are not fully utilised. To avoid this problem, adaptive approaches are used to dynamically increase/decrease computational resources depending on the real needs. A different approach based on Dynamic Voltage and Frequency Scaling (DVFS) is emerging as a possible alternative solution to reduce energy consumption of idle CPUs by lowering their frequencies. In this work, we propose to tackle the problem in stream parallel computations by using both the classic adaptivity concepts and the possibility provided by modern CPUs to dynamically change their frequency. We validate our approach showing a real network application that performs Deep Packet Inspection over network traffic. We are able to manage bandwidth changing over time, guaranteeing minimal packet loss during reconfiguration and minimal energy consumption. Marco Danelutto, Daniele De Sensi, Massimo Torquati |
PDP | 3 |
| 2015 | A Green Perspective on Structured Parallel ProgrammingabstractStructured parallel programming, and in particular programming models using the algorithmic skeleton or parallel design pattern concepts, are increasingly considered to be the only viable means of supporting effective development of scalable and efficient parallel programs. Structured parallel programming models have been assessed in a number of works in the context of performance. In this paper we consider how the use of structured parallel programming models allows knowledge of the parallel patterns present to be harnessed to address both performance and energy consumption. We consider different features of structured parallel programming that may be leveraged to impact the performance/energy trade-off and we discuss a preliminary set of experiments validating our claims. Marco Danelutto, Massimo Torquati, Peter Kilpatrick |
PDP | 2 |
| 2014 | Loop Parallelism: A New Skeleton Perspective on Data Parallel PatternsabstractTraditionally, skeleton based parallel programming frameworks support data parallelism by providing the programmer with a comprehensive set of data parallel skeletons, based on different variants of map and reduce patterns. On the other side, more conventional parallel programming frameworks provide application programmers with the possibility to introduce parallelism in the execution of loops with a relatively small programming effort. In this work, we discuss a "ParallelFor" skeleton provided within the FastFlow framework and aimed at filling the usability and expressivity gap between the classical data parallel skeleton approach and the loop parallelisation facilities offered by frameworks such as OpenMP and Intel TBB. By exploiting the low run-time overhead of the FastFlow parallel skeletons and the new facilities offered by the C++11 standard, our ParallelFor skeleton succeeds to obtain comparable or better performance than both OpenMP and TBB on the Intel Phi many-core and Intel Nehalem multi-core for a set of benchmarks considered, yet requiring a comparable programming effort. Marco Danelutto, Massimo Torquati |
PDP | 2 |
| 2014 | Message Passing on InfiniBand RDMA for Parallel Run-Time SupportsabstractInfiniBand networks are commonly used in the high performance computing area. They offer RDMA-based operations that help to improve the performance of communication subsystems. In this paper, we propose a minimal message-passing communication layer providing the programmer with a point-to-point communication channel implemented by way of InfiniBand RDMA features. Differently from other libraries exploiting the InfiniBand features, such as the well-known Message Passing Interface (MPI), the proposed library is a communication layer only rather than a programming model, and can be easily used as building block for high-level parallel programming frameworks. Evaluated on micro-benchmarks, the proposed RDMA-based communication channel implementation achieves a comparable performance with highly optimised MPI/InfiniBand implementations. Eventually, the flexibility of the communication layer is evaluated by integrating it within the FastFlow parallel framework, currently supporting TCP/IP networks (via the ZeroMQ communication library). Alessandro Secco, Muhammad Irfan Uddin, Guilherme Peretti Pezzi, Massimo Torquati |
PDP | 4 |
| 2014 | Parallel stochastic systems biology in the cloudabstractThe stochastic modelling of biological systems, coupled with Monte Carlo simulation of models, is an increasingly popular technique in bioinformatics. The simulation-analysis workflow may result computationally expensive reducing the interactivity required in the model tuning. In this work, we advocate the high-level software design as a vehicle for building efficient and portable parallel simulators for the cloud. In particular, the Calculus of Wrapped Components (CWC) simulator for systems biology, which is designed according to the FastFlow pattern-based approach, is presented and discussed. Thanks to the FastFlow framework, the CWC simulator is designed as a high-level workflow that can simulate CWC models, merge simulation results and statistically analyse them in a single parallel workflow in the cloud. To improve interactivity, successive phases are pipelined in such a way that the workflow begins to output a stream of analysis results immediately after simulation is started. Performance and effectiveness of the CWC simulator are validated on the Amazon Elastic Compute Cloud. Marco Aldinucci, Massimo Torquati, Concetto Spampinato, Maurizio Drocco, Claudia Misale, Cristina Calcagno, Mario Coppo |
Briefings Bioinform. | 2 |
| 2014 | Decision tree building on multi-core using FastFlowabstractSUMMARY The whole computer hardware industry embraced the multi‐core. The extreme optimisation of sequential algorithms is then no longer sufficient to squeeze the real machine power, which can be only exploited via thread‐level parallelism. Decision tree algorithms exhibit natural concurrency that makes them suitable to be parallelised. This paper presents an in‐depth study of the parallelisation of an implementation of the C4.5 algorithm for multi‐core architectures. We characterise elapsed time lower bounds for the forms of parallelisations adopted and achieve close to optimal performance. Our implementation is based on the FastFlow parallel programming environment, and it requires minimal changes to the original sequential code. Copyright © 2013 John Wiley & Sons, Ltd. Marco Aldinucci, Salvatore Ruggieri, Massimo Torquati |
Concurr. Comput. Pract. Exp. | 3 |
| 2014 | Parallel patterns for heterogeneous CPU/GPU architectures: Structured parallelism from cluster to cloud
Sonia Campa, Marco Danelutto, Mehdi Goli 0001, Horacio González-Vélez, Alina Madalina Popescu, Massimo Torquati |
Future Gener. Comput. Syst. | 6 |
| 2013 | Towards The Deployment Of Fastflow On Distributed Virtual ArchitecturesabstractIn this paper we investigate the deployment of FastFlow applications on multi-core virtual platforms. The overhead introduced by the virtual environment has been measured using a well-known application benchmark both in the sequential and in the FastFlow parallel setting. The overhead introduced for both the sequential and the parallel executions of CPU and memory-intensive applications is in the range of 2-30%, while the execution speedup is almost preserved. Additionally, we have ported the FastFlow benchmark to a cloud-based distributed environment in which a task-intensive application has been tested and the performance compared with the corresponding run on a smaller cluster of multi-core machines without virtualisation.
From a parallel programming perspective, we have demonstrated how a unique programming framework based on the structured parallel programming paradigm can cope with very different kind of target architectures without any (or minimal) code intervention. Sonia Campa, Marco Danelutto, Massimo Torquati, Horacio González-Vélez, Alina Madalina Popescu |
ECMS | 3 |
| 2013 | Parallel Stochastic Simulators in System Biology: The Evolution of the SpeciesabstractThe stochastic simulation of biological systems is an increasingly popular technique in Bioinformatics. It is often an enlightening technique, especially for multi-stable systems which dynamics can be hardly captured with ordinary differential equations. To be effective, stochastic simulations should be supported by powerful statistical analysis tools. The simulation-analysis workflow may however result in being computationally expensive, thus compromising the interactivity required in model tuning. In this work we advocate the high-level design of simulators for stochastic systems as a vehicle for building efficient and portable parallel simulators. In particular, the Calculus of Wrapped Components (CWC) simulator, which is designed according to the FastFlow's pattern-based approach, is presented and discussed in this work. FastFlow has been extended to support also clusters of multi-cores with minimal coding effort, assessing the portability of the approach. Marco Aldinucci, Maurizio Drocco, Fabio Tordini, Mario Coppo, Massimo Torquati |
PDP | 5 |
| 2013 | Parallel Patterns for General Purpose Many-CoreabstractEfficient programming of general purpose many-core accelerators poses several challenging problems. The high number of cores available, the peculiarity of the interconnection network, and the complex memory hierarchy organization, all contribute to make efficient programming of such devices difficult. We propose to use parallel design patterns, implemented using algorithmic skeletons, to abstract and hide most of the difficulties related to the efficient programming of many-core accelerators. In particular, we discuss the porting of the FastFlow framework on the Tilera TilePro64 architecture and the results obtained running synthetic benchmarks as well as true application kernels. These results demonstrate the efficiency achieved while using patterns on the TilePro64 both to program stand-alone skeleton-based parallel applications and to accelerate existing sequential code. Daniele Buono, Marco Danelutto, Silvia Lametti, Massimo Torquati |
PDP | 4 |
| 2013 | A RISC Building Block Set for Structured Parallel ProgrammingabstractWe propose a set of building blocks (RISC-pb2l) suitable to build high-level structured parallel programming frameworks. The set is designed following a RISC approach. RISC-pb2l is architecture independent but the implementation of the different blocks may be specialized to make the best usage of the target architecture peculiarities. A number of optimizations may be designed transforming basic building blocks compositions into more efficient compositions, such that parallel application efficiency may be derived by construction rather than by debugging. Marco Danelutto, Massimo Torquati |
PDP | 2 |
| 2012 | An Efficient Unbounded Lock-Free Queue for Multi-core Systems
Marco Aldinucci, Marco Danelutto, Peter Kilpatrick, Massimiliano Meneghin, Massimo Torquati |
Euro-Par | 5 |
| 2012 | Parallel Patterns + Macro Data Flow for Multi-core ProgrammingabstractData flow techniques have been around since the early '70s when they were used in compilers for sequential languages. Shortly after their introduction they were also considered as a possible model for parallel computing, although the impact here was limited. Recently, however, data flow has been identified as a candidate for efficient implementation of various programming models on multi-core architectures. In most cases, however, the burden of determining data flow ``macro'' instructions is left to the programmer, while the compiler/run time system manages only the efficient scheduling of these instructions. We discuss a structured parallel programming approach supporting automatic compilation of programs to macro data flow and we show experimental results demonstrating the feasibility of the approach and the efficiency of the resulting ``object'' code on different classes of state-of-the-art multi-core architectures. The experimental results use different base mechanisms to implement the macro data flow run time support, from plain pthreads with condition variables to more modern and effective lock- and fence-free parallel frameworks. Experimental results comparing efficiency of the proposed approach with those achieved using other, more classical, parallel frameworks are also presented. Marco Aldinucci, L. Anardu, Marco Danelutto, Massimo Torquati, Peter Kilpatrick |
PDP | 4 |
| 2011 | Accelerating Code on Multi-cores with FastFlow
Marco Aldinucci, Marco Danelutto, Peter Kilpatrick, Massimiliano Meneghin, Massimo Torquati |
Euro-Par (2) | 5 |
| 2011 | On Designing Multicore-Aware Simulators for Biological SystemsabstractThe stochastic simulation of biological systems is an increasingly popular technique in bioinformatics. It often is an enlightening technique, which may however result in being computational expensive. We discuss the main opportunities to speed it up on multi-core platforms, which pose new challenges for parallelisation techniques. These opportunities are developed in two general families of solutions involving both the single simulation and a bulk of independent simulations (either replicas of derived from parameter sweep). Proposed solutions are tested on the parallelisation of the CWC simulator (Calculus of Wrapped Compartments) that is carried out according to proposed solutions by way of the Fast Flow programming framework making possible fast development and efficient execution on multi-cores. Marco Aldinucci, Mario Coppo, Ferruccio Damiani, Maurizio Drocco, Massimo Torquati, Angelo Troina |
PDP | 5 |
| 2010 | Resource discovery support for time-critical adaptive applicationsabstractSeveral complex and time-critical applications require the existence of novel distributed and dynamical platforms composed of a variety of fixed and mobile processing nodes and networks. Notable examples of such applications are crisis and emergency management and natural phenomenon prediction. In this scenario we need the development of applications able to adapt their behavior according to the dynamical platform conditions, such as the presence of specific classes of computing resources and the actual network availability. For these reasons such adaptive applications need to interact with a fast and reliable resource discovery support, which ensures required response times by means of an high-degree of reconfigurability and selectivity. In this paper we present an integrated approach between our programming model for distributed adaptive time-critical computations and a suitable resource discovery support. Carlo Bertolli, Daniele Buono, Gabriele Mencagli, Massimo Torquati, Marco Vanneschi, Matteo Mordacchini, Franco Maria Nardini |
IWCMC | 4 |
| 2010 | Efficient Smith-Waterman on Multi-core with FastFlowabstractShared memory multiprocessors have returned to popularity thanks to rapid spreading of commodity multi-core architectures. However, little attention has been paid to supporting effective streaming applications on these architectures. In this paper we describe FastFlow, a low-level programming framework based on lock-free queues explicitly designed to support high-level languages for streaming applications. We compare FastFlow with state-of-the-art programming frameworks such as Cilk, OpenMP, and Intel TBB. We experimentally demonstrate that FastFlow is always more efficient than them on a given real world application: the speedup of FastFlow over other solutions may be substantial for fine grain tasks, for example +35% over OpenMP, +226% over Cilk, +96% over TBB for the alignment of protein P01111 against UniProt DB using the Smith-Waterman algorithm. Marco Aldinucci, Massimiliano Meneghin, Massimo Torquati |
PDP | 3 |
| 2010 | Porting Decision Tree Algorithms to Multicore Using FastFlow
Marco Aldinucci, Salvatore Ruggieri, Massimo Torquati |
ECML/PKDD (1) | 3 |
| 2008 | The VirtuaLinux Storage Abstraction Layer for Ef?cient Virtual ClusteringabstractVirtuaLinux is a meta-distribution that enables a standard Linux distribution to support robust physical and virtualized clusters. VirtuaLinux helps in avoiding the "single point of failure" effect by means of a combination of architectural strategies, including the transparent support for disk-less and master-less cluster configuration. VirtuaLinux supports the creation and management of Virtual Clusters in seamless way: VirtuaLinux Virtual Cluster Manager enables the system administrator to create, save, restore Xen- based Virtual Clusters, and to map and dynamically re-map them onto the nodes of the physical cluster. In this paper we introduce and discuss VirtuaLinux virtualization architecture, features, and tools, and in particular, the novel disk abstraction layer, which permits the fast and space-efficient creation of Virtual Clusters. Marco Aldinucci, Massimo Torquati, Marco Vanneschi, Pierfrancesco Zuccato |
PDP | 2 |
| 2005 | Dynamic Reconfiguration of Grid-Aware Applications in ASSIST
Marco Aldinucci, Alessandro Petrocelli, Edoardo Pistoletti, Massimo Torquati, Marco Vanneschi, Luca Veraldi, Corrado Zoccolo |
Euro-Par | 4 |
| 2004 | Targeting Heterogeneous Architectures in ASSIST: Experimental Results
Marco Aldinucci, Sonia Campa, Massimo Coppola, Silvia Magini, Paolo Pesciullesi, Laura Potiti, Roberto Ravazzolo, Massimo Torquati, Corrado Zoccolo |
Euro-Par | 8 |
| 2004 | Accelerating Apache Farms Through Ad-HOC Distributed Scalable Object Repository
Marco Aldinucci, Massimo Torquati |
Euro-Par | 2 |
| 2003 | ASSIST Demo: A High Level, High Performance Portable, Structured Parallel Programming Environment at Work
Marco Aldinucci, Sonia Campa, Pierpaolo Ciullo, Massimo Coppola, Marco Danelutto, Paolo Pesciullesi, Roberto Ravazzolo, Massimo Torquati, Marco Vanneschi, Corrado Zoccolo |
Euro-Par | 8 |
| 2003 | The Implementation of ASSIST, an Environment for Parallel and Distributed Programming
Marco Aldinucci, Sonia Campa, Pierpaolo Ciullo, Massimo Coppola, Silvia Magini, Paolo Pesciullesi, Laura Potiti, Roberto Ravazzolo, Massimo Torquati, Marco Vanneschi, Corrado Zoccolo |
Euro-Par | 9 |