VLDB 2026 Research / reviewers in the wild / expert
Gabriele Mencagli
dblp:85/5541
· DBLP profile ↗
51ranked-venue papers
15as first author
20since 2021 · last 2026
0000-0002-6263-7723ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 11 first-author · 12 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 3 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enabling large-state stream processing on memory-constrained multi-core systems via key-value stores
Andrea Filippi, Gabriele Mencagli, Dalvan Griebler |
Future Gener. Comput. Syst. | 2 |
| 2026 | Leveraging cutting-edge high performance computing for large-scale applications
Claude Tadonki, Gabriele Mencagli, Leonel Sousa |
Future Gener. Comput. Syst. | 2 |
| 2026 | Scalable join operators over data streams with shared-nothing parallelismabstractStream joins are among the most computationally demanding stateful operators in stream processing. Tuples arriving from different streams must be analyzed on-the-fly to identify pairs that satisfy specific user-defined conditions. Since buffering all tuples from the input streams is infeasible due to memory constraints, stream joins are typically computed over a subset of the received tuples. This subset is often organized either by a specific time interval ( online interval joins ) or by fixed-length temporal windows with a defined slide ( window joins ). In this paper, we present various parallel patterns for stream join computation, aimed at effectively increasing overall query throughput. Our focus is on leveraging shared-nothing parallelism to provide portable parallelization strategies that can be efficiently executed on modern scale-in and scale-out Stream Processing Engines. Among the proposed patterns, the one exhibiting hybrid parallelism emerges as the most promising in terms of performance and load balancing. The experimental evaluation highlights the performance characteristics of the proposed patterns using real-world datasets and diverse key distributions, and compares them with state-of-the-art solutions, confirming the effectiveness of the parallel pattern with hybrid parallelism against the main competitors. Gabriele Mencagli, Yuriy Rymarchuk, Dalvan Griebler |
Inf. Syst. | 1 |
| 2025 | Non-Functional Properties in HPC Systems: Design Exploration of Energy, Power, and ReliabilityabstractModern HPC systems must be designed considering different parameters, which include cost, performance, and throughput, as well as non-functional properties, such as power/energy consumption and reliability. This paper describes the work performed and the results achieved by the partners of the Italian National Research Center for HPC, Big Data and Quantum Computing in the frame of the sub-project dealing with Future HPC architectures and solutions. The work in this subproject focused on advanced design and monitoring techniques for devising energy- and power-efficient, reliable parallel architectures based on open standards (e.g., RISC-V) and design space exploration techniques and tools. This paper provides a summary of the achieved results and developed products stemming from the activities of the different partners. Giovanni Agosta, Enrico Bini, Davide Baroffio, Carlo Brandolese, Michele Castrovilli, Daniele Cattaneo 0002, Daniele Cesarini, William Fornaciari, Andrea Galimberti, Alberto Garfagnini, Arsenii Gavrikov, Francesco Iannone, Marco Lapegna, Tomas Antonio López, Gabriele Magnani, Gabriele Mencagli, Cecilia Metra, Martin Omaña 0001, Filippo Palombi, Federico Reghenzani, Josie E. Rodriguez Condia, A. Serafini, Matteo Sonza Reorda, Davide Zoni, Giuseppe Zummo |
DSD | 16 |
| 2025 | Scalable compute continuumabstractThe Compute Continuum paradigm addresses the challenges of heterogeneous and dynamic computing resources, facilitating distributed application execution while enhancing data locality, performance, availability, adaptability, and energy efficiency. By integrating IoT, edge, and cloud resources into a cohesive continuum, applications can operate closer to data sources and end users. This approach supports refined adaptation strategies tailored to specific infrastructure components, enabling reduced latency, optimized bandwidth use, and improved privacy. To fully realize the Compute Continuum’s potential, autonomous and proactive management is essential, leveraging interdisciplinary methods from optimization theory, control theory, machine learning, and artificial intelligence. This special issue highlights advancements in three key areas: resource characterization and scheduling, middleware for application deployment and reconfiguration, and applications in the Compute Continuum. These contributions highlight innovative solutions for resource optimization, dynamic management, and real-world implementations, showcasing the potential of the Compute Continuum to revolutionize distributed computing across diverse domains. Valeria Cardellini, Patrizio Dazzi, Gabriele Mencagli, Matteo Nardelli 0001, Massimo Torquati |
Future Gener. Comput. Syst. | 3 |
| 2024 | General-purpose data stream processing on heterogeneous architectures with WindFlowabstractMany emerging applications analyze data streams by running graphs of communicating tasks called operators. To develop and deploy such applications, Stream Processing Systems (SPSs) like Apache Storm and Flink have been made available to researchers and practitioners. They exhibit imperative or declarative programming interfaces to develop operators running arbitrary algorithms working on structured or unstructured data streams. In this context, the interest in leveraging hardware acceleration with GPUs has become more pronounced in high-throughput use cases. Unfortunately, GPU acceleration has been studied for relational operators working on structured streams only, while non-relational operators have often been overlooked. This paper presents WindFlow, a library supporting the seamless GPU offloading of general partitioned-stateful operators, extending the range of operators that benefit from hardware acceleration. Its design provides high throughput still exposing a high-level API to users compared with the raw utilization of GPUs in Apache Flink. Gabriele Mencagli, Massimo Torquati, Dalvan Griebler, Alessandra Fais, Marco Danelutto |
J. Parallel Distributed Comput. | 1 |
| 2024 | Boosting general-purpose stream processing with reconfigurable hardwareabstractAbstract Reconfigurable devices such as field-programmable gate arrays (FPGAs) offer flexible solutions to workload acceleration with high energy efficiency. Despite such a potential advantage, they often reveal hard to program by application programmers. High-level synthesis languages have been developed to provide higher-level abstractions, allowing the developers to define the FPGA behavior using an imperative programming approach based on C/C++ languages. However, such approaches still leave the developer with the responsibility to harness the low-level optimizations required to develop efficient FPGA programs. Along this line, this paper introduces , a framework helping programmers to develop FPGA-accelerated data stream processing (DSP) applications. The approach provides a high-level Python API to develop the data-flow graph of operators, which is automatically translated into an efficient Vitis source code targeting Xilinx devices. The execution of the bitstreams implementing two benchmark applications showcases the efficiency of using FPGAs for DSP workloads. In general, provides, with a reasonable time-to-solution, higher performance compared with state-of-the-art DSP frameworks. Alberto Ottimo, Gabriele Mencagli, Marco Danelutto |
J. Supercomput. | 2 |
| 2024 | Springald: GPU-Accelerated Window-Based Aggregates Over Out-of-Order Data StreamsabstractAn increasing number of application domains require high-throughput processing to extract insights from massive data streams. The Data Stream Processing (DSP) paradigm provides formal approaches to analyze structured data streams considered as special, unbounded relations. The most used class of stateful operators in DSP are the ones running sliding-window aggregation, which continuously extracts insights from the most recent portion of the stream. This article presentsSpringald, an efficient sliding-window operator leveraging GPU devices.Springald, incorporated in theWindFlowparallel library, processes out-of-order data streams with watermarks propagation. These two features—GPU processing and out-of-orderliness—makeSpringalda novel contribution to this research area. This article describes the methodology behindSpringald, its design and implementation. We also provide an extensive experimental evaluation to understand the behavior ofSpringalddeeply, and we showcase its superior performance against state-of-the-art competitors. Gabriele Mencagli, Patrizio Dazzi, Massimo Coppola |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | Accelerating Stream Processing Queries with Congestion-aware Scheduling and Real-time Linux ThreadsabstractStream Processing Engines (SPEs) have been used by companies and industries to develop queries able to extract insights from data streams. The Edge/IoT context poses additional challenges, since streaming queries need to run closer to data producers to save latency, i.e., on resource-constrained devices. Lachesis is a middleware helping Linux to schedule more efficiently threads of the SPE, which revealed useful especially for devices with limited CPU resources. Lachesis does not require any architectural change to the SPE implementation. It collects metrics from the SPE, and computes high-level priorities that are converted into hints to the Operating System to affect its actual scheduling of threads. This paper extends the initial contribution of Lachesis in two main directions: i) we optimize the policy assigning to threads a priority proportional to their actual load by accurately studying the implementation of Storm and Flink, two popular SPEs; ii) instead of restricting the OS scheduling to traditional SCHED_OTHER threads as done previously by Lachesis, we leverage the real-time capability of the modern Linux kernel. Our experimental evaluation shows that both enhancements provide important benefits compared with the previous version of Lachesis: we get +9.75% (average) throughput (+19% peak) with --27% latency on average (--40% peak). Fausto Frasca, Vincenzo Gulisano, Gabriele Mencagli, Dimitris Palyvos-Giannas, Massimo Torquati |
CF | 3 |
| 2023 | A Proposal for a Continuum-aware Programming Model: From Workflows to Services Autonomously Interacting in the Compute ContinuumabstractThis paper proposes a continuum-aware programming model enabling the execution of application workflows across the compute continuum: cloud, fog and edge resources. It simplifies the management of heterogeneous nodes while alleviating the burden of programmers and unleashing innovation. This model optimizes the continuum through advanced development experiences by transforming workflows into autonomous service collaborations. It reduces complexity in positioning/interconnecting services across the continuum. A meta-model introduces high-level workflow descriptions as service networks with defined contracts and quality of service, thus enabling the deployment/management of workflows as first-class entities. It also provides automation based on policies, monitoring and heuristics. Tailored mechanisms orchestrate/manage services across the continuum, optimizing performance, cost, data protection and sustainability while managing risks. This model facilitates incremental development with visibility of design impacts and seamless evolution of applications and infrastructures. In this work, we explore this new computing paradigm showing how it can trigger the development of a new generation of tools to support the compute continuum progress. Marco Aldinucci, Robert Birke, Antonio Brogi, Emanuele Carlini 0001, Massimo Coppola, Marco Danelutto, Patrizio Dazzi, Luca Ferrucci, Stefano Forti 0002, Hanna Kavalionak, Gabriele Mencagli, Matteo Mordacchini, Marcelo Pasin, Federica Paganelli, Massimo Torquati |
COMPSAC | 11 |
| 2023 | FastFlow targeting FPGAsabstractWriting good code for FPGA is a challenge “per se”, but also running already existing and optimized FPGA kernels often requires writing specific “host side” code and some target hardware knowledge to achieve good performances. In this work, we describe a FastFlow extension supporting seamless off loading of tasks to FPGA, once an FPGA kernel is available. In particular, we show how kernels implemented in Vitis and running on XILINX Alveo FPGA boards may be integrated to implement “normal” parallel stages (pipeline stages, map/farm workers) in a structured parallel FastFlow computation. Experimental results are shown, demonstrating the feasibility of the approach. Marco Danelutto, Gabriele Mencagli, Alberto Ottimo, Francesco Iannone, Paolo Palazzari |
PDP | 2 |
| 2023 | FSP: a Framework for Data Stream Processing Applications targeting FPGAsabstractFPGA architectures are becoming popular because of their high performance-to-energy ratio. Nonetheless, their effective exploitation is often counterbalanced by a high programming effort, since most of the modern hardware description languages provide only low-level programming abstractions. This paper proposes FSP, a framework to productively support the development of Data Stream Processing applications on CPU+FPGA System-on-Chip devices (SoCs). By exploiting a code generation approach starting from a high-level DSL in Python, FSP generates an efficient OpenCL skeleton implementation of the parallel pipeline on FPGA and the library to be used by host programs to transfer inputs and collect results to/from the FPGA program. The experimental results showcase the effectiveness of FSP on an SoC equipped with an Intel Arria 10 FPGA by running two streaming benchmark applications. Alberto Ottimo, Gabriele Mencagli, Marco Danelutto |
PDP | 2 |
| 2022 | Towards Parallel Data Stream Processing on System-on-Chip CPU+GPU DevicesabstractData Stream Processing is a pervasive computing paradigm with a wide spectrum of applications. Traditional streaming systems exploit the processing capabilities provided by homogeneous Clusters and Clouds. Due to the transition to streaming systems suitable for IoT/Edge environments, there has been the urgent need of new streaming frameworks and tools tailored for embedded platforms, often available as System-onChips composed of a small multicore CPU and an integrated onchip GPU. Exploiting this hybrid hardware requires special care in the runtime system design. In this paper, we discuss the support provided by the WindFlow library, showing its design principles and its effectiveness on the NVIDIA Jetson Nano board. Gabriele Mencagli, Dalvan Griebler, Marco Danelutto |
PDP | 1 |
| 2021 | Lachesis: a middleware for customizing OS scheduling of stream processing queriesabstractData streaming applications in Cyber-Physical Systems enable high-throughput, low-latency transformations of raw data into value. The performance of such applications, run by Stream Processing Engines (SPEs), can be boosted through custom CPU scheduling. Previous schedulers in the literature require alterations to SPEs to control the scheduling through user-level threads. While such alterations allow for fine-grained control, they hinder the adoption of such schedulers due to the high implementation cost and potential limitations in application semantics (e.g., blocking I/O). Dimitris Palyvos-Giannas, Gabriele Mencagli, Marina Papatriantafilou, Vincenzo Gulisano |
Middleware | 2 |
| 2021 | Towards On-the-fly Self-Adaptation of Stream Parallel PatternsabstractStream processing applications compute streams of data and provide insightful results in a timely manner, where parallel computing is necessary for accelerating the application executions. Considering that these applications are becoming increasingly dynamic and long-running, a potential solution is to apply dynamic runtime changes. However, it is challenging for humans to continuously monitor and manually self-optimize the executions. In this paper, we propose self-adaptiveness of the parallel patterns used, enabling flexible on-the-fly adaptations. The proposed solution is evaluated with an existing programming framework and running experiments with a synthetic and a real-world application. The results show that the proposed solution is able to dynamically self-adapt to the most suitable parallel pattern configuration and achieve performance competitive with the best static cases. The feasibility of the proposed solution encourages future optimizations and other applicabilities. Adriano Vogel, Gabriele Mencagli, Dalvan Griebler, Marco Danelutto, Luiz Gustavo Fernandes |
PDP | 2 |
| 2021 | The 4th International Workshop on Autonomic Solutions for Parallel and Distributed Data Stream Processing (Auto-DaSP 2021)abstractThe organizers of the 4th International Workshop on Autonomic Solutions for Parallel and Distributed Data Stream Processing (Auto-DaSP 2021) are delighted to welcome you to the workshop proceedings as part of the ICPE 2021 conference companion. Valeria Cardellini, Gabriele Mencagli, Massimo Torquati |
ICPE | 2 |
| 2021 | Novel parallel processing techniques for IoT-based machine learning applicationsabstractThis special issue of Concurrency and Computation: Practice and Experience comprises four papers extending the original workshop publications accepted and presented at MPP 2019 (8th Workshop on Parallel Programming Models – Special Issue on IoT and Machine Learning), held in Rio de Janeiro in conjunction with IPDPS 2019. The papers represent interesting research ideas of parallel programming models, tools, and optimizations suited for being applied to IoT-based applications. The paper titled “Enabling Heterogeneous Ray-Tracing Acceleration in Edge/Cloud Architectures” represents a very interesting research on reconfigurable accelerators for Ray-Tracing, specialized in computing ray-triangle intersections at the network edge of a heterogeneous cloud computing environment. The authors validated their approach on the Xilinx Zynq FPGA platform. The paper titled “An Incremental Reinforcement Learning Scheduling Strategy for Data-Intensive Scientific Workflows in the Cloud” is a research work proposing a new scheduling algorithm based on Reinforcement Learning for scientific workflows on HPC distributed resources and Clouds. The paper titled “Latency-aware Adaptive Micro-Batching Techniques for Streamed Data Compression on GPUs” proposes a set of Autonomic Computing strategies for configuring the optimal parallelism degree and batch size on streaming data compression parallel applications on GPU architectures. Finally, the paper titled “Gamma – General Abstract Model for Multiset mAnipulation and Dynamic Dataflow Model: an Equivalence Study” is an interesting research study on new applications of the Gamma Model (General Abstract Model for Multiset mAnipulation) and its joint utilization with the Dataflow programming model. As Guest Editors, we would like to express our gratitude for the valuable contributions made by all authors. We would like to thank all the reviewers who helped us during the thorough review process based on several review rounds. Finally, we would like to thank all the editorial board members and all staff of Concurrency and Computation: Practice and Experience for allowing us to publish our Special Issue and for their invaluable support during the whole publication process. Cristiana Bentes, Felipe M. G. França, Leandro A. J. Marzulo, Gabriele Mencagli, Maurício L. Pilla |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | Latency-aware adaptive micro-batching techniques for streamed data compression on graphics processing unitsabstractSummary Stream processing is a parallel paradigm used in many application domains. With the advance of graphics processing units (GPUs), their usage in stream processing applications has increased as well. The efficient utilization of GPU accelerators in streaming scenarios requires to batch input elements in microbatches, whose computation is offloaded on the GPU leveraging data parallelism within the same batch of data. Since data elements are continuously received based on the input speed, the bigger the microbatch size the higher the latency to completely buffer it and to start the processing on the device. Unfortunately, stream processing applications often have strict latency requirements that need to find the best size of the microbatches and to adapt it dynamically based on the workload conditions as well as according to the characteristics of the underlying device and network. In this work, we aim at implementing latency‐aware adaptive microbatching techniques and algorithms for streaming compression applications targeting GPUs. The evaluation is conducted using the Lempel‐Ziv‐Storer‐Szymanski compression application considering different input workloads. As a general result of our work, we noticed that algorithms with elastic adaptation factors respond better for stable workloads, while algorithms with narrower targets respond better for highly unbalanced workloads. Charles Michael Stein, Dinei A. Rockenbach, Dalvan Griebler, Massimo Torquati, Gabriele Mencagli, Marco Danelutto, Luiz Gustavo Fernandes |
Concurr. Comput. Pract. Exp. | 5 |
| 2021 | The NAS Parallel Benchmarks for evaluating C++ parallel programming frameworks on shared-memory architectures
Junior Loff, Dalvan Griebler, Gabriele Mencagli, Gabriell Alves de Araujo, Massimo Torquati, Marco Danelutto, Luiz Gustavo Fernandes |
Future Gener. Comput. Syst. | 3 |
| 2021 | WindFlow: High-Speed Continuous Stream Processing With Parallel Building BlocksabstractNowadays, we are witnessing the diffusion of Stream Processing Systems (SPSs) able to analyze data streams in near realtime. Traditional SPSs likeStormandFlinktarget distributed clusters and adopt thecontinuous streaming model, where inputs are processed as soon as they are available while outputs are continuously emitted. Recently, there has been a great focus on SPSs for scale-up machines. Some of them (e.g.,BriskStream) still use the continuous model to achieve low latency. Others optimize throughput with batching approaches that are, however, often inadequate to minimize latency for live-streaming applications. Our contribution is to show a novel software engineering approach to design the runtime system of SPSs targeting multicores, with the aim of providing a uniform solution able to optimize throughput and latency. The approach has a formal nature based on the assembly of components calledbuilding blocks, whose composition allows optimizations to be easily expressed in a compositional manner. We use this methodology to build a new SPS calledWindFlow. Our evaluation showcases the benefits ofWindFlow: it provides lower latency than SPSs for continuous streaming, and can be configured to optimize throughput, to perform similarly and even better than batch-based scale-up SPSs. Gabriele Mencagli, Massimo Torquati, Andrea Cardaci, Alessandra Fais, Luca Rinaldi, Marco Danelutto |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | Data stream processing in HPC systems: New frameworks and architectures for high-frequency streaming
Marco Aldinucci, Valeria Cardellini, Gabriele Mencagli, Massimo Torquati |
Parallel Comput. | 3 |
| 2020 | Challenging the abstraction penalty in parallel patterns libraries
José Daniel García, David del Rio Astorga, Marco Aldinucci, Fabio Tordini, Marco Danelutto, Gabriele Mencagli, Massimo Torquati |
J. Supercomput. | 6 |
| 2019 | Accelerating Actor-Based Applications with Parallel PatternsabstractParallel programmers mandate high-level parallel programming tools allowing to reduce the effort of the efficient parallelization of their applications. Parallel programming leveraging parallel patterns has recently received renovated attention thanks to their clear functional and parallel semantics. In this work, we propose a synergy between the well-known Actors-based programming model and the pattern-based parallelization methodology. We present our preliminary results in that direction, discussing and assessing the implementation of the Map parallel pattern by using an Actor-based software accelerator abstraction that seamlessly integrates within the C++ Actor Framework (CAF). The results obtained on the Intel Xeon Phi KNL platform demonstrate good performance figures achieved with negligible programming efforts. Luca Rinaldi, Massimo Torquati, Gabriele Mencagli, Marco Danelutto, Tullio Menga |
PDP | 3 |
| 2019 | Power-aware pipelining with automatic concurrency controlabstractSummary Continuous streaming computations are usually composed of different modules, exchanging data through shared message queues. The selection of the algorithm used to access such queues (ie, theconcurrency control) is a critical aspect both for performance and power consumption. In this paper, we describe the design of automatic concurrency control algorithm for implementing power‐efficient communications on shared‐memory multicores. The algorithm automatically switches between nonblocking and blocking concurrency protocols, getting the best from the two worlds, ie, obtaining the same throughput offered by the nonblocking implementation and the same power efficiency of the blocking concurrency protocol. We demonstrate the effectiveness of our approach using two micro‐benchmarks and two real streaming applications. Massimo Torquati, Daniele De Sensi, Gabriele Mencagli, Marco Aldinucci, Marco Danelutto |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | New Landscapes of the Data Stream Processing in the era of Fog Computing
Valeria Cardellini, Gabriele Mencagli, Domenico Talia, Massimo Torquati |
Future Gener. Comput. Syst. | 2 |
| 2019 | On dynamic memory allocation in sliding-window parallel patterns for streaming analytics
Massimo Torquati, Gabriele Mencagli, Maurizio Drocco, Marco Aldinucci, Tiziano De Matteis, Marco Danelutto |
J. Supercomput. | 2 |
| 2018 | SpinStreams: a Static Optimization Tool for Data Stream Processing ApplicationsabstractThe ubiquity of data streams in different fields of computing has led to the emergence of Stream Processing Systems (SPSs) used to program applications that extract insights from unbounded sequences of data items. Streaming applications demand various kinds of optimizations. Most of them are aimed at increasing throughput and reducing processing latency, and need cost models used to analyze the steady-state performance by capturing complex aspects like backpressure and bottleneck detection. In those systems, the tendency is to support dynamic optimizations of running applications which, although with a substantial run-time overhead, are unavoidable in case of unpredictable workloads. As an orthogonal direction, this paper proposes SpinStreams, a static optimization tool able to leverage cost models that programmers can use to detect and understand the inefficiencies of an initial application design. SpinStreams suggests optimizations for restructuring applications by generating code to be run on the SPS. We present the theory behind our optimizations, which cover more general classes of application structures than the ones studied in the literature so far. Then, we assess the accuracy of our models in Akka, an actor-based streaming framework providing a Java and Scala API. Gabriele Mencagli, Patrizio Dazzi, Nicolò Tonci |
Middleware | 1 |
| 2018 | Efficient NAS Benchmark Kernels with C++ Parallel ProgrammingabstractBenchmarking is a way to study the performance of new architectures and parallel programming frameworks. Well-established benchmark suites such as the NAS Parallel Benchmarks (NPB) comprise legacy codes that still lack portability to C++ language. As a consequence, a set of high-level and easy-to-use C++ parallel programming frameworks cannot be tested in NPB. Our goal is to describe a C++ porting of the NPB kernels and to analyze the performance achieved by different parallel implementations written using the Intel TBB, OpenMP and FastFlow frameworks for Multi-Cores. The experiments show an efficient code porting from Fortran to C++ and an efficient parallelization on average. Dalvan Griebler, Junior Loff, Gabriele Mencagli, Marco Danelutto, Luiz Gustavo Fernandes |
PDP | 3 |
| 2018 | Reducing Message Latency and CPU Utilization in the CAF Actor FrameworkabstractIn this work, we consider the C++ Actor Framework (CAF), a recent proposal that revamped the interest in building concurrent and distributed applications using the actor programming model in C++. CAF has been optimized for high-throughput computing, whereas message latency between actors is greatly influenced by the message data rate: at low and moderate rates the latency is higher than at high data rates. To this end, we propose a modification of the polling strategies in the work-stealing CAF scheduler, which can reduce message latency at low and moderate data rates up to two orders of magnitude without compromising the overall throughput and message latency at maximum pressure. The technique proposed uses a lightweight event notification protocol that is general enough to be used used to optimize the runtime of other frameworks experiencing similar issues. Massimo Torquati, Tullio Menga, Tiziano De Matteis, Daniele De Sensi, Gabriele Mencagli |
PDP | 5 |
| 2018 | Elastic-PPQ: A two-level autonomic system for spatial preference query processing over dynamic data streams
Gabriele Mencagli, Massimo Torquati, Marco Danelutto |
Future Gener. Comput. Syst. | 1 |
| 2018 | The home-forwarding mechanism to reduce the cache coherence overhead in next-generation CMPs
Gabriele Mencagli, Marco Vanneschi, Silvia Lametti |
Future Gener. Comput. Syst. | 1 |
| 2018 | Harnessing sliding-window execution semantics for parallel stream processing
Gabriele Mencagli, Massimo Torquati, Fabio Lucattini, Salvatore Cuomo, Marco Aldinucci |
J. Parallel Distributed Comput. | 1 |
| 2018 | Data stream processing via code annotations
Marco Danelutto, Tiziano De Matteis, Gabriele Mencagli, Massimo Torquati |
J. Supercomput. | 3 |
| 2017 | Elastic Scaling for Distributed Latency-Sensitive Data Stream OperatorsabstractHigh-volume data streams are straining the limits of stream processing frameworks which need advanced parallel processing capabilities to withstand the actual incoming bandwidth. Parallel processing must be synergically integrated with elastic features in order dynamically scale the amount of utilized resources by accomplishing the Quality of Service goals in a cost-effective manner. This paper proposes a control-theoretic strategy to drive the elastic behavior of latency-sensitive streaming operators in distributed environments. The strategy takes scaling decisions in advance by relying on a predictive model-based approach. Our ideas have been experimentally evaluated on a cluster using a real-world streaming application fed by synthetic and real datasets. The results show that our approach takes the strictly necessary reconfigurations while providing reduced resource consumption. Furthermore, it allows the operator to meet desired average latency requirements with a significant reduction in the experienced latency jitter. Tiziano De Matteis, Gabriele Mencagli |
PDP | 2 |
| 2017 | Proactive elasticity and energy awareness in data stream processing
Tiziano De Matteis, Gabriele Mencagli |
J. Syst. Softw. | 2 |
| 2017 | Bringing Parallel Patterns Out of the Corner: The P3 ARSEC Benchmark SuiteabstractHigh-level parallel programming is an active research topic aimed at promoting parallel programming methodologies that provide the programmer with high-level abstractions to develop complex parallel software with reduced time to solution. Pattern-based parallel programming is based on a set of composable and customizable parallel patterns used as basic building blocks in parallel applications. In recent years, a considerable effort has been made in empowering this programming model with features able to overcome shortcomings of early approaches concerning flexibility and performance. In this article, we demonstrate that the approach is flexible and efficient enough by applying it on 12 out of 13 PARSEC applications. Our analysis, conducted on three different multicore architectures, demonstrates that pattern-based parallel programming has reached a good level of maturity, providing comparable results in terms of performance with respect to both other parallel programming methodologies based on pragma-based annotations (i.e., Open mp and O mp S s ) and native implementations (i.e., P threads ). Regarding the programming effort, we also demonstrate a considerable reduction in lines of code and code churn compared to P threads and comparable results with respect to other existing implementations. Daniele De Sensi, Tiziano De Matteis, Massimo Torquati, Gabriele Mencagli, Marco Danelutto |
ACM Trans. Archit. Code Optim. | 4 |
| 2017 | Parallel Continuous Preference Queries over Out-of-Order and Bursty Data StreamsabstractTechniques to handle traffic bursts and out-of-order arrivals are of paramount importance to provide real-time sensor data analytics in domains like traffic surveillance, transportation management, healthcare and security applications. In these systems the amount of raw data coming from sensors must be analyzed by continuous queries that extract value-added information used to make informed decisions in real-time. To perform this task with timing constraints, parallelism must be exploited in the query execution in order to enable the real-time processing on parallel architectures. In this paper we focus on continuous preference queries, a representative class of continuous queries for decision making, and we propose a parallel query model targeting the efficient processing over out-of-order and bursty data streams. We study how to integrate punctuation mechanisms in order to enable out-of-order processing. Then, we present advanced scheduling strategies targeting scenarios with different burstiness levels, parameterized using the index of dispersion quantity. Extensive experiments have been performed using synthetic datasets and real-world data streams obtained from an existing real-time locating system. The experimental evaluation demonstrates the efficiency of our parallel solution and its effectiveness in handling the out-of-orderness degrees and burstiness levels of real-world applications. Gabriele Mencagli, Massimo Torquati, Marco Danelutto, Tiziano De Matteis |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | Keep calm and react with foresight: strategies for low-latency and energy-efficient elastic data stream processingabstractThis paper addresses the problem of designing scaling strategies for elastic data stream processing. Elasticity allows applications to rapidly change their configuration on-the-fly (e.g., the amount of used resources) in response to dynamic workload fluctuations. In this work we face this problem by adopting the Model Predictive Control technique, a control-theoretic method aimed at finding the optimal application configuration along a limited prediction horizon in the future by solving an online optimization problem. Our control strategies are designed to address latency constraints, using Queueing Theory models, and energy consumption by changing the number of used cores and the CPU frequency through the Dynamic Voltage and Frequency Scaling (DVFS) support available in the modern multicore CPUs. The proactive capabilities, in addition to the latency- and energy-awareness, represent the novel features of our approach. To validate our methodology, we develop a thorough set of experiments on a high-frequency trading application. The results demonstrate the high-degree of flexibility and configurability of our approach, and show the effectiveness of our elastic scaling strategies compared with existing state-of-the-art techniques used in similar scenarios. Tiziano De Matteis, Gabriele Mencagli |
PPoPP | 2 |
| 2016 | Continuous skyline queries on multicore architecturesabstractSummary The emergence of real‐time decision‐making applications in domains like high‐frequency trading, emergency management, and service level analysis in communication networks has led to the definition of new classes of queries.Skyline queriesare a notable example. Their results consist of all the tuples whose attribute vector is not dominated (in the Pareto sense) by one of any other tuple. Because of their popularity, skyline queries have been studied in terms of both sequential algorithms and parallel implementations for multiprocessors and clusters. Within theData Stream Processingparadigm, traditional database queries on static relations have been revised in order to operate on continuous data streams. Most of the past papers propose sequential algorithms for continuous skyline queries, whereas there exist very few works targeting implementations on parallel machines. This paper contributes to fill this gap by proposing a parallel implementation for multicore architectures. We propose (i) a parallelization of theeageralgorithm based on the notion ofSkyline Influence Time, (ii) optimizations of the reduce phase and load‐balancing strategies to achieve near‐optimal speedup, and (iii) a set of experiments with both synthetic benchmarks and a real dataset in order to show our implementation effectiveness. Copyright © 2016 John Wiley & Sons, Ltd. Tiziano De Matteis, Salvatore Di Girolamo, Gabriele Mencagli |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | Adaptive model predictive control of autonomic distributed parallel computations with variable horizons and switching costsabstractSummary Autonomic computing is a paradigm for building systems capable of adapting their operation when external changes occur, such as workload variations, load surges and changes in the resource availability. The optimal configuration in terms of the number of computing resources assigned to each component must be automatically adjusted to the new environmental conditions. To accomplish the execution goals with the desired Quality of Service, decision‐making strategies should be in charge of selecting the best reconfigurations by taking into account metrics like performance, efficiency (avoiding wasting resources), number and frequency of reconfigurations, and their amplitude (performing minimal modifications of the current configuration). This paper presents a decision‐making strategy that merges the potential of Model Predictive Control with a cooperative optimization framework. After a description of our approach, we investigate the effect of different switching costs to model the resource allocation problem. We use a control method in which our proactive decision‐making strategy (designed to use future prediction horizons) is made adaptive itself by dynamically changing the horizon length on the basis of the prediction errors. Simulations have been used to exemplify our approach and to discuss the effectiveness of the variable‐horizon strategy in achieving the best trade‐offs between reconfiguration metrics. Copyright © 2015 John Wiley & Sons, Ltd. Gabriele Mencagli |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | A Game-Theoretic Approach for Elastic Distributed Data Stream ProcessingabstractDistributed data stream processing applications are structured as graphs of interconnected modules able to ingest high-speed data and to transform them in order to generate results of interest. Elasticity is one of the most appealing features of stream processing applications. It makes it possible to scale up/down the allocated computing resources on demand in response to fluctuations of the workload. On clouds, this represents a necessary feature to keep the operating cost at affordable levels while accommodating user-defined QoS requirements. In this article, we study this problem from a game-theoretic perspective. The control logic driving elasticity is distributed among local control agents capable of choosing the right amount of resources to use by each module. In a first step, we model the problem as a noncooperative game in which agents pursue their self-interest. We identify the Nash equilibria and we design a distributed procedure to reach the best equilibrium in the Pareto sense. As a second step, we extend the noncooperative formulation with a decentralized incentive-based mechanism in order to promote cooperation by moving the agreement point closer to the system optimum. Simulations confirm the results of our theoretical analysis and the quality of our strategies. Gabriele Mencagli |
ACM Trans. Auton. Adapt. Syst. | 1 |
| 2015 | A Multicore Parallelization of Continuous Skyline Queries on Data Streams
Tiziano De Matteis, Salvatore Di Girolamo, Gabriele Mencagli |
Euro-Par | 3 |
| 2014 | A High-Throughput and Low-Latency Parallelization of Window-Based Stream Joins on MulticoresabstractData Stream Processing (DaSP) is a paradigm characterized by on-line (often real-time) applications working on unlimited data streams whose elements must be processed efficiently "on the fly". DaSP computations are characterized by data-flow graphs of operators connected via streams and working on the received elements according to high throughput and low latency requirements. To achieve these constraints, high-performance DaSP operators requires advanced parallelism models, as well related design and implementation techniques targeting multi-core architectures. In this paper we focus on the parallelization of the window-based stream join, an important operator that raises challenging issues in terms of parallel windows management. We review the state-of-the-art solutions about the stream join parallelization and we propose our novel parallel strategy and its implementation on multicores. As demonstrated by experimental results, our parallel solution introduces two important advantages with respect to the existing solutions: (i) it features an high-degree of configurability in order to address the symmetricity/asymmetricity of input streams (in terms of their arrival rate and window length), (ii) our parallelization provides a high throughput and it is definitely better than the compared solutions in terms of latency, providing an efficient way to perform stream joins on latency-sensible applications. Daniele Buono, Tiziano De Matteis, Gabriele Mencagli |
ISPA | 3 |
| 2014 | Optimizing Message-Passing on Multicore Architectures Using Hardware Multi-threadingabstractShared-memory and message-passing are two opposite models to develop parallel computations. The shared-memory model, adopted by existing frameworks such as OpenMP, represents a de-facto standard on multi-/many-core architectures. However, message-passing deserves to be studied for its inherent properties in terms of portability and flexibility as well as for its better ease of debugging. Achieving good performance from the use of messages in shared-memory architectures requires an efficient implementation of the run-time support. This paper investigates the definition of a delegation mechanism on multi-threaded architectures able to: (i) overlap communications with calculation phases, (ii) parallelize distribution and collective operations. Our ideas have been exemplified using two parallel benchmarks on the Intel Phi, showing that in these applications our message-passing support outperforms MPI and reaches similar performance compared to standard OpenMP implementations. Daniele Buono, Tiziano De Matteis, Gabriele Mencagli, Marco Vanneschi |
PDP | 3 |
| 2014 | A Cooperative Predictive Control Approach to Improve the Reconfiguration Stability of Adaptive Distributed Parallel ApplicationsabstractAdaptiveness in distributed parallel applications is a key feature to provide satisfactory performance results in the face of unexpected events such as workload variations and time-varying user requirements. The adaptation process is based on the ability to change specific characteristics of parallel components (e.g., their parallelism degree) and to guarantee that such modifications of the application configuration are effective and durable. Reconfigurations often incur a cost on the execution (a performance overhead and/or an economic cost). For this reason advanced adaptation strategies have become of paramount importance. Effective strategies must achieve properties like control optimality (making decisions that optimize the global application QoS), reconfiguration stability expressed in terms of the average time between consecutive reconfigurations of the same component, and optimizing the reconfiguration amplitude (number of allocated/deallocated resources). To control such parameters, in this article we propose a method based on a Cooperative Model-based Predictive Control approach in which application controllers cooperate to make optimal reconfigurations and taking account of the durability and amplitude of their control decisions. The effectiveness and the feasibility of the methodology is demonstrated through experiments performed in a simulation environment and by comparing it with other existing techniques. Gabriele Mencagli, Marco Vanneschi, Emanuele Vespa |
ACM Trans. Auton. Adapt. Syst. | 1 |
| 2013 | Reconfiguration Stability of Adaptive Distributed Parallel Applications through a Cooperative Predictive Control Approach
Gabriele Mencagli, Marco Vanneschi, Emanuele Vespa |
Euro-Par | 1 |
| 2011 | QoS-control of Structured Parallel Computations: A Predictive Control ApproachabstractA central issue for parallel applications executed on heterogeneous distributed platforms (e.g. Grids and Clouds) is assuring that performance and cost parameters are optimized throughout the execution. A solution is based on providing application components with adaptation strategies able to select at run-time the best component configuration. In this paper we will introduce a preliminary work concerning the exploitation of control-theoretic techniques for controlling parallel computations. In particular we will demonstrate how a predictive control approach can be used based on first-principle performance models of structured parallelism schemes. We will also evaluate the viability of our approach on a first experimental scenario. Gabriele Mencagli, Marco Vanneschi |
CloudCom | 1 |
| 2011 | Consistent reconfiguration protocols for adaptive high-performance applicationsabstractProgramming models for Pervasive Computing applications typically include the possibility of specifying software components according to multiple alternative versions, each optimized for a certain class of computing and communication technologies. A main mechanism provided by these programming models permits to dynamically select one of the alternative versions for the execution. This reconfiguration activity may be critical, from a performance point of view, when considering High-Performance Pervasive Computing applications, especially if the reconfiguration must be performed in such a way that the application semantics is respected (i.e. the reconfiguration is consistent). In this paper we show how to introduce consistent reconfiguration protocols for the ASSISTANT programming model, we exemplify two general protocols and we show experimental results for one of them. Carlo Bertolli, Gabriele Mencagli, Marco Vanneschi |
IWCMC | 2 |
| 2010 | Resource discovery support for time-critical adaptive applicationsabstractSeveral complex and time-critical applications require the existence of novel distributed and dynamical platforms composed of a variety of fixed and mobile processing nodes and networks. Notable examples of such applications are crisis and emergency management and natural phenomenon prediction. In this scenario we need the development of applications able to adapt their behavior according to the dynamical platform conditions, such as the presence of specific classes of computing resources and the actual network availability. For these reasons such adaptive applications need to interact with a fast and reliable resource discovery support, which ensures required response times by means of an high-degree of reconfigurability and selectivity. In this paper we present an integrated approach between our programming model for distributed adaptive time-critical computations and a suitable resource discovery support. Carlo Bertolli, Daniele Buono, Gabriele Mencagli, Massimo Torquati, Marco Vanneschi, Matteo Mordacchini, Franco Maria Nardini |
IWCMC | 3 |
| 2010 | Analyzing Memory Requirements for Pervasive Grid ApplicationsabstractPervasive Grid Computing Platforms include centralized computing nodes (e. g. parallel servers) as well as decentralized and mobile devices. Pervasive Grid applications include data- and computing-intensive components which can be mapped also onto decentralized and mobile nodes. The effective and practical success of this mapping resides also in deriving proper configurations of applications which consider the limited memory capabilities of those resources. In this paper we target this issue by showing how we can study and configure the memory requirements of an Emergency Management application. We present our solutions by using the ASSISTANT programming model for Pervasive Grid applications. Carlo Bertolli, Gabriele Mencagli, Marco Vanneschi |
PDP | 2 |
| 2009 | Next generation grids and wireless communication networks: towards a novel integrated approachabstractAbstract One of the most promising trends for next generation networks is to consider an integrated approach to the communication infrastructure and the processing layer. In particular, the introduction of broadband and reliable wireless networks allows the interaction of a huge number of devices all creating a single network. On the other hand, the grid paradigm is considered as one of the most promising approach for pervasive and dynamic applications. Aim of this paper is to present a novel integrated approach between grid paradigm and wireless networks by highlighting the main advantages of their cooperation. In particular, it will be shown here how a wireless heterogeneous network can be exploited for implementing a pervasive and dynamic grid (mobile grid) and, on the other hand, a mobile grid allows the optimization of the communication infrastructure. The integrated approach can be an effective method for solving applications, such as emergency management, where a huge amount of data derived from a wireless infrastructure needs to be processed efficiently and adaptively, and the traffic flow in the wide area wireless networks needs to be coordinated and optimized. Copyright © 2008 John Wiley & Sons, Ltd. Romano Fantacci, Marco Vanneschi, Carlo Bertolli, Gabriele Mencagli, Daniele Tarchi |
Wirel. Commun. Mob. Comput. | 4 |