VLDB 2026 Research / reviewers in the wild / expert
Alessandro Pellegrini 0001
dblp:15/2894
· DBLP profile ↗
69ranked-venue papers
8as first author
24since 2021 · last 2026
0000-0002-0179-9868ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 15 · 8 since 2021Artificial intelligence and machine learning · 14 · 7 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Operator Rebinding for Stream Processing on NUMA MachinesabstractABSTRACT Introduction Modern stream processing engines are increasingly deployed on high‐core‐count servers with Non‐Uniform Memory Access (NUMA) architectures, where the cost of inter‐socket memory access poses a significant challenge to achieving low latency and high throughput. Existing approaches to operator placement either rely on static assignments that degrade under workload variations or employ dynamic migrations that incur excessive overhead due to blocking synchronization or global barriers. Methods This paper introduces a lock‐free, NUMA‐aware operator rebinding mechanism that dynamically reallocates operator tasks across threads with minimal disruption. The mechanism uses an autonomic controller to detect imbalance in per‐thread queues and enacts rebinding via control messages and atomic updates, ensuring correctness without stalling execution. A two‐level policy is proposed, combining NUMA‐level partitioning with intra‐node thread‐level refinements, triggered by latency thresholds. Results Extensive experiments using a 300‐query urban traffic analytics workload demonstrate that the proposed method achieves non‐negligible throughput improvement and reduces latency compared to state‐of‐the‐art static and METIS‐based approaches. Furthermore, it reduces latency variance by an order of magnitude, illustrating the importance of fine‐grained NUMA‐aware scheduling in memory‐bound stream processing. Xiaorui Du, Andrea Piccione, Adriano Pimpini, Stefano Bortoli, Alessandro Pellegrini 0001, Alois C. Knoll |
Softw. Pract. Exp. | 5 |
| 2025 | Model-Driven Parallel and Distributed Stochastic Simulation of Chemical Reaction Networks
Simone Bauco, Federica Montesano, Adriano Pimpini, Romolo Marotta, Alessandro Pellegrini 0001 |
DS-RT | 5 |
| 2025 | DESL: A Literate Programming Language Framework for Interoperable Parallel Discrete Event SimulationabstractSimulation is indispensable for advanced scientific research, enabling accurate explorations of complex phenomena and supporting evidence-based decision-making across interdisciplinary boundaries. Parallel Discrete Event Simulation (PDES) provides substantial advantages in modelling large-scale systems by distributing computational tasks among multiple processors, enhancing scalability. However, exploiting it is extremely challenging due to obstacles in model efficiency, concurrency control, reproducibility, and maintainability. Furthermore, the large number of available PDES run-time environments makes it difficult to explore their (performance) capabilities for some specific model, hindering the identification of the best-suited technology for a certain simulation study. To address these limitations, we introduce a unified framework grounded in literate programming and model-driven engineering, integrating interwoven documentation and model logic within a single source. This design enhances intrinsic consistency between model logic and explanatory content, while enabling the generation of model implementations tailored to multiple runtime environments, thus allowing simulationists to focus on model development without being locked in to any specific technology or environment. This facilitates model reuse and performance comparisons across diverse execution environments. We show the viability of this approach by providing the first-ever experimental comparison across three different simulators, starting from the same model implementation. Simone Bauco, Romolo Marotta, Alessandro Pellegrini 0001 |
SIGSIM-PADS | 3 |
| 2025 | Test of Time Award: Transparently Mixing Undo Logs and Software Reversibility for State Recovery in Optimistic PDESabstractThe paper "Transparently Mixing Undo Logs and Software Reversibility for State Recovery in Optimistic PDES" introduced a seminal hybrid rollback technique that efficiently and transparently combines checkpointing and reverse computation methods through runtime-generated undo instructions. This short abstract, which accompanies the Test of Time Award received at PADS 2025, summarises the fundamental challenges presented in the original paper and reflects on its impact in the decade after its publication. Davide Cingolani, Alessandro Pellegrini 0001, Francesco Quaglia |
SIGSIM-PADS | 2 |
| 2025 | RBlockSim: Parallel and Distributed Simulation for Blockchain BenchmarkingabstractSince the introduction of Bitcoin in 2008, blockchain technology's popularity has been rapidly on the rise.This has led to and benefited from the introduction of numerous new implementations addressing specific needs regarding performance, scalability, privacy, etc.Developing new implementations requires making design choices that greatly influence the blockchain's behaviour.While deploying full-scale networks for evaluation is impractical and costly, simulation offers a safe and convenient environment to test and benchmark different implementations against the requirements.Although single-threaded simulation of blockchain networks is available, it can incur long execution times when simulating largescale scenarios.Additionally, it is constrained by the machine's RAM, limiting the size of networks one may study.In this paper, we introduce and benchmark RBlockSim, a simulation approach that leverages optimistic parallel and distributed discrete-event simulation to overcome the above limitations and provides a modular, high-performance test-bed for blockchain evaluation. Adriano Pimpini, Alessandro Pellegrini 0001 |
SIGSIM-PADS | 2 |
| 2025 | Would you mind hiding my malware? Building malicious Android apps with StegoPackabstractThis paper empirically explores the resilience of the current Android ecosystem against stegomalware, which involves both Java/Kotlin and native code. To this aim, we rely on a methodology that goes beyond traditional approaches by hiding malicious Java code and extending it to encoding and dynamically loading native libraries at runtime. By merging app resources, steganography, and repackaging, the methodology seamlessly embeds malware samples into the assets of a host app, making detection significantly more challenging. We implemented the methodology in a tool, StegoPack, which allows the extraction and execution of the payload at runtime through reverse steganography. We used StegoPack to embed well-known DEX and native malware samples over 14 years into real Android host apps. We then challenged top-notch antivirus engines, which previously had high detection rates on the original malware, to detect the embedded samples. Our results reveal a significant reduction in the number of detections (up to zero in most cases), indicating that current detection techniques, while thorough in analyzing app code, largely disregard app assets, leading us to believe that steganographic adversaries are not even included in the adversary models of most deployed defensive analysis systems. Thus, we propose potential countermeasures for StegoPack to detect steganographic data in the app assets and the dynamic loader used to execute malware. Danilo Dell'Orco, Giorgio Bernardinetti, Giuseppe Bianchi 0001, Alessio Merlo, Alessandro Pellegrini 0001 |
Pervasive Mob. Comput. | 5 |
| 2024 | HUILLY: A Non-Blocking Ingestion Buffer for Timestepped Simulation AnalyticsabstractWe present HUILLY, a non-blocking data ingestion buffer designed for parallel applications built relying on time-stepped, fork-join computational paradigm. It provides complete data separation, reducing the intricacies of multi-threaded data structures and provides high operational efficiency. The effectiveness of HUILLY as a non-blocking ingestion buffer is demonstrated through an extensive experimental evaluation, providing significant improvements in throughput and latency for time-stepped, fork-join applications. Xiaorui Du, Andrea Piccione, Adriano Pimpini, Stefano Bortoli, Alois C. Knoll, Alessandro Pellegrini 0001 |
CCGrid | 6 |
| 2024 | Sampling Policies for Near-Optimal Device Choice in Parallel Simulations on CPU/GPU PlatformsabstractHeterogeneous hardware platforms comprised of CPUs, GPUs, and other accelerators offer the opportunity to choose the best-suited device for executing a given scientific simulation in order to minimize execution time and energy consumption. To this end, the recently proposed "Follow the Leader" approach dynamically selects a suitable device based on runtime performance measurements during speculative discrete-event simulations. A currently active "leader" device is periodically challenged by a "follower" device in order to negotiate the new leader. The optimality of the device choices and the associated overhead depends critically on the challenge frequency and timing. Here, we explore policies to schedule challenges with the goal of attaining Pareto-optimal combinations of execution time and energy consumption. Several heuristics are first evaluated in an abstract fashion using a "meta-simulation" by mimicking the progress and energy consumption of an idealized co-execution. In this setting, we optimize the heuristics’ tuning parameters to assess their relative merits in near-optimal configurations when compared to challenge timings based on perfect knowledge. We find that under challenging stochastic workloads based on a class of mean-reverting random walks, the best heuristics can closely approximate the execution time and energy consumption achievable under an optimal device choice. Empirical support for this observation is given by measurements of a CPU/GPU co-execution of the Time Warp algorithm on physical hardware. Philipp Andelfinger, Alessandro Pellegrini 0001, Romolo Marotta |
DS-RT | 2 |
| 2024 | Online Analytics with Local Operator Rebinding for Simulation Data Stream ProcessingabstractLeveraging multiple threads to process high volumes of simulation data is a prevalent strategy in modern streaming data processing systems. Statically binding operators to specific threads is the most common design employed due to its simplicity in implementation and initial system configuration. However, this approach often fails to effectively account for the inherently dynamic nature of simulation data, potentially leading to inefficient resource utilisation and processing bottlenecks. To address these limitations, we present a novel mechanism for stream-processing operator rebinding that enables lock-free, dynamic workload rebalancing between worker threads. The rebinding is driven by an autonomic policy that captures workload imbalance in the stream-processing pipeline when multiple queries are computed and reacts to it by moving computation around the different threads. We evaluate our proposal using data generated from large-scale traffic simulations on which multiple queries are executed. The volume and organisation of the data we feed to the stream-processing pipeline significantly change over time, providing excellent grounds to evaluate our rebinding policy. The evaluation confirms that the performance of stream processing pipelines can be greatly improved using local operator rebinding. Xiaorui Du, Andrea Piccione, Adriano Pimpini, Stefano Bortoli, Alessandro Pellegrini 0001, Alois C. Knoll |
DS-RT | 5 |
| 2024 | Follow the Leader: Alternating CPU/GPU Computations in PDESabstractDespite the successes of graphics processing units (GPUs) in accelerating simulations in several research fields, their use is largely restricted to domain-specific workloads that consistently offer the large degree of inherent parallelism and computational intensity at which GPUs excel. When targeting generic discrete-event simulations, whose dynamics can vary wildly over time, a static choice between a GPU-based and traditional CPU-based execution is likely to be suboptimal. Here, we explore a parallel discrete-event (PDES) execution scheme for CPU-GPU platforms that aims to approximate an optimal dynamic device choice. Starting from an intermediate model state, a current “leader” device running the simulation is periodically challenged by a brief concurrent run on another device starting from an intermediate model state. Based on the gathered performance measurements, a forecasting scheme determines the leader for the next period. The execution time and power consumption of this scheme hinge on 1) an efficient mechanism for providing the “follower” device with a consistent model state, and 2) robust performance forecasting to justify the device choices. We present these building blocks, their implementation combining the existing CPU and GPU simulators ROOT-Sim and GPUTW, and measurement results demonstrating substantially reduced execution time without increasing energy consumption over a static device choice. Romolo Marotta, Alessandro Pellegrini 0001, Philipp Andelfinger |
SIGSIM-PADS | 2 |
| 2024 | Efficient Non-Blocking Event Management for Speculative Parallel Discrete Event SimulationabstractParallel Discrete Event Simulation (PDES) is a modelling technique that takes advantage of concurrent computing resources. However, its asynchronous nature can present challenges for efficient execution. This paper proposes a new non-blocking management system for handling messages and anti-messages in Time Warp simulations. This approach exploits the benefits of non-blocking algorithms to surpass the limitations of existing blocking mechanisms, resulting in more efficient and scalable simulations. Specifically, the approach relies on efficient atomic fetch-and-add operations provided by modern computer architectures for evaluating and updating the status of the event. Andrea Piccione, Alessandro Pellegrini 0001 |
SIGSIM-PADS | 2 |
| 2023 | Incremental Checkpointing of Large State Simulation Models with Write-Intensive Events via Memory Update Correlation on Buddy PagesabstractCheckpointing techniques for speculative parallel simulation of discrete event models have been widely studied in the literature. However, there has been a very marginal attempt to exploit operating system page-protection services, which have instead been largely exploited in the context of checkpointing for fault tolerance. In this article, we discuss how these services can effectively manage simulation models with large states and write-intensive events in zones of the state layout. In particular, we present a solution where the correlation of write operations on buddy pages in the state layout can be exploited to achieve effective incremental checkpointing support, which allows scaling down the costs of operating system services. Our solution does not require any instrumentation of the simulation application code and is usable on any Posix-compliant operating system. We also discuss its integration within the USE (Ultimate-Share-Everything) open-source speculative simulation package and report some experimental data for its assessment. Romolo Marotta, Federica Montesano, Alessandro Pellegrini 0001, Francesco Quaglia |
DS-RT | 3 |
| 2023 | Practical Tie-Breaking for Parallel/Distributed SimulationsabstractIn this paper, we discuss a tie-breaking strategy based on a bitwise comparison of event payload that allows parallel and distributed discrete-event simulations to observe a deterministic order in the execution of events, even in the presence of event ties. This approach provides practical usability whenever model-assisted tie-breaking is unavailable, thus ensuring that multiple simulation executions provide deterministic behaviour and repeatable results. Moreover, it ensures that the selected order of events is also consistent with sequential executions. We discuss the theory behind this strategy and experimentally show that the performance drop is imputable to event queue management when relying on tie-breaking strategies like the ones discussed in this work. Andrea Piccione, Alessandro Pellegrini 0001 |
DS-RT | 2 |
| 2023 | Hybrid Speculative Synchronisation for Parallel Discrete Event SimulationabstractParallel discrete-event simulation (PDES) is a well-established family of methods to accelerate discrete-event simulations. However, the available algorithms vary substantially in the performance achievable for different models, largely preventing generic solutions applicable by modellers without expert knowledge. For instance, in Time Warp, the processing elements execute events asynchronously and speculatively with high aggressiveness, leading to frequent and costly rollbacks if misspeculations occur often. In contrast, synchronous approaches such as the new Window Racer algorithm exhibit a more cautious form of speculation. In the present paper, we combine these two fundamentally different algorithms within a single runtime environment, allowing for a choice of the best algorithm for different model segments. We describe the architecture and the algorithmic considerations to support the efficient coexistence and interaction of the algorithms without violating the correctness of the simulation. Our experiments using a synthetic benchmark and an epidemics model show that the hybrid algorithm is less sensitive to its configuration and can deliver substantially higher performance in models with varying degrees of coupling among entities compared to each algorithm on its own. Andrea Piccione, Philipp Andelfinger, Alessandro Pellegrini 0001 |
SIGSIM-PADS | 3 |
| 2023 | What makes test programs similar in microservices applications?abstractThe emergence of microservices architecture calls for novel methodologies and technological frameworks that support the design, development, and maintenance of applications structured according to this new architectural style. In this paper, we consider the issue of designing suitable strategies for the governance of testing activities within the microservices paradigm. We focus on the problem of discovering implicit relations between test programs that help to avoid re-running all the available test suites each time one of its constituents evolves. We propose a dynamic analysis technique and its supporting framework that collects information about the invocations of local and remote APIs. Information on test program execution is obtained in two ways: instrumenting the test program code or running a symbolic execution engine. The extracted information is processed by a rule-based automated reasoning engine, which infers implicit similarities among test programs. We show that our analysis technique can be used to support the reduction of test suites, and therefore has good application potential in the context of regression test optimisation. The proposed approach has been validated against two real-world microservices applications. Emanuele De Angelis, Guglielmo De Angelis, Alessandro Pellegrini 0001, Maurizio Proietti |
J. Syst. Softw. | 3 |
| 2023 | Strategies and software support for the management of hardware performance countersabstractAbstract Hardware performance counters (HPCs) are facilities offered by most off‐the‐shelf CPU architectures. They are a vital support to post‐mortem performance profiling and are exploited by standard tools such as Linux or Intel V‐Tune. Nevertheless, an increasing number of application domains (e.g., simulation, task‐based high‐performance computing, or cybersecurity) are exploiting them to perform different activities, such as self‐tuning, autonomic optimization, and/or system inspection. This repurposing of HPCs can be difficult, for example, because of the overhead for extracting relevant information. This overhead might render any online or self‐tuning activity ineffective. This article discusses various practical strategies to exploit HPCs beyond post‐mortem profiling, suitable for different application contexts. The presented strategies are accompanied by a general primer on HPCs usage on Linux. We also provide reference x86 (both Intel and AMD) implementations targeting the Linux kernel, upon which we present an experimental assessment of the viability of our proposals. Stefano Carnà, Romolo Marotta, Alessandro Pellegrini 0001, Francesco Quaglia |
Softw. Pract. Exp. | 3 |
| 2022 | Comparing Speculative Synchronization Algorithms for Continuous-Time Agent-Based SimulationsabstractContinuous-time agent-based models often represent tightly-coupled systems in which an agent’s state transitions occur in close interaction with neighboring agents. Without artificial discretization, the potential for near-instantaneous propagation of effects across the model presents a challenge to parallelizing their execution. Although existing algorithms can tackle the largely unpredictable nature of such simulations through speculative execution, they are subject to trade-offs concerning the degree of optimism, the probability and cost of rollbacks, and the exploitation of locality. This paper is aimed at understanding the suitability of asynchronous and synchronous parallel simulation algorithms when executing continuous-time agent-based models with rate-driven stochastic transitions. We present extensive measurement results comparing optimized implementations under various configurations of a parametrizable simulation model of the epidemic spread of disease. Our results show that the amount of locality in the agent interactions is the decisive factor for the relative performance of the approaches. Based on profiling results, we identify remaining hurdles for higher simulation performance with the two classes of algorithms and outline potential refinements. Philipp Andelfinger, Andrea Piccione, Alessandro Pellegrini 0001, Adelinde M. Uhrmacher |
DS-RT | 3 |
| 2022 | On the Accuracy and Performance of Spiking Neural Network SimulationsabstractSpiking Neural Networks (SNNs) are a class of Artificial Neural Networks that show a time behaviour that cannot be computed with single one-shot functions. Therefore, to study their evolution over time, simulations are typically employed. Typical simulation approaches rely on time-stepped simulations, while more recent works have highlighted the opportunity to rely on Parallel Discrete Event Simulation (PDES) for improved accuracy. In particular, Speculative PDES has been shown to be a suitable simulation paradigm to deal with the peculiar temporal domain of SNNs. In this paper, we perform an experimental evaluation of these two different approaches, showing the implications on both simulation performance and accuracy. Our assessment showcases that Parallel Discrete Event Simulation can deliver good scaling on parallel architectures while offering more accurate results. Adriano Pimpini, Andrea Piccione, Alessandro Pellegrini 0001 |
DS-RT | 3 |
| 2022 | Speculative Distributed Simulation of Very Large Spiking Neural NetworksabstractSpiking Neural Networks are a class of Artificial Neural Networks that closely mimic biological neural networks. They are particularly interesting because of their potential to advance research in several fields, both because of better insights on neural behaviour (benefiting medicine, neuroscience, psychology) and the potential in Artificial Intelligence. Their ability to run on a low energy budget once implemented in hardware makes them even more appealing. However, because of their behaviour that evolves with time, when a hardware implementation is not available, their output cannot simply be computed with a one-shot function (however complex), but instead they need to be simulated. Adriano Pimpini, Andrea Piccione, Bruno Ciciani, Alessandro Pellegrini 0001 |
SIGSIM-PADS | 4 |
| 2022 | Design and implementation of a fully transparent partial abort support for software transactional memoryabstractAbstract Software transactional memory (STM) provides synchronization support to ensure atomicity and isolation when threads access shared data in concurrent applications. With STM, shared data accesses are encapsulated within transactions automatically handled by the STM layer. Hence, programmers are not requested to use code‐synchronization mechanisms explicitly, like locking. In this article, we present our experience in designing and implementing a partial abort scheme for STM. The objective of our work is threefold: (1) enabling STM to undo only part of the transaction execution in the case of conflict, (2) designing a scheme that is fully transparent to programmers, thus also allowing to run existing STM applications without modifications, and (3) providing a scheme that can be easily integrated within existing STM runtime environments without altering their internal structure. The scheme we designed is based on automated software instrumentation, which injects into the application capabilities to undo the required portions of transaction executions. Further, it can correctly undo also non‐transactional operations executed on the stack and the heap during a transaction. This capability allows programmers to write transactional code without concerns about the side effects of aborted transactions on both shared and thread‐private data. We integrated and evaluated our partial abort scheme within the TinySTM open‐source library. We analyze the experimental results we achieved with common STM benchmark applications, focusing on the advantages and disadvantages of the proposed solutions for implementing our scheme's different components. Hence, we highlight the appropriate choices and possible solutions to improve partial abort schemes further. Alessandro Pellegrini 0001, Pierangelo di Sanzo, Andrea Piccione, Francesco Quaglia |
Softw. Pract. Exp. | 1 |
| 2022 | NBBS: A Non-Blocking Buddy System for Multi-Core MachinesabstractCommon implementations of core memory allocation components handle concurrent allocation/release requests by synchronizing threads via spin-locks. This approach is not prone to scale, a problem that has been addressed in the literature by introducing layered allocation services or replicating the core allocators—the bottom-most ones within the layered architecture. Both these solutions tend to reduce the pressure of actual concurrent accesses to each individual core allocator. In this article, we explore an alternative approach to scalability of memory allocation/release, which can be still combined with those literature proposals. We present a fully non-blocking buddy system, where threads performing concurrent allocations/releases do not undergo any spin-lock based synchronization. Our solution allows threads to proceed in parallel, and commit their allocations/releases unless a conflict is materialized while handling the allocator metadata—memory fragmentation and coalescing are also carried out in a fully non-blocking manner. Conflict detection relies in our solution on atomic Read-Modify-Write (RMW) machine instructions, guaranteed to execute atomically by the processor firmware. We also provide a proof of the correctness of our non-blocking buddy system and show the results of an experimental study that outlines the effectiveness of our solution. Romolo Marotta, Mauro Ianni, Alessandro Pellegrini 0001, Francesco Quaglia |
IEEE Trans. Computers | 3 |
| 2022 | Effective Runtime Management of Tasks and Priorities in GNU OpenMP ApplicationsabstractOpenMP has become a reference standard for the design of parallel applications. This standard is evolving quickly, thus offering new opportunities to the application programmers. However, OpenMP runtime environments are often not fully aligned with the actual requirements imposed by the evolution of such a standard. Among the main lacks, we find: (a) a limited capability to effectively cope with task priorities, and (b) the inadequacy in guaranteeing core properties while processing tasks such as the so-calledwork-conservativeness—the ability of the OpenMP runtime environment to fully exploit the underlying multi-processor/multi-core machine through the avoidance of thread-blocking phases. In this article, we present the design of extensions to the GNU OpenMP (GOMP) implementation, integrated intogcc, which allow the effective management of tasks and their priorities. Our proposal is based on a user-space library—modularly combined with the one already offered byGOMP—and an external kernel-level Linux module—offering the opportunity to exploit raising hardware facilities for task/priority management. We also provide experimental results showing the effectiveness of our proposal, achieved by running either OpenMP common benchmarks or a new benchmark application (Hashtag-Text) that we explicitly devised to stress the runtime environment in relation to the above-mentioned task/priority management aspects. Emiliano Silvestri, Alessandro Pellegrini 0001, Pierangelo di Sanzo, Francesco Quaglia |
IEEE Trans. Computers | 2 |
| 2021 | On power capping and performance optimization of multithreaded applicationsabstractSummary Multithreaded applications facilitate the exploitation of the computing power of multicore architectures. On the other hand, these applications can become extremely energy‐intensive, in contrast with the need for limiting the energy usage of computing systems. In this article, we explore the design of techniques enabling multithreaded applications to maximize their performance under a power cap. We consider two control parameters: the number of cores used by the application, and the core power state. We target the design of an autotuning power‐capping technique with minimal intrusiveness and high portability, which is agnostic about the workload profile of the application. We investigate two different approaches for building the strategy for selecting the best configuration of the parameters under control, namely a heuristic approach and a model‐based approach. Through an extensive experimental study, we evaluate the effectiveness of the proposed technique considering two different selection strategies, and we compare them with existing solutions. Stefano Conoci, Pierangelo di Sanzo, Alessandro Pellegrini 0001, Bruno Ciciani, Francesco Quaglia |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | Autonomic rejuvenation of cloud applications as a countermeasure to software anomaliesabstractSummary Failures in computer systems can be often tracked down to software anomalies of various kinds. In many scenarios, it might be difficult, unfeasible, or unprofitable to carry out extensive debugging activity to spot the cause of anomalies and remove them. In other cases, taking corrective actions may led to undesirable service downtime. In this article, we propose an alternative approach to cope with the problem of software anomalies in cloud‐based applications, and we present the design of a distributed autonomic framework that implements our approach. It exploits the elastic capabilities of cloud infrastructures, and relies on machine learning models, proactive rejuvenation techniques, and a new load balancing approach. By putting together all these elements, we show that it is possible to improve both availability and performance of applications deployed to heterogeneous cloud regions and subject to frequent failures. Overall, our study demonstrates the viability of our approach, thus opening the way towards its adoption, and encouraging further studies and practical experiences to evaluate and improve it. Pierangelo di Sanzo, Dimiter R. Avresky, Alessandro Pellegrini 0001 |
Softw. Pract. Exp. | 3 |
| 2020 | Agent-based Modeling and Simulation for Emergency Scenarios: A Holistic ApproachabstractAgent-based Modeling and Simulation is a powerful technique which allows to study the interactions in complex systems, and allows to explore or even foresee the emergence of more complicated properties or behaviors related to the interaction among the simpler agents in the environment. In the context of emergency or crisis scenarios, Agent-based Modeling and Simulation can allow to effectively study emergency plans, with the goal of assessing their viability, also with respect to the number of possible fatalities. In this paper, we analyze Agent-based Modeling and Simulation for crisis scenarios from a methodological and empirical point of view, with the goal of identifying what are the behavioral parameters that a model should encompass, in order for the results of the simulation to be useful for emergency plan assessment and/or compilation. We also experimentally provide a characterization of the effects of such behavioral parameters. Andrea Piccione, Alessandro Pellegrini 0001 |
DS-RT | 2 |
| 2020 | NUMA-Aware Non-Blocking Calendar QueueabstractModern computing platforms are based on multi-processor/multi-core technology. This allows running applications with a high degree of hardware parallelism. However, medium-to-high end machines pose a problem related to the asymmetric delays threads experience when accessing shared data. Specifically, Non-Uniform-Memory-Access (NUMA) is the dominating technology-thanks to its capability for scaled-up memory bandwidth-which however imposes asymmetric distances between CPU-cores and memory banks, making an access by a thread to data placed on a far NUMA node severely impacting performance. In this article, we tackle this problem in the context of shared event-pool management, a relevant aspect in many fields, like parallel discrete event simulation. Specifically, we present a NUMA-aware calendar queue, which also has the advantage of making concurrent threads coordinate via a non-blocking scalable approach. Our proposal is based on work deferring combined with dynamic re-binding of the calendar queue operations (insertions/extractions) to the best suited among the concurrent threads hosted by the underlying computing platform. This changes the locality of the operations by threads in a way positively reflected onto NUMA tasks at the hardware level. We report the results of an experimental study, demonstrating the capability of our solution to achieve the order of 15% better performance compared to state-of-the-art solutions already suited for multicore environments. Maryan Rab, Romolo Marotta, Mauro Ianni, Alessandro Pellegrini 0001, Francesco Quaglia |
DS-RT | 4 |
| 2020 | Autonomic Power Management in Speculative Simulation Runtime EnvironmentsabstractWhile transitioning to exascale systems, it has become clear that power management plays a fundamental role to support a viable utilization of the underlying hardware, also performance-wise. To meet power restrictions imposed by future exascale supercomputers, runtime environments will be required to enforce self-tuning schemes to run dynamic workloads under an imposed power cap. Literature results show that, for a wide class of multi-threaded applications, tuning both the degree of parallelism and frequency/voltage of cores allows a more effective use of the budget, compared to techniques that use only one of these mechanisms in isolation. In this paper, we explore the issues associated with applying these techniques on speculative Time-Warp based simulation runtime environments. We discuss how the differences in two antithetical Time Warp-based simulation environments impact the obtained results. Our assessment confirms that the performance gains achieved through a proper allocation of the power budget can be significant. We also identify the research challenges that would make these form of self-tuning more broadly applicable. Stefano Conoci, Mauro Ianni, Romolo Marotta, Alessandro Pellegrini 0001 |
SIGSIM-PADS | 4 |
| 2020 | Approximated RollbacksabstractA rollback operation in a speculative parallel discrete event simulator has traditionally targeted the perfect reconstruction of the state to be restored after a timestamp-order violation. This imposes that the rollback support entails specific capabilities and consequently pays given costs. In this article we propose approximated rollbacks, which allow a simulation object to perfectly realign its virtual time to the timestamp of the state to be restored, but lead the reconstructed state to be an approximation of what it should really be. The advantage is an important reduction of the cost for managing the state restore task in a rollback phase, as well as for managing the activities (i.e. state saving) that actually enable rollbacks to be executed. Our proposal is suited for stochastic simulations, and explores a tradeoff between the statistical representativeness of the outcome of the simulation run and the execution performance. We provide mechanisms that enable the application programmer to control this tradeoff, as well as simulation-platform level mechanisms that constitute the basis for managing approximate rollbacks in general simulation scenarios. A study on the aforementioned tradeoff is also presented. Matteo Principe, Andrea Piccione, Alessandro Pellegrini 0001, Francesco Quaglia |
SIGSIM-PADS | 3 |
| 2020 | Exploiting Inter-Processor-Interrupts for Virtual-Time Coordination in Speculative Parallel Discrete Event SimulationabstractReducing the waste of resource usage (e.g., CPU-cycles) when a causality error occurs in speculative parallel discrete event simulation (PDES) is still a core objective. In this article, we target this objective in the context of speculative PDES run on top of shared-memory machines. We propose an Operating System approach that is based on the exploitation of the Inter-Processor-Interrupt (IPI) facility offered by off-the-shelf hardware chipsets, which enables cross-CPU-core control of the execution flow of threads. As soon as a thread T produces a new event placed in the past virtual time of a simulation object currently run by another thread T', our IPI-based support allows T to change the execution flow of T'---with very minimal delay---so to enable the early squash of the currently processed (and no longer consistent) event. Our solution is fully transparent to the application level code, and is coupled with a lightweight heuristic-based mechanism that determines the actual goodness of killing thread T' via the IPI (rather than skipping the IPI send) depending on the expected residual execution time of the incorrect event being processed. We integrated our proposal within the speculative open-source USE (Ultimate Share Everything) PDES package, and we report experimental results obtained by running various PDES models on top of two shared-memory hardware architectures equipped with 32 and 24 (48 Hyper-threads) CPU-cores, which demonstrate the effectiveness of our proposal. Emiliano Silvestri, Cristian Milia, Romolo Marotta, Alessandro Pellegrini 0001, Francesco Quaglia |
SIGSIM-PADS | 4 |
| 2020 | Mutable locks: Combining the best of spin and sleep locksabstractSummary In this article, we present mutable locks, a synchronization construct with the same semantic of traditional locks (such as spin locks or sleep locks), but with a self‐tuned optimized trade‐off between responsiveness and CPU‐time usage during threads' wait phases. Mutable locks tackle the need for efficient synchronization supports in the era of multicore machines, where the run‐time performance should be optimized while reducing resource usage. This goal should be achieved with no intervention by the programmers. Our proposal is intended for exploitation in generic concurrent applications, where scarce or no knowledge is available about the underlying software/hardware stack and the workload. This is an adverse scenario for static choices between spinning and sleeping, which is tackled by our mutable locks thanks to their hybrid waiting phase and self‐tuning capabilities. Romolo Marotta, Davide Tiriticco, Pierangelo di Sanzo, Alessandro Pellegrini 0001, Bruno Ciciani, Francesco Quaglia |
Concurr. Comput. Pract. Exp. | 4 |
| 2020 | Adaptive Model-Based Scheduling in Software Transactional MemoryabstractSoftware Transactional Memory (STM) stands as powerful concurrent programming paradigm, enabling atomicity, and isolation while accessing shared data. On the downside, STM may suffer from performance degradation due to excessive conflicts among concurrent transactions, which cause waste of CPU-cycles and energy because of transaction aborts. An approach to cope with this issue consists of putting in place smart scheduling strategies which temporarily suspend the execution of some transaction in order to reduce the transaction conflict rate. In this article, we present an adaptive model-based transaction scheduling technique relying on a Markov Chain-based performance model of STM systems. Our scheduling technique is adaptive in a twofold sense: (i) It controls the execution of transactions depending on throughput predictions by the model as a function of the current system state. (ii) It re-tunes on-line the Markov Chain-based model to adapt it-and the outcoming transaction scheduling decisions-to dynamic variations of the workload. We have been able to achieve the latter target thanks to the fact that our performance model is extremely lightweight. In fact, to be recomputed, it requires a reduced set of input parameters, whose values can be estimated via a few on-line samples related to the current workload dynamics. We also present a scheduler that implements our adaptive technique, which we integrated within the open source TinySTM package. Further, we report the results of an experimental study based on the STAMP benchmark suite, which has been aimed at assessing both the accuracy of our performance model in predicting the actual system throughput and the advantages of the adaptive scheduling policy over literature techniques. Pierangelo di Sanzo, Alessandro Pellegrini 0001, Marco Sannicandro, Bruno Ciciani, Francesco Quaglia |
IEEE Trans. Computers | 2 |
| 2019 | NBBS: A Non-Blocking Buddy System for Multi-core MachinesabstractCommon implementations of core memory allocation components, like the Linux buddy system, handle concurrent allocation/release requests by synchronizing threads via spin-locks. This approach is not prone to scale, a problem that has been addressed in the literature by introducing layered allocation services or replicating the core allocators-the bottom most ones within the layered architecture. Both these solutions tend to reduce the pressure of actual concurrent accesses to each individual core allocator. In this article we explore an alternative approach to scalability of memory allocation/release, which can be still combined with those literature proposals. We present a fully non-blocking buddy-system, where threads performing concurrent allocations/releases do not undergo any spin-lock based synchronization. Our solution allows threads to proceed in parallel, and commit their allocations/releases unless a conflict is materialized while handling the allocator metadata. Conflict detection relies on atomic Read-Modify-Write (RMW) machine instructions. Beyond improving scalability and performance, our solution can also avoid wasting clock cycles for spin-lock operations by threads that could in principle carry out their memory allocations/releases in full concurrency. Romolo Marotta, Mauro Ianni, Andrea Scarselli, Alessandro Pellegrini 0001, Francesco Quaglia |
CCGRID | 4 |
| 2019 | An Agent-Based Simulation API for Speculative PDES Runtime EnvironmentsabstractAgent-Based Modeling and Simulation (ABMS) is an effective paradigm to model systems exhibiting complex interactions, also with the goal of studying the emergent behavior of these systems. While ABMS has been effectively used in many disciplines, many successful models are still run only sequentially. Relying on simple and easy-to-use languages such as NetLogo limits the possibility to benefit from more effective runtime paradigms, such as speculative Parallel Discrete Event Simulation (PDES). In this paper, we discuss a semantically-rich API allowing to implement Agent-Based Models in a simple and effective way. We also describe the critical points which should be taken into account to implement this API in a speculative PDES environment, to scale up simulations on distributed massively-parallel clusters. We present an experimental assessment showing how our proposal allows to implement complicated interactions with a reduced complexity, while delivering a non-negligible performance increase. Andrea Piccione, Matteo Principe, Alessandro Pellegrini 0001, Francesco Quaglia |
SIGSIM-PADS | 3 |
| 2019 | Cross-state events: A new approach to parallel discrete event simulation and its speculative runtime support
Alessandro Pellegrini 0001, Francesco Quaglia |
J. Parallel Distributed Comput. | 1 |
| 2019 | Anonymous Readers Counting: A Wait-Free Multi-Word Atomic Register Algorithm for Scalable Data Sharing on Multi-Core MachinesabstractIn this article we present Anonymous Readers Counting (ARC), a multi-word atomic (1,N) register algorithm for multi-core machines. ARC exploits Read-Modify-Write (RMW) instructions to coordinate the writer and reader threads in a wait-free manner and enables large-scale data sharing by admitting up to$(2^{32}-2)$concurrent readers on off-the-shelf 64-bit machines, as opposed to the most advanced RMW-based approach which is limited to 58 readers on the same kind of machines. Further, ARC avoids multiple copies of the register content when accessing it—this is a problem that affects classical register algorithms based on atomic read/write operations on single words. Thus it allows for higher scalability with respect to the register size. Moreover, ARC explicitly reduces the overall power consumption, via a proper limitation of RMW instructions in case of read operations re-accessing a still-valid snapshot of the register content, and by showing constant time for read operations and amortized constant time for write operations. Our proposal has therefore a strong focus on real-world off-the-shelf architectures, allowing us to capture properties which benefit both performance and power consumption. A proof of correctness of our register algorithm is also provided, together with experimental data for a comparison with literature proposals. Beyond assessing ARC on physical platforms, we carry out as well an experimentation on virtualized infrastructures, which shows the resilience of wait-free synchronization as provided by ARC with respect to CPU-steal times, proper of modern paradigms such as cloud computing. Finally, we discuss how to extend ARC for scenarios with multiple writers and multiple readers—the so called (M,N) register. This is achieved not by changing the operations (and their wait-free nature) executed along the critical path of the threads, rather only changing the ratio between the number of buffers keeping the register snapshots and the number of threads to coordinate, as well as the number of bits used for counting readers within a 64-bit mask accessed via RMW instructions—just depending on the target balance between the number of readers and the number of writers to be supported. Mauro Ianni, Alessandro Pellegrini 0001, Francesco Quaglia |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | A Non-blocking Buddy System for Scalable Memory Allocation on Multi-core MachinesabstractCommon implementations of core memory allocation components handle concurrent allocation/release requests by synchronizing threads via spin-locks. This approach is not prone to scale with large thread counts, a problem that has been addressed in the literature by introducing layered allocation services or replicating the core allocators-the bottom most ones within the layered architecture. Both these solutions tend to reduce the pressure of actual concurrent accesses to each individual core allocator. In this article we explore an alternative approach to scalability of memory allocation/release, which can be still combined with those literature proposals. We present a fully non-blocking buddy-system, that allows threads to proceed in parallel, and commit their allocations/releases unless a conflict is materialized while handling its metadata. Beyond improving scalability and performance it is resilient to performance degradation in face of concurrent accesses independently of the current level of fragmentation of the handled memory blocks. Romolo Marotta, Mauro Ianni, Andrea Scarselli, Alessandro Pellegrini 0001, Francesco Quaglia |
CLUSTER | 4 |
| 2018 | Model-Based Proactive Read-Validation in Transaction Processing SystemsabstractConcurrency control protocols based on read-validation schemes allow transactions which are doomed to abort to still run until a subsequent validation check reveals them as invalid. These late aborts do not favor the reduction of wasted computation and can penalize performance. To counteract this problem, we present an analytical model that predicts the abort probability of transactions handled via read-validation schemes. Our goal is to determine what are the suited points-along a transaction lifetime-to carry out a validation check. This may lead to early aborting doomed transactions, thus saving CPU time. We show how to exploit the abort probability predictions returned by the model in combination with a threshold-based scheme to trigger read-validations. We also show how this approach can definitely improve performance-leading up to 14 % better turnaround-as demonstrated by some experiments carried out with a port of the TPC-C benchmark to Software Transactional Memory. Simone Economo, Emiliano Silvestri, Pierangelo di Sanzo, Alessandro Pellegrini 0001, Francesco Quaglia |
ICPADS | 4 |
| 2018 | A Power Cap Oriented Time Warp ArchitectureabstractControlling power usage has become a core objective in modern computing platforms. In this article we present an innovative Time Warp architecture oriented to efficiently run parallel simulations under a power cap. Our architectural organization considers power usage as a foundational design principle, as opposed to classical power-unaware Time Warp design. We provide early experimental results showing the potential of our proposal. Stefano Conoci, Davide Cingolani, Pierangelo di Sanzo, Bruno Ciciani, Alessandro Pellegrini 0001, Francesco Quaglia |
SIGSIM-PADS | 5 |
| 2018 | The Ultimate Share-Everything PDES SystemabstractThe share-everything PDES (Parallel Discrete Event Simulation) paradigm is based on fully sharing the possibility to process any individual event across concurrent threads, rather than binding Logical Processes (LPs) and their events to threads. It allows concentrating, at any time, the computing power---the CPU-cores on board of a shared-memory machine---towards the unprocessed events that stand closest to the current commit horizon of the simulation run. This fruitfully biases the delivery of the computing power towards the hot portion of the model execution trajectory. In this article we present an innovative share-everything PDES system that provides (1) fully non-blocking coordination of the threads when accessing shared data structures and (2) fully speculative processing capabilities---Time Warp style processing---of the events. As we show via an experimental study, our proposal can cope with hard workloads where both classical Time Warp systems---based on LPs to threads binding---and previous share-everything proposals---not able to exploit fully speculative processing of the events---tend to fail in delivering adequate performance. Mauro Ianni, Romolo Marotta, Davide Cingolani, Alessandro Pellegrini 0001, Francesco Quaglia |
SIGSIM-PADS | 4 |
| 2018 | Porting Event &Cross-State Synchronization to the CloudabstractAlong the years, Parallel Discrete Event Simulation (PDES) has been enriched with programming facilities to bypass state disjointness across the concurrent Logical Processes (LPs). New supports have been proposed, offering the programmer approaches alternative to message passing to code complex LPs' relations. Along this path we find Event &Cross-State (ECS), which allows writing event handlers which can perform in-place accesses to the state of any LP, by simply relying on pointers. This programming model has been shipped with a runtime support enabling concurrent speculative execution of LPs limited to shared-memory machines. In this paper, we present the design of a middleware layer that allows ECS to be ported to distributed-memory clusters of machines. A core application of our middleware is to let ECS-coded models be hosted on top of (low-cost) resources from the Cloud. Overall, ECS-coded models no longer demand for powerful shared-memory machines to execute in reasonable time. Thanks to our solution, we retain indeed the possibility to rely on the enriched ECS programming model while still enabling deployments of PDES models on convenient (Cloud-based) infrastructures. An experimental assessment of our proposal is also provided. Matteo Principe, Tommaso Tocci, Alessandro Pellegrini 0001, Francesco Quaglia |
SIGSIM-PADS | 3 |
| 2017 | Preemptive Software Transactional MemoryabstractIn state-of-the-art Software Transactional Memory (STM) systems, threads carry out the execution of transactions as non-interruptible tasks. Hence, a thread can react to the injection of a higher priority transactional task and take care of its processing only at the end of the currently executed transaction. In this article we pursue a paradigm shift where the execution of an in-memory transaction is carried out as a preemptable task, so that a thread can start processing a higher priority transactional task before finalizing its current transaction. We achieve this goal in an application-transparent manner, by only relying on Operating System facilities we include in our preemptive STM architecture. With our approach we are able to re-evaluate CPU assignment across transactions along a same thread every few tens of microseconds. This is mandatory for an effective priority-aware architecture given the typically finer-grain nature of in-memory transactions compared to their counterpart in database systems. We integrated our preemptive STM architecture with the TinySTM package, and released it as open source. We also provide the results of an experimental assessment of our proposal based on running a port of the TPC-C benchmark to the STM environment. Emiliano Silvestri, Simone Economo, Pierangelo di Sanzo, Alessandro Pellegrini 0001, Francesco Quaglia |
CCGrid | 4 |
| 2017 | A Wait-Free Multi-word Atomic (1, N) Register for Large-Scale Data Sharing on Multi-core MachinesabstractWe present a multi-word atomic (1,N) register for multi-core machines exploiting Read-Modify-Write (RMW) instructions to coordinate the writer and the readers in a wait-free manner. Our proposal, called Anonymous Readers Counting (ARC), enables large-scale data sharing by admitting up to 2^{32}-2 concurrent readers on off-the-shelf 64-bit machines, as opposed to the most advanced RMW-based approach which is limited to 58 readers. Further, ARC avoids multiple copies of the register content while accessing it-this affects classical register's algorithms based on atomic read/write operations on single words. Thus, ARC allows for higher scalability with respect to the register size. Mauro Ianni, Alessandro Pellegrini 0001, Francesco Quaglia |
CLUSTER | 2 |
| 2017 | A non-blocking global virtual time algorithm with logarithmic number of memory operationsabstractThe increasing diffusion of shared-memory multi-core machines has given rise to a change in the design of Parallel Discrete Event Simulation (PDES) platforms. In particular, the possibility to share large amounts of memory by many worker threads has lead to a boost in the adoption of non-blocking coordination algorithms, which have been proven to offer higher scalability when compared to their blocking counterparts based on critical sections. In this article we present an innovative non-blocking algorithm for computing Global Virtual Time (GVT) - namely, the current commit horizon-in multi-thread PDES engines to be run on top of multi-core machines. Beyond being non-blocking, our proposal has the advantage of providing a logarithmic (rather than linear) number of per-thread memory operations - read/write operations of values involved in the reduction for computing the GVT value-vs the amount of threads participating in the GVT computation. This allows for keeping low the actual CPU time that is required for determining the new GVT value. We compare our algorithm with a literature solution, still based on the non-blocking approach, but entailing a linear number of memory operations, quantifying the advantages from our proposal especially for very large numbers of threads participating in the GVT computation. Mauro Ianni, Romolo Marotta, Alessandro Pellegrini 0001, Francesco Quaglia |
DS-RT | 3 |
| 2017 | Towards a fully non-blocking share-everything PDES platformabstractShared-memory multi-core platforms are changing the nature of Parallel Discrete Event Simulation (PDES) because of the possibility to fully share the workload of events to be processed across threads. In this context, one rising PDES paradigm - referred to as share-everything PDES - is no longer based on the concept of (temporary) biding of simulation objects to worker threads. Rather, each worker threads can - at any time - pick from a fully shared event pool an event to process which can be destined to whatever simulation object. While attention has been posed on the design of concurrent shared pools, allowing non-blocking parallel operations, the scenario where two (or more) threads pick events destined to the same simulation object still lacks adequate synchronization support. In fact, these events are currently sequentialized and processed in a critical section touching the simulation object state, thus leading threads to mutually block each other. In this article we present the design of a share-everything speculative PDES engine that prevents mutual thread blocks because of the access to a same object state. In our design, the non-blocking property is seen as a vertical attribute of the engine (not only of the event pool). This vertical view demands for innovative event-dispatching schemes and, at the same time, innovative interactions with (and management of) the fully-shared event pool, which are features that we embed in our innovative design. Mauro Ianni, Romolo Marotta, Alessandro Pellegrini 0001, Francesco Quaglia |
DS-RT | 3 |
| 2017 | ORCHESTRA: An asynchronous wait-free distributed GVT algorithmabstractTaking advantage of computing capabilities offered by modern parallel and distributed architectures is fundamental to run large-scale simulation models based on the Parallel Discrete Event Simulation (PDES) paradigm. By relying on this computing organization, it is possible to effectively overcome both the power and the memory wall, which are core limiting aspects to deliver high-performance simulations. This is even more the case when relying on the speculative Time Warp synchronization protocol, which could be particularly memory greedy. At the same time, some form of coordination, such as the computation of the Global Virtual Time (GVT), is required by Time Warp Systems. These coordination points could easily become the bottleneck of large-scale simulations, hindering an efficient exploitation of the computing power offered by large supercomputing facilities. In this paper we present ORCHESTRA, a coordination algorithm which is both wait-free and asynchronous. The nature of this algorithm allows any computing node to carry on simulation activities while the global agreement is reached, thus offering an effective building block to achieve scalable PDES. We claim that the general organization of ORCHESTRA could be adopted by different high-performance computing applications, thus paving the way to a more effective usage of modern computing infrastructures. Tommaso Tocci, Alessandro Pellegrini 0001, Francesco Quaglia, Josep Casanovas, Toyotaro Suzumura |
DS-RT | 2 |
| 2017 | Machine learning-based management of cloud applications in hybrid clouds: A Hadoop case studyabstractThis paper illustrates the effort to integrate a machine learning-based framework which can predict the remaining time to failure of computing nodes with Hadoop applications. This work is part of a larger effort targeting the development of a cloud-oriented autonomic framework to increase the availability of applications subject to software anomalies, and to jointly improve their performance. The framework uses machine-learning, software rejuvenation, and load distribution techniques to proactively prevent failures. We believe that this work allows to set a possible path towards the definition of best practices for the development of systems to support autonomic management of cloud applications, illustrating what are the issues that should be addressed by the research community. Indeed, given the scale and the complexity of modern computing infrastructures, effective autonomic management approaches of cloud applications are becoming mandatory. Dimiter R. Avresky, Alessandro Pellegrini 0001, Pierangelo di Sanzo |
NCA | 2 |
| 2017 | Prompt application-transparent transaction revalidation in software transactional memoryabstractSoftware Transactional Memory (STM) allows encapsulating shared-data accesses within transactions, executed with atomicity and isolation guarantees. The assessment of the consistency of a running transaction is performed by the STM layer at specific points of its execution, such as when a read or write access to a shared object occurs, or upon a commit attempt. However, performance and energy efficiency issues may arise when no shared-data read/write operation occurs for a while along a thread running a transaction. In this scenario, the STM layer may not regain control for a considerable amount of time, thus not being able to early detect if such transaction has become inconsistent in the meantime. To tackle this problem we present an STM architecture that, thanks to a lightweight operating system support, is able to perform a fine-grain periodic (hence prompt) revalidation of running transactions. Our proposal targets Linux and x86 systems and has been integrated with the open source TinySTM package. Experimental results with a port of the TPC-C benchmark to STM environments show the effectiveness of our solution. Simone Economo, Emiliano Silvestri, Pierangelo di Sanzo, Alessandro Pellegrini 0001, Francesco Quaglia |
NCA | 4 |
| 2017 | Dealing with Reversibility of Shared Libraries in PDESabstractState recoverability is a crucial aspect of speculative Time Warp-based Parallel Discrete Event Simulation. In the literature, we can identify three major classes of techniques to support the correct restoration of a previous simulation state upon the execution of a rollback operation: state checkpointing/restore, manual reverse computation and automatic reverse computation. The latter class has been recently supported by relying either on binary code instrumentation or on source-to-source code transformation. Nevertheless, both solutions are not intrinsically meant to support a reversible execution of third-party shared libraries, which can be pretty useful when implementing complex simulation models. Davide Cingolani, Alessandro Pellegrini 0001, Markus Schordan, Francesco Quaglia, David R. Jefferson |
SIGSIM-PADS | 2 |
| 2017 | A Conflict-Resilient Lock-Free Calendar Queue for Scalable Share-Everything PDES PlatformsabstractEmerging share-everything Parallel Discrete Event Simulation (PDES) platforms rely on worker threads fully sharing the workload of events to be processed. These platforms require efficient event pool data structures enabling high concurrency of extraction/insertion operations. Non-blocking event pool algorithms are raising as promising solutions for this problem. However, the classical non-blocking paradigm leads concurrent conflicting operations, acting on a same portion of the event pool data structure, to abort and then retry. In this article we present a conflict-resilient non-blocking calendar queue that enables conflicting dequeue operations, concurrently attempting to extract the minimum element, to survive, thus improving the level of scalability of accesses to the hot portion of the data structure---namely the bucket to which the current locality of the events to be processed is bound. We have integrated our solution within an open source share-everything PDES platform and report the results of an experimental analysis of the proposed concurrent data structure compared to some literature solutions. Romolo Marotta, Mauro Ianni, Alessandro Pellegrini 0001, Francesco Quaglia |
SIGSIM-PADS | 3 |
| 2016 | OS-Based NUMA Optimization: Tackling the Case of Truly Multi-thread Applications with Non-partitioned Virtual Page AccessesabstractA common approach to improve memory access in NUMA machines exploits operating system (OS) page protection mechanisms to induce faults to determine which pages are accessed by what thread, so as to move the thread and its working-set of pages to the same NUMA node. However, existing proposals do not fully fit the requirements of truly multi-thread applications with non-partitioned accesses to virtual pages. In fact, these proposals exploit (induced) faults on a same page-table for all the threads of a same process to determine the access pattern. Hence, the fault by one thread (and the consequent re-opening of the access to the corresponding page) would mask those by other threads on the same page. This may lead to inaccuracy in the estimation of the working-set of individual threads. We overcome this drawback by presenting a lightweight operating system support for Linux, referred to as multi-view address space, explicitly targeting accuracy of per-thread working-set estimation in truly multi-thread applications with non-partitioned accesses, and an associated thread/data migration policy. Our solution is fully transparent to user-space code. It is embedded in a Linux/x86_64 module that installs any required modification to the original kernel image by solely relying on dynamic patching. A motivated case study in the context of HPC is also presented for an assessment of our proposal. Ilaria Di Gennaro, Alessandro Pellegrini 0001, Francesco Quaglia |
CCGrid | 2 |
| 2016 | A Lock-Free O(1) Event Pool and Its Application to Share-Everything PDES PlatformsabstractThe large diffusion of highly-parallel shared-memory multi-core machines has led Parallel Discrete Event Simulation (PDES) platforms to a shift towards a share-everything model. This model is based on loose coupling between simulation objects and threads, lasting (as an extreme) no more than the lifetime of individual events. Concurrent threads can therefore CPU-dispatch events destined to any object at any point in time, thus fully sharing the workload of events to be processed on a fine grain basis. This demands for efficient mechanisms to share the overall pool of pending events by enabling parallelism in insertion and extraction operations. In this article we present a lock-free event pool which also provides amortized O(1) time complexity for both insertions and extractions. It can sustain highly concurrent accesses, while not leading to noticeable performance degradation when scaling up the thread count. Experimental results demonstrate that our solution stands as a core facility capable of further raising up the pragmatical impact of such an emerging share-everything PDES paradigm. Romolo Marotta, Mauro Ianni, Alessandro Pellegrini 0001, Francesco Quaglia |
DS-RT | 3 |
| 2016 | Configurable and Efficient Memory Access Tracing via Selective Expression-Based x86 Binary InstrumentationabstractMemory access tracing is a program analysis technique with many different applications, ranging from architectural simulation to (on-line) data placement optimization and security enforcement. In this article we propose a memory access tracing approach based on static x86 binary instrumentation. Unlike non-selective schemes, which instrument all the memory access instructions, our proposal selectively instruments a subset of those instructions that are the most (or fully) representative of the actual memory access pattern. The selection of the memory access instructions to be instrumented is based on a new method, which clusters instructions on the basis of their compile/link-time observable address expressions and selects representatives of these clusters. This allows for reducing the runtime cost for running instrumented code, while still enabling high accuracy in the determination of memory accesses. The trade-off between overhead and precision of the tracing process is user-tunable, so that it can be set depending on the final objective of memory access tracing (say on-line vs off-line exploitation). Additionally, our approach can track memory access at different granularity (e.g., virtual-pages or cache line-sized buffers), thus having applications in a variety of different contexts. The effectiveness of our proposal is demonstrated via experiments with applications taken from the PARSEC benchmark suite. Simone Economo, Davide Cingolani, Alessandro Pellegrini 0001, Francesco Quaglia |
MASCOTS | 3 |
| 2016 | Message from the program chairsabstractIt is with great pleasure that we welcome you to the 15thedition of IEEE NCA. Over the years, NCA has become a successful series of conferences that serves as a large international forum for presenting and sharing recent research results and technological developments in the fields of Network and Cloud Computing. This edition of NCA features a lively, interesting, and stimulating program with a lot of opportunities for discussing new results, on-going projects, and the future trend in our fields. Aris Gkoulalas-Divanis, Alessandro Pellegrini 0001, Pierangelo di Sanzo |
NCA | 2 |
| 2016 | Granular Time Warp ObjectsabstractA recent trend has shown the relevance of PDES paradigms where simulation objects are no longer seen as fully disjoint entities only interacting via events' scheduling. Particularly, mutual cross-state access (as a form of state sharing) can represent an approach enabling the simplification of the programmer's job. In this article, we present a multi-core oriented Time Warp platform supporting so called granular objects, where cross-state access is transparently enabled jointly with the dynamic clustering (granulation) of objects into groups depending on the volume of mutual state accesses along phases of the model execution. Each group represents an island where activities are sequentially dispatched in timestamp order. Concurrency is still preserved by enabling the optimistic execution of the different islands. Granulated objects do not pay synchronization costs due to mutual causal inconsistencies. Also, the underlying Time Warp platform does not pay memory management (e.g. memory access tracing) overheads to determine that mutual accesses are taking place within a group. Overall, the platform transparently (and dynamically) determines a well-suited granulation of the overall model state, and a corresponding level of concurrency, depending on the actual state access pattern by the simulation code. As far as we know, this is the first study where the problem of clustering Time Warp simulation objects is addressed for the case of in-place cross-object state accesses by the application code, and where dynamic granulation of multiple objects in a larger one is supported in a fully transparent manner. We integrated our proposal in the open source ROOT-Sim platform. Nazzareno Marziale, Francesco Nobilia, Alessandro Pellegrini 0001, Francesco Quaglia |
SIGSIM-PADS | 3 |
| 2016 | Mixing Hardware and Software Reversibility for Speculative Parallel Discrete Event Simulation
Davide Cingolani, Mauro Ianni, Alessandro Pellegrini 0001, Francesco Quaglia |
RC | 3 |
| 2015 | Hardware-Transactional-Memory Based Speculative Parallel Discrete Event Simulation of Very Fine Grain ModelsabstractThis article presents an innovative runtime support for speculative parallel processing of discrete event simulation models on multi-core architectures, which exploits Hardware-Transactional-Memory (HTM) facilities for the purpose of state recoverability. In this proposal, the speculative updates on the state of the simulation model are executed as concurrent HTM-based transactions that are also in charge of detecting whether the update is consistent with the advancement of logical-time along model execution. Our proposal is fully transparent to the application code. Hence, our HTM-based run-time support can host conventionally developed discrete event models relying on the concept of event-handlers to be dispatched by an underlying simulation engine. Experimental data show that our proposal provides 75% to 92% of the ideal speedup on an Intel Haswell based platform (equipped with 4 physical cores and HTM support) for discrete event models with event granularity ranging between 2 and 12 microseconds. The data also show that these same models cannot be executed efficiently on top of a last generation parallel discrete event simulation platform employing software-based recoverability. Emanuele Santini, Mauro Ianni, Alessandro Pellegrini 0001, Francesco Quaglia |
HiPC | 3 |
| 2015 | Proactive Scalability and Management of Resources in Hybrid Clouds via Machine LearningabstractIn this paper, we present a novel framework for supporting the management and optimization of application subject to software anomalies and deployed on large scale cloud architectures, composed of different geographically distributed cloud regions. The framework uses machine learning models for predicting failures caused by accumulation of anomalies. It introduces a novel workload balancing approach and a proactive system scale up/scale down technique. We developed a prototype of the framework and present some experiments for validating the applicability of the proposed approaches. Dimiter R. Avresky, Pierangelo di Sanzo, Alessandro Pellegrini 0001, Bruno Ciciani, Luca Forte |
NCA | 3 |
| 2015 | Transparently Mixing Undo Logs and Software Reversibility for State Recovery in Optimistic PDESabstractThe rollback operation is a fundamental building block to support the correct execution of a speculative Time Warp-based Parallel Discrete Event Simulation. In the literature, several solutions to reduce the execution cost of this operation have been proposed, either based on the creation of a checkpoint of previous simulation state images, or on the execution of negative copies of simulation events which are able to undo the updates on the state. In this paper, we explore the practical design and implementation of a state recoverability technique which allows to restore a previous simulation state either relying on checkpointing or on the reverse execution of the state updates occurred while processing events in forward mode. Differently from other proposals, we address the issue of executing backward updates in a fully-transparent and event granularity-independent way, by relying on static software instrumentation (targeting the x86 architecture and Linux systems) to generate at runtime reverse update code blocks (not to be confused with reverse events, proper of the reverse computing approach). These are able to undo the effects of a forward execution while minimizing the cost of the undo operation. We also present experimental results related to our implementation, which is released as free software and fully integrated into the open source ROOT-Sim (ROme OpTimistic Simulator) package. The experimental data support the viability and effectiveness of our proposal. Davide Cingolani, Alessandro Pellegrini 0001, Francesco Quaglia |
SIGSIM-PADS | 2 |
| 2015 | Time-Sharing Time Warp via Lightweight Operating System SupportabstractThe order according to which the different tasks are carried out within a Time Warp platform has a direct impact on performance, given that event processing is speculative, thus being subject to the possibility of being rolled-back. It is typically recognized that not-yet-executed events having lower timestamps should be given higher CPU-schedule priority, since this contributes to keep low the amount of rollbacks. However, common Time Warp platforms usually execute events as atomic actions. Hence control is bounced back to the underlying simulation platform only at the end of the current event processing routine. In other words, CPU-scheduling of events resembles classical batch-multitasking scheduling, which is recognized not to promptly react to variations of the priority of pending tasks (e.g. associated with the injection of new events in the system). In this article we present the design and implementation of a time-sharing Time Warp platform, to be run on multi-core machines, where the platform-level software is allowed to take back control on a periodical basis (with fine grain period), and to possibly preempt any ongoing event processing activity in favor of dispatching (along the same thread) any other event that is revealed to have higher priority. Our proposal is based on an ad-hoc kernel module for Linux, which implements a fine grain timer-interrupt mechanism with lightweight management, which is fully integrated with the modern top/bottom-half timer-interrupt Linux architecture, and which does not induce any bias in terms of relative CPU-usage planning across Time Warp vs non-Time Warp threads running on the machine. Our time-sharing architecture has been integrated within the open source ROOT-Sim optimistic simulation package, and we also report some experimental data for an assessment of our proposal. Alessandro Pellegrini 0001, Francesco Quaglia |
SIGSIM-PADS | 1 |
| 2015 | NUMA Time WarpabstractIt is well known that Time Warp may suffer from large usage of memory, which may hamper the efficiency of the memory hierarchy. To cope with this issue, several approaches have been devised, mostly based on the reduction of the amount of used virtual memory, e.g., by the avoidance of checkpointing and the exploitation of reverse computing. In this article we present an orthogonal solution aimed at optimizing the latency for memory access operations when running Time Warp systems on Non-Uniform Memory Access (NUMA) multi-processor/multi-core computing systems. More in detail, we provide an innovative Linux-based architecture allowing per simulation-object management of memory segments made up by disjoint sets of pages, and supporting both static and dynamic binding of the memory pages reserved for an individual object to the different NUMA nodes, depending on what worker thread is in charge of running that simulation object along a given wall-clock-time window. Our proposal not only manages the virtual pages used for the live state image of the simulation object, rather, it also copes with memory pages destined to keep the simulation object's event buffers and any recoverability data. Further, the architecture allows memory access optimization for data (messages) exchanged across the different simulation objects running on the NUMA machine. Our proposal is fully transparent to the application code, thus operating in a seamless manner. Also, a free software release of our NUMA memory manager for Time Warp has been made available within the open source ROOT-Sim simulation platform. Experimental data for an assessment of our innovative proposal are also provided in this article. Alessandro Pellegrini 0001, Francesco Quaglia |
SIGSIM-PADS | 1 |
| 2015 | Autonomic State Management for Optimistic Simulation PlatformsabstractWe present the design and implementation of an autonomic state manager (ASM) tailored for integration within optimistic parallel discrete event simulation (PDES) environments based on the C programming language and the executable and linkable format (ELF), and developed for execution on ×86_64 architectures. With ASM, the state of any logical process (LP), namely the individual (concurrent) simulation unit being part of the simulation model, is allowed to be scattered on dynamically allocated memory chunks managed via standard API (e.g., malloc/free). Also, the application programmer is not required to provide any serialization/ deserialization module in order to take a checkpoint of the LP state, or to restore it in case a causality error occurs during the optimistic run, or to provide indications on which portions of the state are updated by event processing, so to allow incremental checkpointing. All these tasks are handled by ASM in a fully transparent manner via (A) runtime identification (with chunk-level granularity) of the memory map associated with the LP state, and (B) runtime tracking of the memory updates occurring within chunks belonging to the dynamic memory map. The co-existence of the incremental and non-incremental log/restore modes is achieved via dual versions of the same application code, transparently generated by ASM via compile/link time facilities. Also, the dynamic selection of the best suited log/ restore mode is actuated by ASM on the basis of an innovative modeling/optimization approach which takes into account stability of each operating mode with respect to variations of the model/environmental execution parameters. Alessandro Pellegrini 0001, Roberto Vitali, Francesco Quaglia |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | Transparent multi-core speculative parallelization of DES models with event and cross-state dependenciesabstractIn this article we tackle transparent parallelization of Discrete Event Simulation (DES) models to be run on top of multi-core machines according to speculative schemes. The innovation in our proposal lies in that we consider a more general programming and execution model, compared to the one targeted by state of the art PDES platforms, where the boundaries of the state portion accessible while processing an event at a specific simulation object do not limit access to the actual object state, or to shared global variables. Rather, the simulation object is allowed to access (and alter) the state of any other object, thus causing what we term cross-state dependency. We note that this model exactly complies with typical (easy to manage) sequential-style DES programming, where a (dynamically-allocated) state portion of object A can be accessed by object B in either read or write mode (or both) by, e.g., passing a pointer to B as the payload of a scheduled simulation event. However, while read/write memory accesses performed in the sequential run are always guaranteed to observe (and to give rise to) a consistent snapshot of the state of the simulation model, consistency is not automatically guaranteed in case of parallelization and concurrent execution of simulation objects with cross-state dependencies. We cope with such a consistency issue, and its application-transparent support, in the context of parallel and optimistic executions. This is achieved by introducing an advanced memory management architecture, able to efficiently detect read/write accesses by concurrent objects to whichever object state in an application transparent manner, together with advanced synchronization mechanisms providing the advantage of exploiting parallelism in the underlying multi-core architecture while transparently handling both cross-state and traditional event-based dependencies. Our proposal targets Linux and has been integrated with the ROOT-Sim open source optimistic simulation platform, although its design principles, and most parts of the developed software, are of general relevance. Alessandro Pellegrini 0001, Francesco Quaglia |
SIGSIM-PADS | 1 |
| 2014 | Wait-Free Global Virtual Time Computation in Shared Memory TimeWarp SystemsabstractGlobal Virtual Time (GVT) is a powerful abstraction used to discriminate what events belong (and what do not belong) to the past history of a parallel/distributed computation. For high performance simulation systems based on the Time Warp synchronization protocol, where concurrent simulation objects are allowed to process their events speculatively and causal consistency is achieved via rollback/recovery techniques, GVT is used to determine which portion of the simulation can be considered as committed. Hence it is the base for actuating memory recovery (e.g. of obsolete logs that were taken in order to support state recoverability) and nonrevocable operations (e.g. I/O). For shared memory implementations of simulation platforms based on the Time Warp protocol, the reference GVT algorithm is the one presented by Fujimoto and Hybinette [1]. However, this algorithm relies on critical sections that make it non-wait-free, and which can hamper scalability. In this article we present a waitfree shared memory GVT algorithm that requires no critical section. Rather, correct coordination across the processes while computing the GVT value is achieved via memory atomic operations, namely compare-and-swap. The price paid by our proposal is an increase in the number of GVT computation phases, as opposed to the single phase required by the proposal in [1]. However, as we show via the results of an experimental study, the wait-free nature of the phases carried out in our GVT algorithm pays-off in reducing the actual cost incurred by the proposal in [1]. Alessandro Pellegrini 0001, Francesco Quaglia |
SBAC-PAD | 1 |
| 2013 | Transparent Support for Partial Rollback in Software Transactional Memories
Alice Porfirio, Alessandro Pellegrini 0001, Pierangelo di Sanzo, Francesco Quaglia |
Euro-Par | 2 |
| 2013 | Consistent and efficient output-streams management in optimistic simulation platformsabstractOptimistic synchronization is considered an effective means for supporting Parallel Discrete Event Simulations. It relies on a speculative approach, where concurrent processes execute simulation events regardless of their safety, and consistency is ensured via proper rollback mechanisms, upon the a-posteriori detection of causal inconsistencies along the events' execution path. Interactions with the outside world (e.g. generation of output streams) are a well-known problem for rollback-based systems, since the outside world may have no notion of rollback. In this context, approaches for allowing the simulation modeler to generate consistent output rely on either the usage of ad-hoc APIs (which must be provided by the underlying simulation kernel) or temporary suspension of processing activities in order to wait for the final outcome (commit/rollback) associated with a speculatively-produced output. Francesco Antonacci, Alessandro Pellegrini 0001, Francesco Quaglia |
SIGSIM-PADS | 2 |
| 2012 | A load-sharing architecture for high performance optimistic simulations on multi-core machinesabstractIn Parallel Discrete Event Simulation (PDES), the simulation model is partitioned into a set of distinct Logical Processes (LPs) which are allowed to concurrently execute simulation events. In this work we present an innovative approach to load-sharing on multi-core/multiprocessor machines, targeted at the optimistic PDES paradigm, where LPs are speculatively allowed to process simulation events with no preventive verification of causal consistency, and actual consistency violations (if any) are recovered via rollback techniques. In our approach, each simulation kernel instance, in charge of hosting and executing a specific set of LPs, runs a set of worker threads, which can be dynamically activated/deactivated on the basis of a distributed algorithm. The latter relies in turn on an analytical model that provides indications on how to reassign processor/core usage across the kernels in order to handle the simulation workload as efficiently as possible. We also present a real implementation of our load-sharing architecture within the ROme OpTimistic Simulator (ROOT-Sim), namely an open-source C-based simulation platform implemented according to the PDES paradigm and the optimistic synchronization approach. Experimental results for an assessment of the validity of our proposal are presented as well. Roberto Vitali, Alessandro Pellegrini 0001, Francesco Quaglia |
HiPC | 2 |
| 2012 | Transparent and Efficient Shared-State Management for Optimistic Simulations on Multi-core MachinesabstractTraditionally, Logical Processes (LPs) forming a simulation model store their execution information into disjoint simulations states, forcing events exchange to communicate data between each other. In this work we propose the design and implementation of an extension to the traditional Time Warp (optimistic) synchronization protocol for parallel/distributed simulation, targeted at shared-memory/multicore machines, allowing LPs to share parts of their simulation states by using global variables. In order to preserve optimism's intrinsic properties, global variables are transparently mapped to multi-version ones, so to avoid any form of safety predicate verification upon updates. Execution's consistency is ensured via the introduction of a new rollback scheme which is triggered upon the detection of an incorrect global variable's read. At the same time, efficiency in the execution is guaranteed by the exploitation of non-blocking algorithms in order to manage the multi-version variables' lists. Furthermore, our proposal is integrated with the simulation model's code through software instrumentation, in order to allow the application-level programmer to avoid using any specific API to mark or to inform the simulation kernel of updates to global variables. Thus we support full transparency. An assessment of our proposal, comparing it with a traditional message-passing implementation of variables' multi-version is provided as well. Alessandro Pellegrini 0001, Roberto Vitali, Sebastiano Peluso, Francesco Quaglia |
MASCOTS | 1 |
| 2010 | Autonomic Log/Restore for Advanced Optimistic Simulation SystemsabstractIn this paper we address state recoverability in optimistic simulation systems by presenting an autonomic log/restore architecture. Our proposal is unique in that it jointly provides the following features: (i) log/restore operations are carried out in a completely transparent manner to the application programmer, (ii) the simulation-object state can be scattered across dynamically allocated non-contiguous memory chunks, (iii) two differentiated operating modes, incremental vs non-incremental, coexist via transparent, optimized run-time management of dual versions of the same application layer, with dynamic selection of the best suited operating mode in different phases of the optimistic simulation run, and (iv) determination of the best suited mode for any time frame is carried out on the basis of an innovative modeling/optimization approach that takes into account stability of each operating mode vs variations of the model execution parameters. Roberto Vitali, Alessandro Pellegrini 0001, Francesco Quaglia |
MASCOTS | 2 |
| 2009 | Benchmarking Memory Management Capabilities within ROOT-SimabstractIn parallel discrete event simulation techniques, the simulation model is partitioned into objects, concurrently executing events on different CPUs and/or multiple CPU-Cores.In such a context, run-time supports for logical time synchronization across the different simulation objects play a central role in determining the effectiveness of the specific parallel simulation environment. In this paper we present an experimental evaluation of the memory management capabilities offered by the ROme OpTimistic Simulator (ROOT-Sim). This is an open source parallel simulation environment transparently supporting optimistic synchronization via recoverability (based on incremental log/restore techniques) of any type of memory operation affecting the state of simulation objects, i.e., memory allocation, deallocation and update operations. The experimental study is based on a synthetic benchmark which mimics different read/write patterns inside the dynamic memory map associated with the state of simulation objects. This allows sensibility analysis of time and space effects due to the memory management subsystem while varying the type and the locality of the accesses associated with event processing. Roberto Vitali, Alessandro Pellegrini 0001, Francesco Quaglia |
DS-RT | 2 |