VLDB 2026 Research / reviewers in the wild / expert
Romolo Marotta
dblp:129/1238
· DBLP profile ↗
29ranked-venue papers
15as first author
17since 2021 · last 2025
0000-0001-7589-9274ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 10 · 5 first-author · 6 since 2021Systems, architecture and hardware · 8 · 5 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Bootstrapping Technique for Reducing the Costs of Machine Learning Models for Predicting Execution Times in IaaS CloudsabstractMachine Learning (ML) emerged as a powerful tool for predicting task execution times across the variety of VM types offered by Infrastructure-as-a-Service (IaaS) clouds. However, training ML models to ensure accurate predictions can often become uneconomical for users due to the high costs—in terms of both time and money—for collecting samples, especially when an IaaS cloud offers a wide choice of VM types. This paper investigates a ML model bootstrapping technique that leverages analytical modeling to reduce the cost of collecting training samples while maintaining robust performance predictions. Complementarily, the technique can be used to improve the accuracy of ML models in the case of limited availability of training samples. Experimental results highlighted the potential of the proposed technique with various workloads and with a large set of VM types, paving the way for more cost-effective ML-based performance prediction in IaaS clouds. Romolo Marotta, Gabriele Russo Russo, Francesco Quaglia, Pierangelo di Sanzo |
SoCC | 1 |
| 2025 | Model-Driven Parallel and Distributed Stochastic Simulation of Chemical Reaction Networks
Simone Bauco, Federica Montesano, Adriano Pimpini, Romolo Marotta, Alessandro Pellegrini 0001 |
DS-RT | 4 |
| 2025 | Longer (Not Longest) Processing - Time First in Constant Global Lookahead PDES Engines
Romolo Marotta, Dissan Uddin Ahmed, Francesco Quaglia |
DS-RT | 1 |
| 2025 | Comparing the Run-Time Behavior of Modern PDES Engines on PowerPC and x86 Architectures
Romolo Marotta, Francesco Quaglia |
DS-RT | 1 |
| 2025 | DESL: A Literate Programming Language Framework for Interoperable Parallel Discrete Event SimulationabstractSimulation is indispensable for advanced scientific research, enabling accurate explorations of complex phenomena and supporting evidence-based decision-making across interdisciplinary boundaries. Parallel Discrete Event Simulation (PDES) provides substantial advantages in modelling large-scale systems by distributing computational tasks among multiple processors, enhancing scalability. However, exploiting it is extremely challenging due to obstacles in model efficiency, concurrency control, reproducibility, and maintainability. Furthermore, the large number of available PDES run-time environments makes it difficult to explore their (performance) capabilities for some specific model, hindering the identification of the best-suited technology for a certain simulation study. To address these limitations, we introduce a unified framework grounded in literate programming and model-driven engineering, integrating interwoven documentation and model logic within a single source. This design enhances intrinsic consistency between model logic and explanatory content, while enabling the generation of model implementations tailored to multiple runtime environments, thus allowing simulationists to focus on model development without being locked in to any specific technology or environment. This facilitates model reuse and performance comparisons across diverse execution environments. We show the viability of this approach by providing the first-ever experimental comparison across three different simulators, starting from the same model implementation. Simone Bauco, Romolo Marotta, Alessandro Pellegrini 0001 |
SIGSIM-PADS | 2 |
| 2025 | Spin/Sleep Proactive-Awakening Locks for Alternative Performance/Energy Trade-OffsabstractABSTRACT Locking plays a crucial role since it ensures synchronized access by concurrent threads to shared resources—like shared data structures to be managed in critical sections. Traditional sleep locks—based on blocking operating system services—adopt a reactive approach (e.g., upon lock release) to waking up waiting threads, which might introduce additional latency on the critical path. On the opposite side, non‐blocking locks, like spinlocks, allow threads to wait while still using CPU cycles for checking and updating the lock variable, which causes the waste of both cycles and energy. In this article, we present a new locking algorithm, called SSPA (Spin/Sleep Proactive‐Awakening)—and its implementation for Linux systems—which combines spin and sleep waiting phases via the introduction of an innovative proactive wake‐up mechanism that exploits the SoftIRQ daemon of the Linux kernel. Our solution allows threads to be awakened from their sleep phases on time to be already CPU dispatched when the lock is really released. This provides the opportunity to quickly access the critical section while at the same time enabling control over the actual amount of CPU cycles that are spent by spinning wait phases. As we show via experimental data, our solution allows exploring new trade‐offs between responsiveness and CPU/energy efficiency in concurrent applications, hence rising as an interesting alternative to literature solutions. Matteo Federico, Romolo Marotta, Francesco Quaglia |
Concurr. Comput. Pract. Exp. | 2 |
| 2024 | Sampling Policies for Near-Optimal Device Choice in Parallel Simulations on CPU/GPU PlatformsabstractHeterogeneous hardware platforms comprised of CPUs, GPUs, and other accelerators offer the opportunity to choose the best-suited device for executing a given scientific simulation in order to minimize execution time and energy consumption. To this end, the recently proposed "Follow the Leader" approach dynamically selects a suitable device based on runtime performance measurements during speculative discrete-event simulations. A currently active "leader" device is periodically challenged by a "follower" device in order to negotiate the new leader. The optimality of the device choices and the associated overhead depends critically on the challenge frequency and timing. Here, we explore policies to schedule challenges with the goal of attaining Pareto-optimal combinations of execution time and energy consumption. Several heuristics are first evaluated in an abstract fashion using a "meta-simulation" by mimicking the progress and energy consumption of an idealized co-execution. In this setting, we optimize the heuristics’ tuning parameters to assess their relative merits in near-optimal configurations when compared to challenge timings based on perfect knowledge. We find that under challenging stochastic workloads based on a class of mean-reverting random walks, the best heuristics can closely approximate the execution time and energy consumption achievable under an optimal device choice. Empirical support for this observation is given by measurements of a CPU/GPU co-execution of the Time Warp algorithm on physical hardware. Philipp Andelfinger, Alessandro Pellegrini 0001, Romolo Marotta |
DS-RT | 3 |
| 2024 | Out-of-Order Discrete Event Simulation: Fighting Memory Boundedness while Running DES ModelsabstractIn this article we present Out-of-order Discrete Event Simulation (ODES), a solution for sequential style execution of DES models not following timestamp order. ODES ensures anyway the same identical simulation results as timestamp ordered execution, thanks to fully correct maintenance of the simulation model data flow. At the same time, it drastically reduces the memory boundedness—namely, the impact of cache misses and of stale CPU cycles—on the simulation model execution speed. Beyond presenting foundational concepts, we also discuss our ODES-engine implementation, based on the c programming language. Additionally, we report experimental data for comparing ODES with the classical timestamp ordered execution of simulation models according to conventional sequential simulation. The relevance of ODES compared to classical timestamp ordered sequential DES not only stands in its benefits on performance, rather ODES can also assume the role of new reference for determining the speedup achievable via parallel/distributed discrete event simulation systems, compared to the single thread execution. Also, thanks to its improvements in the interaction with RAM, ODES constitutes a new framework for effective parallel replication of simulation experiments on multi-processor/multi-core machines. Romolo Marotta, Francesco Quaglia |
DS-RT | 1 |
| 2024 | Lightweight Operating System Services for Incremental Checkpointing in Speculative Discrete Event Simulation on Linux PlatformsabstractOne way for supporting incremental checkpointing is the exploitation of classical memory protection services—in particular the mprotect (…) system call offered by Posix compliant operating systems—for intercepting memory-writes and identifying dirty pages in the address space. However, this solution involves Inter-Processor-Interrupt (IPI) and the associated handling mechanisms, which show costs that increase when scaling up the level of parallelism in the underlying hardware architecture. For HPC contexts like speculative (aka optimistic) parallel discrete event simulation, where checkpointing and state restore massively take place in order to maintain causality among the concurrent simulation objects, these costs may impact performance in a non-negligible manner. In this work, we present the design of operating system services that enable write-protection and the tracing of dirtied pages via per CPU setup of the MMU (Memory Management Unit), completely avoiding the usage of IPI and their management. Hence, we provide a solution where the cost for setting up the incremental checkpointing support of a simulation object processed by a specific thread at a given time is definitely limited, compared to the aforementioned classical case. Our design has been devised for Linux, although it can be ported to other operating systems, and has been integrated for its testing in the USE (Ultimate Share Everything) open source parallel simulation environment. Federica Montesano, Romolo Marotta, Francesco Quaglia |
ICPADS | 2 |
| 2024 | Follow the Leader: Alternating CPU/GPU Computations in PDESabstractDespite the successes of graphics processing units (GPUs) in accelerating simulations in several research fields, their use is largely restricted to domain-specific workloads that consistently offer the large degree of inherent parallelism and computational intensity at which GPUs excel. When targeting generic discrete-event simulations, whose dynamics can vary wildly over time, a static choice between a GPU-based and traditional CPU-based execution is likely to be suboptimal. Here, we explore a parallel discrete-event (PDES) execution scheme for CPU-GPU platforms that aims to approximate an optimal dynamic device choice. Starting from an intermediate model state, a current “leader” device running the simulation is periodically challenged by a brief concurrent run on another device starting from an intermediate model state. Based on the gathered performance measurements, a forecasting scheme determines the leader for the next period. The execution time and power consumption of this scheme hinge on 1) an efficient mechanism for providing the “follower” device with a consistent model state, and 2) robust performance forecasting to justify the device choices. We present these building blocks, their implementation combining the existing CPU and GPU simulators ROOT-Sim and GPUTW, and measurement results demonstrating substantially reduced execution time without increasing energy consumption over a static device choice. Romolo Marotta, Alessandro Pellegrini 0001, Philipp Andelfinger |
SIGSIM-PADS | 1 |
| 2023 | Incremental Checkpointing of Large State Simulation Models with Write-Intensive Events via Memory Update Correlation on Buddy PagesabstractCheckpointing techniques for speculative parallel simulation of discrete event models have been widely studied in the literature. However, there has been a very marginal attempt to exploit operating system page-protection services, which have instead been largely exploited in the context of checkpointing for fault tolerance. In this article, we discuss how these services can effectively manage simulation models with large states and write-intensive events in zones of the state layout. In particular, we present a solution where the correlation of write operations on buddy pages in the state layout can be exploited to achieve effective incremental checkpointing support, which allows scaling down the costs of operating system services. Our solution does not require any instrumentation of the simulation application code and is usable on any Posix-compliant operating system. We also discuss its integration within the USE (Ultimate-Share-Everything) open-source speculative simulation package and report some experimental data for its assessment. Romolo Marotta, Federica Montesano, Alessandro Pellegrini 0001, Francesco Quaglia |
DS-RT | 1 |
| 2023 | RCR Report of "Zero Lookahead? Zero Problem. The Window Racer Algorithm"abstractThe artifact evaluated in this report is relevant to the paper “Zero Lookahead? Zero Problem. The Window Racer Algorithm”. In fact, it allows to run the experiments, reproduce figures and tables. Dependencies are well documented. The process to regenerate data presented in the article completes correctly, and the results are reproducible. Additionally, the authors have uploaded their artifact on permanent repositories, which ensures a long-term retention. Thus, this paper can receive the Artifacts Available, Artifacts Evaluated—Functional, and Results Reproduced badges. Romolo Marotta |
SIGSIM-PADS | 1 |
| 2023 | Effective Access to the Committed Global State in Speculative Parallel Discrete Event Simulation on Multi-core MachinesabstractOutput production and predicate detection are critical in speculative parallel discrete event simulation, since they need to take place accessing past state values—which have become committed—rather than the current state of the simulation objects, which is possibly affected by causality errors related to speculative event processing. In this article, we present an architecture that enables an effective management of the access to the committed state of any simulation object while still guaranteeing: (i) minimal impact on the forward execution of the simulation in terms of synchronization (and rollback generation) and (ii) highly balanced distribution of the tasks among all the threads running the simulation application. Our architecture is devised for speculative simulation engines running on top of shared-memory parallel machines, where worker threads full share the simulation workload. We exploit kernel-level facilities—targeting the Linux operating system—and user level ones, which work together for enabling a suited wall-clock-time collocation of the threads’ activities for the access to the committed global state of the simulation. We integrated our proposal within the USE (Ultimate Share-Everything) open-source simulation platform, and provide an experimental assessment of it. Romolo Marotta, Federica Montesano, Francesco Quaglia |
SIGSIM-PADS | 1 |
| 2023 | Strategies and software support for the management of hardware performance countersabstractAbstract Hardware performance counters (HPCs) are facilities offered by most off‐the‐shelf CPU architectures. They are a vital support to post‐mortem performance profiling and are exploited by standard tools such as Linux or Intel V‐Tune. Nevertheless, an increasing number of application domains (e.g., simulation, task‐based high‐performance computing, or cybersecurity) are exploiting them to perform different activities, such as self‐tuning, autonomic optimization, and/or system inspection. This repurposing of HPCs can be difficult, for example, because of the overhead for extracting relevant information. This overhead might render any online or self‐tuning activity ineffective. This article discusses various practical strategies to exploit HPCs beyond post‐mortem profiling, suitable for different application contexts. The presented strategies are accompanied by a general primer on HPCs usage on Linux. We also provide reference x86 (both Intel and AMD) implementations targeting the Linux kernel, upon which we present an experimental assessment of the viability of our proposals. Stefano Carnà, Romolo Marotta, Alessandro Pellegrini 0001, Francesco Quaglia |
Softw. Pract. Exp. | 2 |
| 2022 | Spatial/Temporal Locality-based Load-sharing in Speculative Discrete Event Simulation on Multi-core MachinesabstractThe recent literature has reshuffled the architectural organization of speculative parallel discrete event simulation systems for shared-memory multi-core machines. A core aspect has been the full sharing of the workload at the level of individual simulation events, which enables keeping the rollback incidence minimal. However, making each worker thread continuously switch its execution between events destined to different simulation objects does not favor locality. In this article, we propose a workload-sharing algorithm where the worker threads can have short-term binding with specific simulation objects to favor spatial locality and caching effectiveness. Also, new bindings—carried out when a thread decides to switch its execution to other simulation objects—are based on the timeline according to which the object states have passed through the caching hierarchy. At the same time, our solution still enables the worker threads to focus their activities on the events to be processed whose timestamps are closer to the simulation commit horizon—hence we exploit temporal locality along virtual time and keep the rollback incidence minimal. In our design we exploit lock-free constructs to support scalable thread synchronization while accessing the shared event pool. Furthermore, we exploit a multi-view approach of the event pool content, which additionally favors local accesses to the parts of the event pool that are currently relevant for the thread activity. Our solution has been released as an integration within the USE open source speculative simulation platform available to the community. Furthermore, in this article we report the results of an experimental study that shows the effectiveness of our proposal. Federica Montesano, Romolo Marotta, Francesco Quaglia |
SIGSIM-PADS | 2 |
| 2022 | NBBS: A Non-Blocking Buddy System for Multi-Core MachinesabstractCommon implementations of core memory allocation components handle concurrent allocation/release requests by synchronizing threads via spin-locks. This approach is not prone to scale, a problem that has been addressed in the literature by introducing layered allocation services or replicating the core allocators—the bottom-most ones within the layered architecture. Both these solutions tend to reduce the pressure of actual concurrent accesses to each individual core allocator. In this article, we explore an alternative approach to scalability of memory allocation/release, which can be still combined with those literature proposals. We present a fully non-blocking buddy system, where threads performing concurrent allocations/releases do not undergo any spin-lock based synchronization. Our solution allows threads to proceed in parallel, and commit their allocations/releases unless a conflict is materialized while handling the allocator metadata—memory fragmentation and coalescing are also carried out in a fully non-blocking manner. Conflict detection relies in our solution on atomic Read-Modify-Write (RMW) machine instructions, guaranteed to execute atomically by the processor firmware. We also provide a proof of the correctness of our non-blocking buddy system and show the results of an experimental study that outlines the effectiveness of our solution. Romolo Marotta, Mauro Ianni, Alessandro Pellegrini 0001, Francesco Quaglia |
IEEE Trans. Computers | 1 |
| 2021 | PECS'21: The First Workshop on Performance and Energy-efficiency of Concurrent SystemsabstractConcurrent systems, based on (distributed) multi/many-core processing units, are the nowadays reference computing architecture. The (continuously-growing) level of hardware parallelism they offer has led these platforms to play a central role at any scale, ranging from data centers, to personal (mobile) devices. Optimizing performance and/or ensuring energy efficiency when running complex software stacks on top of these systems is extremely challenging due to several aspects, like data dependencies or resource sharing (and interference) among application threads, as well as VMs. Furthermore, hardware accelerators like GPGPUs or FPGAs introduce a level of heterogeneity that can potentially offer further opportunities for combined gain in performance and energy efficiency, if correctly exploited. Romolo Marotta, Francesco Quaglia |
ICPE | 1 |
| 2020 | NUMA-Aware Non-Blocking Calendar QueueabstractModern computing platforms are based on multi-processor/multi-core technology. This allows running applications with a high degree of hardware parallelism. However, medium-to-high end machines pose a problem related to the asymmetric delays threads experience when accessing shared data. Specifically, Non-Uniform-Memory-Access (NUMA) is the dominating technology-thanks to its capability for scaled-up memory bandwidth-which however imposes asymmetric distances between CPU-cores and memory banks, making an access by a thread to data placed on a far NUMA node severely impacting performance. In this article, we tackle this problem in the context of shared event-pool management, a relevant aspect in many fields, like parallel discrete event simulation. Specifically, we present a NUMA-aware calendar queue, which also has the advantage of making concurrent threads coordinate via a non-blocking scalable approach. Our proposal is based on work deferring combined with dynamic re-binding of the calendar queue operations (insertions/extractions) to the best suited among the concurrent threads hosted by the underlying computing platform. This changes the locality of the operations by threads in a way positively reflected onto NUMA tasks at the hardware level. We report the results of an experimental study, demonstrating the capability of our solution to achieve the order of 15% better performance compared to state-of-the-art solutions already suited for multicore environments. Maryan Rab, Romolo Marotta, Mauro Ianni, Alessandro Pellegrini 0001, Francesco Quaglia |
DS-RT | 2 |
| 2020 | Autonomic Power Management in Speculative Simulation Runtime EnvironmentsabstractWhile transitioning to exascale systems, it has become clear that power management plays a fundamental role to support a viable utilization of the underlying hardware, also performance-wise. To meet power restrictions imposed by future exascale supercomputers, runtime environments will be required to enforce self-tuning schemes to run dynamic workloads under an imposed power cap. Literature results show that, for a wide class of multi-threaded applications, tuning both the degree of parallelism and frequency/voltage of cores allows a more effective use of the budget, compared to techniques that use only one of these mechanisms in isolation. In this paper, we explore the issues associated with applying these techniques on speculative Time-Warp based simulation runtime environments. We discuss how the differences in two antithetical Time Warp-based simulation environments impact the obtained results. Our assessment confirms that the performance gains achieved through a proper allocation of the power budget can be significant. We also identify the research challenges that would make these form of self-tuning more broadly applicable. Stefano Conoci, Mauro Ianni, Romolo Marotta, Alessandro Pellegrini 0001 |
SIGSIM-PADS | 3 |
| 2020 | Exploiting Inter-Processor-Interrupts for Virtual-Time Coordination in Speculative Parallel Discrete Event SimulationabstractReducing the waste of resource usage (e.g., CPU-cycles) when a causality error occurs in speculative parallel discrete event simulation (PDES) is still a core objective. In this article, we target this objective in the context of speculative PDES run on top of shared-memory machines. We propose an Operating System approach that is based on the exploitation of the Inter-Processor-Interrupt (IPI) facility offered by off-the-shelf hardware chipsets, which enables cross-CPU-core control of the execution flow of threads. As soon as a thread T produces a new event placed in the past virtual time of a simulation object currently run by another thread T', our IPI-based support allows T to change the execution flow of T'---with very minimal delay---so to enable the early squash of the currently processed (and no longer consistent) event. Our solution is fully transparent to the application level code, and is coupled with a lightweight heuristic-based mechanism that determines the actual goodness of killing thread T' via the IPI (rather than skipping the IPI send) depending on the expected residual execution time of the incorrect event being processed. We integrated our proposal within the speculative open-source USE (Ultimate Share Everything) PDES package, and we report experimental results obtained by running various PDES models on top of two shared-memory hardware architectures equipped with 32 and 24 (48 Hyper-threads) CPU-cores, which demonstrate the effectiveness of our proposal. Emiliano Silvestri, Cristian Milia, Romolo Marotta, Alessandro Pellegrini 0001, Francesco Quaglia |
SIGSIM-PADS | 3 |
| 2020 | Mutable locks: Combining the best of spin and sleep locksabstractSummary In this article, we present mutable locks, a synchronization construct with the same semantic of traditional locks (such as spin locks or sleep locks), but with a self‐tuned optimized trade‐off between responsiveness and CPU‐time usage during threads' wait phases. Mutable locks tackle the need for efficient synchronization supports in the era of multicore machines, where the run‐time performance should be optimized while reducing resource usage. This goal should be achieved with no intervention by the programmers. Our proposal is intended for exploitation in generic concurrent applications, where scarce or no knowledge is available about the underlying software/hardware stack and the workload. This is an adverse scenario for static choices between spinning and sleeping, which is tackled by our mutable locks thanks to their hybrid waiting phase and self‐tuning capabilities. Romolo Marotta, Davide Tiriticco, Pierangelo di Sanzo, Alessandro Pellegrini 0001, Bruno Ciciani, Francesco Quaglia |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | NBBS: A Non-Blocking Buddy System for Multi-core MachinesabstractCommon implementations of core memory allocation components, like the Linux buddy system, handle concurrent allocation/release requests by synchronizing threads via spin-locks. This approach is not prone to scale, a problem that has been addressed in the literature by introducing layered allocation services or replicating the core allocators-the bottom most ones within the layered architecture. Both these solutions tend to reduce the pressure of actual concurrent accesses to each individual core allocator. In this article we explore an alternative approach to scalability of memory allocation/release, which can be still combined with those literature proposals. We present a fully non-blocking buddy-system, where threads performing concurrent allocations/releases do not undergo any spin-lock based synchronization. Our solution allows threads to proceed in parallel, and commit their allocations/releases unless a conflict is materialized while handling the allocator metadata. Conflict detection relies on atomic Read-Modify-Write (RMW) machine instructions. Beyond improving scalability and performance, our solution can also avoid wasting clock cycles for spin-lock operations by threads that could in principle carry out their memory allocations/releases in full concurrency. Romolo Marotta, Mauro Ianni, Andrea Scarselli, Alessandro Pellegrini 0001, Francesco Quaglia |
CCGRID | 1 |
| 2018 | A Non-blocking Buddy System for Scalable Memory Allocation on Multi-core MachinesabstractCommon implementations of core memory allocation components handle concurrent allocation/release requests by synchronizing threads via spin-locks. This approach is not prone to scale with large thread counts, a problem that has been addressed in the literature by introducing layered allocation services or replicating the core allocators-the bottom most ones within the layered architecture. Both these solutions tend to reduce the pressure of actual concurrent accesses to each individual core allocator. In this article we explore an alternative approach to scalability of memory allocation/release, which can be still combined with those literature proposals. We present a fully non-blocking buddy-system, that allows threads to proceed in parallel, and commit their allocations/releases unless a conflict is materialized while handling its metadata. Beyond improving scalability and performance it is resilient to performance degradation in face of concurrent accesses independently of the current level of fragmentation of the handled memory blocks. Romolo Marotta, Mauro Ianni, Andrea Scarselli, Alessandro Pellegrini 0001, Francesco Quaglia |
CLUSTER | 1 |
| 2018 | The Ultimate Share-Everything PDES SystemabstractThe share-everything PDES (Parallel Discrete Event Simulation) paradigm is based on fully sharing the possibility to process any individual event across concurrent threads, rather than binding Logical Processes (LPs) and their events to threads. It allows concentrating, at any time, the computing power---the CPU-cores on board of a shared-memory machine---towards the unprocessed events that stand closest to the current commit horizon of the simulation run. This fruitfully biases the delivery of the computing power towards the hot portion of the model execution trajectory. In this article we present an innovative share-everything PDES system that provides (1) fully non-blocking coordination of the threads when accessing shared data structures and (2) fully speculative processing capabilities---Time Warp style processing---of the events. As we show via an experimental study, our proposal can cope with hard workloads where both classical Time Warp systems---based on LPs to threads binding---and previous share-everything proposals---not able to exploit fully speculative processing of the events---tend to fail in delivering adequate performance. Mauro Ianni, Romolo Marotta, Davide Cingolani, Alessandro Pellegrini 0001, Francesco Quaglia |
SIGSIM-PADS | 2 |
| 2017 | A non-blocking global virtual time algorithm with logarithmic number of memory operationsabstractThe increasing diffusion of shared-memory multi-core machines has given rise to a change in the design of Parallel Discrete Event Simulation (PDES) platforms. In particular, the possibility to share large amounts of memory by many worker threads has lead to a boost in the adoption of non-blocking coordination algorithms, which have been proven to offer higher scalability when compared to their blocking counterparts based on critical sections. In this article we present an innovative non-blocking algorithm for computing Global Virtual Time (GVT) - namely, the current commit horizon-in multi-thread PDES engines to be run on top of multi-core machines. Beyond being non-blocking, our proposal has the advantage of providing a logarithmic (rather than linear) number of per-thread memory operations - read/write operations of values involved in the reduction for computing the GVT value-vs the amount of threads participating in the GVT computation. This allows for keeping low the actual CPU time that is required for determining the new GVT value. We compare our algorithm with a literature solution, still based on the non-blocking approach, but entailing a linear number of memory operations, quantifying the advantages from our proposal especially for very large numbers of threads participating in the GVT computation. Mauro Ianni, Romolo Marotta, Alessandro Pellegrini 0001, Francesco Quaglia |
DS-RT | 2 |
| 2017 | Towards a fully non-blocking share-everything PDES platformabstractShared-memory multi-core platforms are changing the nature of Parallel Discrete Event Simulation (PDES) because of the possibility to fully share the workload of events to be processed across threads. In this context, one rising PDES paradigm - referred to as share-everything PDES - is no longer based on the concept of (temporary) biding of simulation objects to worker threads. Rather, each worker threads can - at any time - pick from a fully shared event pool an event to process which can be destined to whatever simulation object. While attention has been posed on the design of concurrent shared pools, allowing non-blocking parallel operations, the scenario where two (or more) threads pick events destined to the same simulation object still lacks adequate synchronization support. In fact, these events are currently sequentialized and processed in a critical section touching the simulation object state, thus leading threads to mutually block each other. In this article we present the design of a share-everything speculative PDES engine that prevents mutual thread blocks because of the access to a same object state. In our design, the non-blocking property is seen as a vertical attribute of the engine (not only of the event pool). This vertical view demands for innovative event-dispatching schemes and, at the same time, innovative interactions with (and management of) the fully-shared event pool, which are features that we embed in our innovative design. Mauro Ianni, Romolo Marotta, Alessandro Pellegrini 0001, Francesco Quaglia |
DS-RT | 2 |
| 2017 | A Conflict-Resilient Lock-Free Calendar Queue for Scalable Share-Everything PDES PlatformsabstractEmerging share-everything Parallel Discrete Event Simulation (PDES) platforms rely on worker threads fully sharing the workload of events to be processed. These platforms require efficient event pool data structures enabling high concurrency of extraction/insertion operations. Non-blocking event pool algorithms are raising as promising solutions for this problem. However, the classical non-blocking paradigm leads concurrent conflicting operations, acting on a same portion of the event pool data structure, to abort and then retry. In this article we present a conflict-resilient non-blocking calendar queue that enables conflicting dequeue operations, concurrently attempting to extract the minimum element, to survive, thus improving the level of scalability of accesses to the hot portion of the data structure---namely the bucket to which the current locality of the events to be processed is bound. We have integrated our solution within an open source share-everything PDES platform and report the results of an experimental analysis of the proposed concurrent data structure compared to some literature solutions. Romolo Marotta, Mauro Ianni, Alessandro Pellegrini 0001, Francesco Quaglia |
SIGSIM-PADS | 1 |
| 2016 | A Lock-Free O(1) Event Pool and Its Application to Share-Everything PDES PlatformsabstractThe large diffusion of highly-parallel shared-memory multi-core machines has led Parallel Discrete Event Simulation (PDES) platforms to a shift towards a share-everything model. This model is based on loose coupling between simulation objects and threads, lasting (as an extreme) no more than the lifetime of individual events. Concurrent threads can therefore CPU-dispatch events destined to any object at any point in time, thus fully sharing the workload of events to be processed on a fine grain basis. This demands for efficient mechanisms to share the overall pool of pending events by enabling parallelism in insertion and extraction operations. In this article we present a lock-free event pool which also provides amortized O(1) time complexity for both insertions and extractions. It can sustain highly concurrent accesses, while not leading to noticeable performance degradation when scaling up the thread count. Experimental results demonstrate that our solution stands as a core facility capable of further raising up the pragmatical impact of such an emerging share-everything PDES paradigm. Romolo Marotta, Mauro Ianni, Alessandro Pellegrini 0001, Francesco Quaglia |
DS-RT | 1 |
| 2014 | Estimating the Empirical Cost Function of Routines with Dynamic Workloads
Emilio Coppa, Camil Demetrescu, Irene Finocchi, Romolo Marotta |
CGO | 4 |