EDBT 2026 Demo / reviewers in the wild / expert
Paolo Romano 0002
dblp:r/PaoloRomano0
· DBLP profile ↗
101ranked-venue papers
17as first author
12since 2021 · last 2026
0000-0001-7026-7446ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 66 · 9 first-author · 9 since 2021Software engineering, systems software and programming languages · 16 · 1 first-author · 1 since 2021Security and privacy · 10 · 2 first-authorArtificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FUR: Fast and Unlimited Reads on Persistent Memory TransactionsabstractDespite the recent improvements in supporting Persistent Hardware Transactions (PHTs) on emerging persistent memories (PM), they have largely overlooked the poor performance of Read-Only (RO) transactions, which suffer from two crucial bottlenecks: i) the considerable post-commit delays required to ensure consistency with concurrent update transactions; and ii) the well-known tight read capacity limits of the commercially available HTM implementations. João Barreto 0001, Daniel Castro 0004, Paolo Romano 0002, Alexandro Baldassin |
EuroSys | 3 |
| 2026 | Accelerating Transactional Execution via Processing-In-Memory
André Lopes, Daniel Castro 0004, Paolo Romano 0002 |
EuroSys | 3 |
| 2026 | SpecPool: Storage-Layer Speculation for Parallel Smart Contract Execution
Francisco Rola, Miguel Matos, Michal Nazarewicz, Paolo Romano 0002 |
ICDCS | 4 |
| 2025 | Uncertainty Estimation by Human Perception versus Neural Models
Paolo Romano 0002, David Garlan |
PRICAI (4) | 2 |
| 2024 | PIM-STM: Software Transactional Memory for Processing-In-Memory SystemsabstractProcessing-In-Memory (PIM) is a novel approach that augments existing DRAM memory chips with lightweight logic. By allowing to offload computations to the PIM system, this architecture allows for circumventing the data-bottleneck problem that affects many modern workloads. André Lopes, Daniel Castro 0004, Paolo Romano 0002 |
ASPLOS (2) | 3 |
| 2024 | Error-Driven Uncertainty Aware TrainingabstractNeural networks are often overconfident about their predictions, which undermines their reliability and trustworthiness. In this work, we present a novel technique, named Error-Driven Uncertainty Aware Training (EUAT), which aims to enhance the ability of neural classifiers to estimate their uncertainty correctly, namely to be highly uncertain when they output inaccurate predictions and low uncertain when their output is accurate. The EUAT approach operates during the model’s training phase by selectively employing two loss functions depending on whether the training examples are correctly or incorrectly predicted by the model. This allows for pursuing the twofold goal of i) minimizing model uncertainty for correctly predicted inputs and ii) maximizing uncertainty for mispredicted inputs, while preserving the model’s misprediction rate. We evaluate EUAT using diverse neural models and datasets in the image recognition domains considering both non-adversarial and adversarial settings. The results show that EUAT outperforms existing approaches for uncertainty estimation (including other uncertainty-aware training techniques, calibration, ensembles, and DEUP) by providing uncertainty estimates that not only have higher quality when evaluated via statistical metrics (e.g., correlation with residuals) but also when employed to build binary classifiers that decide whether the model’s output can be trusted or not and under distributional data shifts. Paolo Romano 0002, David Garlan |
ECAI | 2 |
| 2024 | Self-adapting Machine Learning-based Systems via a Probabilistic Model Checking FrameworkabstractThis article focuses on the problem of optimizing the system utility of Machine Learning (ML)-based systems in the presence of ML mispredictions. This is achieved via the use of self-adaptive systems and through the execution of adaptation tactics, such as model retraining , which operate at the level of individual ML components. To address this problem, we propose a probabilistic modeling framework that reasons about the cost/benefit tradeoffs associated with adapting ML components. The key idea of the proposed approach is to decouple the problems of estimating (1) the expected performance improvement after adaptation and (2) the impact of ML adaptation on overall system utility. We apply the proposed framework to engineer a self-adaptive ML-based fraud detection system, which we evaluate using a publicly available, real fraud detection dataset. We initially consider a scenario in which information on the model’s quality is immediately available. Next, we relax this assumption by integrating (and extending) state-of-the-art techniques for estimating the model’s quality in the proposed framework. We show that by predicting the system utility stemming from retraining an ML component, the probabilistic model checker can generate adaptation strategies that are significantly closer to the optimal, as compared against baselines such as periodic or reactive retraining. Maria Casimiro, Diogo Soares, David Garlan, Luís E. T. Rodrigues, Paolo Romano 0002 |
ACM Trans. Auton. Adapt. Syst. | 5 |
| 2023 | HyperJump: Accelerating HyperBand via Risk ModellingabstractIn the literature on hyper-parameter tuning, a number of recent solutions rely on low-fidelity observations (e.g., training with sub-sampled datasets) to identify promising configurations to be tested via high-fidelity observations (e.g., using the full dataset). Among these, HyperBand is arguably one of the most popular solutions, due to its efficiency and theoretically provable robustness. In this work, we introduce HyperJump, a new approach that builds on HyperBand’s robust search strategy and complements it with novel model-based risk analysis techniques that accelerate the search by skipping the evaluation of low risk configurations, i.e., configurations that are likely to be eventually discarded by HyperBand. We evaluate HyperJump on a suite of hyper-parameter optimization problems and show that it provides over one-order of magnitude speed-ups, both in sequential and parallel deployments, on a variety of deep-learning, kernel-based learning and neural architectural search problems when compared to HyperBand and to several state-of-the-art optimizers. Maria Casimiro, Paolo Romano 0002, David Garlan |
AAAI | 3 |
| 2023 | CSMV: A highly scalable multi-versioned software transactional memory for GPUs
Diogo Nunes, Daniel Castro 0004, Paolo Romano 0002 |
J. Parallel Distributed Comput. | 3 |
| 2022 | CSMV: A Highly Scalable Multi-Versioned Software Transactional Memory for GPUsabstractGPUs have traditionally focused on streaming applications with regular parallelism. Over the last years, though, GPUs have also been successfully used to accelerate irregular applications in a number of application domains by using fine grained synchronization schemes. Unfortunately, fine-grained synchronization strategies are notoriously complex and error-prone. This has motivated the search for alternative paradigms aimed to simplify concurrent programming and, among these, Transactional Memory (TM) is probably one of the most prominent proposals. This paper introduces CSMV (Client Server Multiversioned), a multi-versioned Software TM (STM) for GPUs that adopts an innovative client-server design. By decoupling the execution of transactions from their commit process, CSMV provides two main benefits: (i) it enables the use of fast on chip memory to access the global metadata used to synchronize transaction (ii) it allows for implementing highly efficient collaborative commit procedures, tailored to take full advantage of the architectural characteristics of GPUs. Via an extensive experimental study, we show that CSMV achieves up to 3 orders of magnitude speed-ups with respect to state of the art STMs for GPUs and that it can accelerate by up to 20× irregular applications running on state of the art STMs for CPUs. Diogo Nunes, Daniel Castro 0004, Paolo Romano 0002 |
IPDPS | 3 |
| 2021 | SPHT: Scalable Persistent Hardware Transactions
Daniel Castro 0004, Alexandro Baldassin, João Barreto 0001, Paolo Romano 0002 |
FAST | 4 |
| 2021 | Investigating the semantics of futures in transactional memory systemsabstractThis paper investigates the problem of integrating two powerful abstractions for concurrent programming, namely futures and transactional memory. Our focus is on specifying the semantics of execution of "transactional futures", i.e., futures that execute as atomic transactions and that are spawned/evaluated by other (plain) transactions or transactional futures. We show that, due to the ability of futures to generate parallel computations with complex dependencies, there exist several plausible (i.e., intuitive) alternatives for defining the isolation and atomicity semantics of transactional futures. The alternative semantics we propose explore different trade-offs between ease of use and efficiency. We have implemented the proposed semantics by introducing a graph-based software transactional memory algorithm, which we integrated with a state of the art JAVA-based Software Transactional Memory (STM). We quantify the performance trade-offs associated with the different semantics using an extensive experimental study encompassing a wide range of diverse workloads. Jingna Zeng, Shady Issa, Paolo Romano 0002, Luís E. T. Rodrigues, Seif Haridi |
PPoPP | 3 |
| 2020 | NV-PhTM: An Efficient Phase-Based Transactional System for Non-volatile Memory
Alexandro Baldassin, Rafael Murari, João P. L. de Carvalho, Guido Araujo, Daniel Castro 0004, João Barreto 0001, Paolo Romano 0002 |
Euro-Par | 7 |
| 2020 | Lynceus: Cost-efficient Tuning and Provisioning of Data Analytic JobsabstractModern data analytic and machine learning jobs find in the cloud a natural deployment platform to satisfy their notoriously large resource requirements. Yet, to achieve cost efficiency, it is crucial to identify a deployment configuration that satisfies user-defined QoS constraints (e.g., on execution time), while avoiding unnecessary over-provisioning.This paper introduces Lynceus, a new approach for the optimization of cloud-based data analytic jobs that improves over state-of-the-art approaches by enabling significant cost savings both in terms of the final recommended configuration and of the optimization process used to recommend configurations.Unlike existing solutions, Lynceus optimizes in a joint fashion both the cloud-related (i.e., which and how many machines to provision) and the application-level (e.g. the hyper-parameters of a machine learning algorithm) parameters. This allows for a reduction of the cost of recommended configurations by up to 3.7× at the 90-th percentile with respect to existing approaches, which treat the optimization of cloud-related and application- level parameters as two independent problems.Further, Lynceus reduces the cost of the optimization process (i.e., the cloud cost incurred for testing configurations) by up to 11×. Such an improvement is achieved thanks to two mechanisms: i) a timeout approach which allows to abort the exploration of configurations that are deemed suboptimal, while still extracting useful information to guide future explorations and to improve its predictive model - differently from recent works, which either incur the full cost for testing suboptimal configurations or are unable to extract any knowledge from aborted runs; ii) a long-sighted and budget-aware technique that determines which configurations to test by predicting the long-term impact of each exploration - unlike state-of-the-art approaches for the optimization of cloud jobs, which adopt greedy optimization methods. Maria Casimiro, Diego Didona, Paolo Romano 0002, Luís E. T. Rodrigues, Willy Zwaenepoel, David Garlan |
ICDCS | 3 |
| 2020 | Exploiting Symbolic Execution to Accelerate Deterministic DatabasesabstractDeterministic databases (DDs) are a promising approach for replicating data across different replicas. A fundamental component of DDs is a deterministic concurrency control algorithm that, given a set of transactions in a specific order, guarantees that their execution always results in the same serial order. State-of-the-art approaches either rely on single threaded execution or on the knowledge of read- and write-sets of transactions to achieve this goal. The former yields poor performance in multi-core machines while the latter requires either manual inputs from the user - a time-consuming and error prone task - or a reconnaissance phase that increases both the latency and abort rates of transactions. In this paper, we present Prognosticator, a novel deterministic database system. Rather than relying on manual transaction classification or an expert programmer, Prognosticator employs Symbolic Execution to build fine-grained transaction profiles (at the key-level). These profiles are then used by Prognosticator's novel deterministic concurrency control algorithm to execute transactions with a high degree of parallelism.Our experimental evaluation, based on both TPC-C and RUBiS benchmarks, shows that Prognosticator can achieve up to 5× higher throughput with respect to state-of-the-art solutions. Shady Issa, Miguel Viegas, Pedro Raminhas, Nuno Machado, Miguel Matos, Paolo Romano 0002 |
ICDCS | 6 |
| 2020 | Bandwidth-Aware Page Placement in NUMAabstractPage placement is a critical problem for memory-intensive applications running on a shared-memory multiprocessor with a non-uniform memory access (NUMA) architecture. State-of-the-art page placement mechanisms interleave pages evenly across NUMA nodes. However, this approach fails to maximize memory throughput in modern NUMA systems, characterized by asymmetric bandwidths and latencies, and sensitive to memory contention and interconnect congestion phenomena.We propose BWAP, a novel page placement mechanism based on asymmetric weighted page interleaving. BWAP combines an analytical performance model of the target NUMA system with on-line iterative tuning of page distribution for a given memory-intensive application. Our experimental evaluation with representative memory-intensive workloads shows that BWAP performs up to 66% better than state-of-the-art techniques. These gains are particularly relevant when multiple co-located applications run in disjoint partitions of a large NUMA machine or when applications do not scale up to the total number of cores. David Gureya, João Neto 0001, Reza Karimi, João Barreto 0001, Pramod Bhatotia, Vivien Quéma, Rodrigo Rodrigues 0001, Paolo Romano 0002, Vladimir Vlassov |
IPDPS | 8 |
| 2020 | TrimTuner: Efficient Optimization of Machine Learning Jobs in the Cloud via Sub-SamplingabstractThis work introduces TrimTuner, the first system for optimizing machine learning jobs in the cloud to exploit sub-sampling techniques to reduce the cost of the optimization process, while keeping into account user-specified constraints. TrimTuner jointly optimizes the cloud and application-specific parameters and, unlike state of the art works for cloud optimization, eschews the need to train the model with the full training set every time a new configuration is sampled. Indeed, by leveraging sub-sampling techniques and data-sets that are up to 60 x smaller than the original one, we show that TrimTuner can reduce the cost of the optimization process by up to 50 x. Further, TrimTuner speeds-up the recommendation process by 65 x with respect to state of the art techniques for hyperparameter optimization that use sub-sampling techniques. The reasons for this improvement are twofold: i) a novel domain specific heuristic that reduces the number of configurations for which the acquisition function has to be evaluated; ii) the adoption of an ensemble of decision trees that enables boosting the speed of the recommendation process by one additional order of magnitude. Maria Casimiro, Paolo Romano 0002, David Garlan |
MASCOTS | 3 |
| 2020 | Giving Future(s) to Transactional Memory
Jingna Zeng, Seif Haridi, Shady Issa, Paolo Romano 0002, Luís E. T. Rodrigues |
SPAA | 4 |
| 2020 | Extending hardware transactional memory capacity via rollback-only transactions and suspend/resume
Shady Issa, Pascal Felber, Alexander Matveev, Paolo Romano 0002 |
Distributed Comput. | 4 |
| 2020 | Transparent speculation in geo-replicated transactional data stores
Zhongmiao Li, Paolo Romano 0002, Peter Van Roy |
J. Parallel Distributed Comput. | 2 |
| 2019 | HeTM: Transactional Memory for Heterogeneous SystemsabstractModern heterogeneous computing architectures, which couple multi-core CPUs with discrete many-core GPUs (or other specialized hardware accelerators), enable unprecedented peak performance and energy efficiency levels. However, developing applications that can take full advantage of the potential of heterogeneous systems is a notoriously hard task. This work takes a step towards reducing the complexity of programming heterogeneous systems by introducing the abstraction of Heterogeneous Transactional Memory (HeTM). HeTM provides programmers with the illusion of a single memory region, shared among the CPUs and the (discrete) GPU(s) of a heterogeneous system, with support for atomic transactions. Besides introducing the abstract semantics and programming model of HeTM, we present the design and evaluation of a concrete implementation of the proposed abstraction, referred herein as Speculative HeTM (SHeTM). SHeTM makes use of a novel design that leverages speculative techniques, which aims at hiding the inherently large communication latency between CPUs and discrete GPUs and at minimizing inter-device synchronization overhead. We demonstrate the efficiency of the SHeTM via an extensive quantitative study based both on synthetic benchmarks and on a popular object caching system. Daniel Castro 0004, Paolo Romano 0002, Aleksandar Ilic, Amin M. Khan |
PACT | 2 |
| 2019 | Sparkle: Speculative Deterministic Concurrency Control for Partially Replicated Transactional StoresabstractModern transactional platforms strive to jointly ensure ACID consistency and high scalability. In order to pursue these antagonistic goals, several recent systems have revisited the classical State Machine Replication (SMR) approach in order to support sharding of application state across multiple data partitions and partial replication. By promoting and exploiting locality principles, these systems, which we call Partially Replicated State Machines (PRSMs), can achieve scalability levels unparalleled by classic SMR. Yet, existing PRSM systems suffer from two major limitations: 1) they rely on a single thread to execute or serialize transactions within a partition, thus failing to fully untap the parallelism of multi-core architectures, and/or 2) they rely on the ability to accurately predict the data items to be accessed by transactions, which is non-trivial for complex applications. This paper proposes Sparkle, an innovative deterministic concurrency control that enhances the throughput of state of the art PRSM systems by more than one order of magnitude under high contention, through the joint use of speculative transaction processing and scheduling techniques. On the one hand, speculation allows Sparkle to take full advantage of modern multi-core micro-processors, while avoiding any assumption on the a-priori knowledge of the transactions' access patterns, which increases its generality and widens the scope of its scalability. Transaction scheduling techniques, on the other hand, are aimed to maximize the efficiency of speculative processing. Zhongmiao Li, Paolo Romano 0002, Peter Van Roy |
DSN | 2 |
| 2019 | Stretching the capacity of hardware transactional memory in IBM POWER architecturesabstractThe hardware transactional memory (HTM) implementations in commercially available processors are significantly hindered by their tight capacity constraints. In practice, this renders current HTMs unsuitable to many real-world workloads of in-memory databases. Ricardo Filipe, Shady Issa, Paolo Romano 0002, João Barreto 0001 |
PPoPP | 3 |
| 2019 | Hardware Transactional Memory meets memory persistency
Daniel Castro 0004, Paolo Romano 0002, João Barreto 0001 |
J. Parallel Distributed Comput. | 2 |
| 2018 | Enhancing Efficiency of Hybrid Transactional Memory Via Dynamic Data Partitioning SchemesabstractTransactional Memory (TM) is an emerging paradigm that promises to significantly ease the development of parallel programs. Hybrid TM (HyTM) is probably the most promising implementation of the TM abstraction, which seeks to combine the high efficiency of hardware implementations (HTM) with the robustness and flexibility of software-based ones (STM). Unfortunately, though, existing Hybrid TM systems are known to suffer from high overheads to guarantee correct synchronization between concurrent transactions executing in hardware and software. This article introduces DMP-TM (Dynamic Memory Partitioning-TM), a novel HyTM algorithm that exploits, to the best of our knowledge for the first time in the literature, the idea of leveraging operating system-level memory protection mechanisms to detect conflicts between HTM and STM transactions. This innovative design allows for employing highly scalable STM implementations, while avoiding instrumentation on the HTM path. This allows DMP-TM to achieve up to ~ 20× speedups compared to state of the art Hybrid TM solutions in uncontended workloads. Further, thanks to the use of simple and lightweight self-tuning mechanisms, DMP-TM achieves robust performance even in unfavourable workload that exhibits high contention between the STM and HTM path. Pedro Raminhas, Shady Issa, Paolo Romano 0002 |
CCGrid | 3 |
| 2018 | Transparent speculation in geo-replicated transactional data storesabstractThis work presents Speculative Transaction Replication (STR), a protocol that exploits transparent speculation techniques to enhance performance of geo-distributed, partially replicated transactional data stores. In addition, we define a new consistency model, Speculative Snapshot Isolation (SPSI), that extends the semantics of Snapshot Isolation (SI) to shelter applications from the subtle anomalies that can arise from using speculative transaction processing. SPSI extends SI in an intuitive and rigorous fashion by specifying desirable atomicity and isolation guarantees that must hold when using speculative execution. Zhongmiao Li, Peter Van Roy, Paolo Romano 0002 |
HPDC | 3 |
| 2018 | Hardware Transactional Memory Meets Memory PersistencyabstractPersistent Memory (PM) and Hardware Transactional Memory (HTM) are two recent architectural developments whose joint usage promises to drastically accelerate the performance of concurrent, data-intensive applications. Unfortunately, combining these two mechanisms using existing architectural supports is far from being trivial. This paper presents NV-HTM, a system that allows the execution of transactions over PM using unmodified commodity HTM implementations. NV-HTM relies on a hardware-software co-design technique, which is based on three key ideas: i) relying on software to persist transactional modifications after they have been committed via HTM; ii) postponing the externalization of commit events to applications until it is ensured, via software, that any data version produced and observed by committed transactions is first logged in PM; ii) pruning the commit logs via checkpointing schemes that not only bound the log space and recovery time, but also implement wear levelling techniques to enhance PM's endurance. By means of an extensive experimental evaluation, we show that NV-HTM can achieve up to 10× speed-ups and up to 11.6× reduced flush operations with respect to state of the art solutions, which, unlike NV-HTM, require custom modifications to existing HTM systems. Daniel Castro 0004, Paolo Romano 0002, João Barreto 0001 |
IPDPS | 2 |
| 2018 | Online Tuning of Parallelism Degree in Parallel Nesting Transactional MemoryabstractThis paper addresses the problem of self-tuning the parallelism degree in Transactional Memory (TM) systems that support parallel nesting (PN-TM). This problem has been long investigated for TMs not supporting nesting, but, to the best of our knowledge, has never been studied in the context of PN-TMs. Indeed, the problem complexity is inherently exacerbated in PN-TMs, since these require to identify the optimal parallelism degree not only for top-level transactions but also for nested sub-transactions. The increase of the problem dimensionality raises new challenges (e.g., increase of the search space, and proneness to suffer from local maxima), which are unsatisfactorily addressed by self-tuning solutions conceived for flat nesting TMs. We tackle these challenges by proposing AUTOPN, an on-line self-tuning system that combines model-driven learning techniques with localized search heuristics in order to pursue a twofold goal: i) enhance convergence speed by identifying the most promising region of the search space via model-driven techniques, while ii) increasing robustness against modeling errors, via a final local search phase aimed at refining the model's prediction. We further address the problem of tuning the duration of the monitoring windows used to collect feedback on the system's performance, by introducing novel, domain-specific, mechanisms aimed to strike an optimal trade-off between latency and accuracy of the self-tuning process. We integrated AUTOPN with a state of the art PN-TM (JVSTM) and evaluated it via an extensive experimental study. The results of this study highlight that AUTOPN can achieve gains of up to 45× in terms of increased accuracy and 4× faster convergence speed, when compared with several on-line optimization techniques (gradient descent, simulated annealing and genetic algorithm), some of which were already successfully used in the context of flat nesting TMs. Jingna Zeng, Paolo Romano 0002, João Barreto 0001, Luís E. T. Rodrigues, Seif Haridi |
IPDPS | 2 |
| 2018 | Speculative Read Write LocksabstractHardware Transactional Memory (HTM) has recently entered the realm of mainstream computing thanks to its integration in processors commercialized by major industrial manufacturers. HTM provides highly-efficient, hardware-assisted synchronization mechanisms for concurrent programs. Unfortunately, though, existing HTM implementations also suffer from severe limitations that are inherently related to their best-effort, hardware-based design. Shady Issa, Paolo Romano 0002, Tiago Lopes |
Middleware | 2 |
| 2018 | CoopREP: Cooperative record and replay of concurrency bugsabstractSummary This paper presents CoopREP, a system that provides support for fault replication of concurrent programs based on cooperative recording and partial log combination. CoopREP uses partial logging to reduce the amount of information that a given program instance is required to store to support deterministic replay. This allows reducing substantially the overhead imposed by the instrumentation of the code, but raises the problem of finding a combination of logs capable of replaying the fault. CoopREP tackles this issue by introducing several innovative statistical analysis techniques aimed at guiding the search of the partial logs to be combined and needed for the replay phase. CoopREP has been evaluated using both standard benchmarks for multithreaded applications and real‐world applications. The results highlight that CoopREP can successfully replay concurrency bugs involving tens of thousands of memory accesses, while reducing recording overhead with respect to state‐of‐the‐art noncooperative logging schemes by up to 13× (and by 2.4× on average). Nuno Machado, Paolo Romano 0002, Luís E. T. Rodrigues |
Softw. Test. Verification Reliab. | 2 |
| 2017 | Exploiting speculation in partially replicated transactional data storesabstractOnline services are often deployed over geographically-scattered data centers (geo-replication), which allows services to be highly available and reduces access latency. On the down side, to provide ACID transactions, global certification (i.e., across data centers) is needed to detect conflicts between concurrent transactions executing at different data centers. The global certification phase reduces throughput because transactions need to hold pre-commit locks, and it increases client-perceived latency because global certification lies in the critical path of transaction execution. Zhongmiao Li, Peter Van Roy, Paolo Romano 0002 |
SoCC | 3 |
| 2017 | An Analytical Model of Hardware Transactional MemoryabstractThis paper investigates the problem of deriving a white box performance model of Hardware Transactional Memory (HTM) systems. The proposed model targets TSX, a popular implementation of HTM integrated in Intel processors starting with the Haswell family in 2013. An inherent difficulty with building white-box models of commercially available HTM systems is that their internals are either vaguely documented or undisclosed by their manufacturers. We tackle this challenge by designing a set of experiments that allow us to shed lights on the internal mechanisms used in TSX to manage conflicts among transactions and to track their readsets and writesets. We exploit the information inferred from this experimental study to build an analytical model of TSX focused on capturing the impact on performance of two key mechanisms: the concurrency control scheme and the management of transactional meta-data in the processor's caches. We validate the proposed model by means of an extensive experimental study encompassing a broad range of workloads executed on a real system. Daniel Castro 0004, Paolo Romano 0002, Diego Didona, Willy Zwaenepoel |
MASCOTS | 2 |
| 2017 | Enhancing throughput of partially replicated state machines via multi-partition operation schedulingabstractState-machine replication (SMR) is a fundamental technique to implement fault-tolerant services. Recently, various works have aimed at enhancing the scalability of SMR by exploiting partial replication techniques. By sharding the state machine across disjoint partitions, and replicating each partition over independent groups of processes, a Partially Replicated State Machine (PRSM) can process operations that involve a single partition by only requiring synchronization among the replicas of that partition - achieving higher scalability than SMR. Unfortunately, though, existing PRSM rely on inefficient mechanisms to coordinate the execution of multi-partition operations, which either impose global coordination across all nodes in the system or require inter-partition synchronization on the critical path of execution of operations. As such, performance and scalability of existing PRSM systems is severely hindered in the presence of even a small fraction of multi-partition operations. This paper tackles this issue by presenting Genepi, a PRSM protocol that introduces a novel, highly efficient mechanism for regulating the execution of multi-partition operations. We show via an experimental evaluation based on both synthetic benchmarks and TPC-C that Genepi can achieve up to 5.5× of throughput gain over existing PRSM systems, with only negligible latency overhead at low load. Zhongmiao Li, Peter Van Roy, Paolo Romano 0002 |
NCA | 3 |
| 2017 | Extending Hardware Transactional Memory Capacity via Rollback-Only Transactions and Suspend/ResumeabstractTransactional memory (TM) aims at simplifying concurrent programming via the familiar abstraction of atomic transactions. Recently, Intel and IBM have integrated hardware based TM (HTM) implementations in commodity processors, paving the way for the mainstream adoption of the TM paradigm. Yet, existing HTM implementations suffer from a crucial limitation, which hampers the adoption of HTM as a general technique for regulating concurrent access to shared memory: the inability to execute transactions whose working sets exceed the capacity of CPU caches. In this paper we propose P8TM, a novel approach that mitigates this limitation on IBM's POWER8 architecture by leveraging a key combination of techniques: uninstrumented read-only transactions, Rollback Only Transaction-based update transactions, HTM-friendly (software-based) read-set tracking, and self-tuning. P8TM can dynamically switch between different execution modes to best adapt to the nature of the transactions and the experienced abort patterns. In-depth evaluation with several benchmarks indicates that P8TM can achieve striking performance gains in workloads that stress the capacity limitations of HTM, while achieving performance on par with HTM even in unfavourable workloads. Shady Issa, Pascal Felber, Alexander Matveev, Paolo Romano 0002 |
DISC | 4 |
| 2017 | Seer: Probabilistic Scheduling for Hardware Transactional MemoryabstractThe ubiquity of multicore processors has led programmers to write parallel and concurrent applications to take advantage of the underlying hardware and speed up their executions. In this context, Transactional Memory (TM) has emerged as a simple and effective synchronization paradigm, via the familiar abstraction of atomic transactions. After many years of intense research, major processor manufacturers (including Intel) have recently released mainstream processors with hardware support for TM (HTM). In this work, we study a relevant issue with great impact on the performance of HTM. Due to the optimistic and inherently limited nature of HTM, transactions may have to be aborted and restarted numerous times, without any progress guarantee. As a result, it is up to the software library that regulates the HTM usage to ensure progress and optimize performance. Transaction scheduling is probably one of the most well-studied and effective techniques to achieve these goals. However, these recent mainstream HTMs have some technical limitations that prevent the adoption of known scheduling techniques: unlike software implementations of TM used in the past, existing HTMs provide limited or no information on which memory regions or contending transactions caused the abort. To address this crucial issue for HTMs, we propose S eer , a software scheduler that addresses precisely this restriction of HTM by leveraging on an online probabilistic inference technique that identifies the most likely conflict relations and establishes a dynamic locking scheme to serialize transactions in a fine-grained manner. The key idea of our solution is to constrain the portions of parallelism that are affecting negatively the whole system. As a result, this not only prevents performance reduction but also in fact unveils further scalability and performance for HTM. Via an extensive evaluation study, we show that S eer improves the performance of the Intel’s HTM by up to 3.6×, and by 65% on average across all concurrency degrees and benchmarks on a large processor with 28 cores. Nuno Diegues, Paolo Romano 0002, Stoyan Garbatov |
ACM Trans. Comput. Syst. | 2 |
| 2016 | ProteusTM: Abstraction Meets Performance in Transactional MemoryabstractThe Transactional Memory (TM) paradigm promises to greatly simplify the development of concurrent applications. This led, over the years, to the creation of a plethora of TM implementations delivering wide ranges of performance across workloads. Yet, no universal implementation fits each and every workload. In fact, the best TM in a given workload can reveal to be disastrous for another one. This forces developers to face the complex task of tuning TM implementations, which significantly hampers their wide adoption. In this paper, we address the challenge of automatically identifying the best TM implementation for a given workload. Our proposed system, ProteusTM, hides behind the TM interface a large library of implementations. Underneath, it leverages a novel multi-dimensional online optimization scheme, combining two popular learning techniques: Collaborative Filtering and Bayesian Optimization. Diego Didona, Nuno Diegues, Anne-Marie Kermarrec, Rachid Guerraoui, Ricardo Neves, Paolo Romano 0002 |
ASPLOS | 6 |
| 2016 | Hardware read-write lock elisionabstractHardware Lock Elision (HLE) represents a promising technique to enhance parallelism of concurrent applications relying on conventional, lock-based synchronization. The idea at the basis of current HLE approaches is to wrap critical sections into hardware transactions: this allows critical sections to be executed in parallel using a speculative approach, while leveraging on conflict detection capabilities provided by hardware transactions to ensure equivalent semantics to pessimistic lock-based synchronization. Pascal Felber, Shady Issa, Alexander Matveev, Paolo Romano 0002 |
EuroSys | 4 |
| 2016 | The Future(s) of Transactional MemoryabstractThis work investigates how to combine two powerful abstractions to manage concurrent programming: Transactional Memory (TM) and futures. The former hides from programmers the complexity of synchronizing concurrent access to shared data, via the familiar abstraction of atomic transactions. The latter serves to schedule and synchronize the parallel execution of computations whose results are not immediately required. While TM and futures are two widely investigated topics, the problem of how to exploit these two abstractions in synergy is still largely unexplored in the literature. This paper fills this gap by introducing Java Transactional Futures (JTF), a Java-based TM implementation that allows programmers to use futures to coordinate the execution of parallel tasks, while leveraging transactions to synchronize accesses to shared data. JTF provides a simple and intuitive semantic regarding the admissible serialization orders of the futures spawned by transactions, by ensuring that the results produced by a future are always consistent with those that one would obtain by executing the future sequentially. Our experimental results show that the use of futures in a TM allows not only to unlock parallelism within transactions, but also to reduce the cost of conflicts among top-level transactions in high contention workloads. Jingna Zeng, João Barreto 0001, Seif Haridi, Luís E. T. Rodrigues, Paolo Romano 0002 |
ICPP | 5 |
| 2016 | STI-BT: A Scalable Transactional IndexabstractDistributed Key-Value (DKV) stores have been intensively used to manage online transaction processing on large data-sets. DKV stores provide simplistic primitives to access data based on the primary key of the stored objects. To help programmers to efficiently retrieve data, some DKV stores provide distributed indexes. Besides that, and also to simplify programming such applications, several proposals have provided strong consistency abstractions via distributed transactions. In this paper we present STI-BT, a highly scalable, transactional index for Distributed Key-Value stores. STI-BT is organized as a distributed B$^+$Tree and adopts an innovative design that allows to achieve high efficiency in large-scale, elastic DKV stores. As such, it provides both the desirable properties identified above, and does so in a far more efficient and scalable way than the few existing state of the art proposals that also enable programmers to have strongly consistent distributed transactional indexes. We have implemented STI-BT on top of an open-source DKV store and deployed it on a public cloud infrastructure. Our extensive study demonstrates scalability in a cluster of$100$machines, and speed ups with respect to state of the art up to$5.4\times$. Nuno Diegues, Paolo Romano 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | GMU: Genuine Multiversion Update-Serializable Partial Data ReplicationabstractIn this article we introduce GMU, a genuine partial replication protocol for transactional systems, which exploits an innovative, highly scalable, distributed multiversioning scheme. Unlike existing multiversion-based solutions, GMU does not rely on any global logical clock, which may represent a contention point and a major impairment to system scalability. Also, GMU never aborts read-only transactions and spares them from undergoing distributed validation schemes. This makes GMU particularly efficient in presence of read-intensive workloads, as typical of a wide range of real-world applications. GMU guarantees the Extended Update Serializability (EUS) isolation level. This consistency criterion is particularly attractive as it is sufficiently strong to ensure correctness even for very demanding applications (such as TPC-C), but is also weak enough to allow efficient and scalable implementations, such as GMU. Further, unlike several relaxed consistency models proposed in literature, EUS shows simple and intuitive semantics, thus being an attractive consistency model for ordinary programmers. We integrated GMU in a popular open source in-memory transactional data grid, namely Infinispan. On the basis of a wide experimental study performed on heterogeneous platforms and using industry standard benchmarks (namely TPC-C and YCSB), we show that GMU achieves almost linear scalability and that it introduces reduced overhead, with respect to solutions ensuring non-serializable semantics, in a wide range of workloads. Sebastiano Peluso, Pedro Ruivo 0002, Paolo Romano 0002, Francesco Quaglia, Luís E. T. Rodrigues |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | Using Analytical Models to Bootstrap Machine Learning Performance PredictorsabstractPerformance modeling is a crucial technique to enable the vision of elastic computing in cloud environments. Conventional approaches to performance modeling rely on two antithetic methodologies: white box modeling, which exploits knowledge on system's internals and capture its dynamics using analytical approaches, and black box techniques, which infer relations among the input and output variables of a system based on the evidences gathered during an initial training phase. In this paper we investigate a technique, which we name Bootstrapping, which aims at reconciling these two methodologies and at compensating the cons of the one with the pros of the other. We analyze the design space of this gray box modeling technique, and identify a number of algorithmic and parametric trade-offs which we evaluate via two realistic case studies, a Key-Value Store and a Total Order Broadcast service. Diego Didona, Paolo Romano 0002 |
ICPADS | 2 |
| 2015 | Green-CM: Energy Efficient Contention Management for Transactional MemoryabstractTransactional memory (TM) is emerging as an attractive synchronization mechanism for concurrent computing. In this work we aim at filling a relevant gap in the TM literature, by investigating the issue of energy efficiency for one crucial building block of TM systems: contention management. Green-CM, the solution proposed in this paper, is the first contention management scheme explicitly designed to jointly optimize both performance and energy consumption. To this end Green-TM combines three key mechanisms: i) it leverages on a novel asymmetric design, which combines different back-off policies in order to take advantage of dynamic frequency and voltage scaling, ii) it introduces an energy efficient design of the back-off mechanism, which combines spin-based and sleep-based implementations, iii) it makes extensive use of self-tuning mechanisms to pursue optimal efficiency across highly heterogeneous workloads. We evaluate Green-CM from both the energy and performance perspectives, and show that it can achieve enhanced efficiency by up to 2.35 times with respect to state of the art contention managers, with an average gain of more than 60% when using 64 threads. Shady Issa, Paolo Romano 0002, Mats Brorsson |
ICPP | 2 |
| 2015 | Enhancing privacy protection in fault replication systemsabstractError reporting systems are valuable mechanisms for enhancing software reliability. Unfortunately, though, conventional error reporting systems are prone to leaking sensitive user information, raising strong privacy concerns. In this work we introduce RE SPA (Recursive Shortest Path-based Anonymizer), a system for generating failure-reproducing, yet anonymized, error reports. RE SPA relies on symbolic execution, executed at client side, in order to identify alternative failure-inducing paths in the program's execution graph, and derive the logical conditions, called path conditions, that define the set of user inputs reproducing these executions. Anonymized failure-inducing inputs are then synthesized using any (random) solution satisfying the path conditions. The search for alternative failure-inducing executions is based on an innovative algorithm that exploits three key ideas: i) ReSPA relies on binary search to determine, in an efficient way, which portions of the original execution should be preserved in the alternative one; ii) in order to identify alternative execution paths with low information leakage, ReSPA explores the execution graph by leveraging on the Djikstra's shortest path algorithm with information leakage as the distance metric; iii) ReSPA ensures provable non-reversibility of the alternative inputs it produces via a recursive technique that anonymizes the alternative inputs found after running the algorithm. We show via an evaluation based on six large, widely used applications and real bugs that ReSPA reduces information leak-age up to 99.76%, and on average by 93.92%. This corresponds to an average increase in privacy by 40% with respect to state-of-the-art systems, with gains that extend up to almost 20x. João Garcia 0001, Paolo Romano 0002 |
ISSRE | 3 |
| 2015 | Q-OPT: Self-tuning Quorum System for Strongly Consistent Software Defined StorageabstractThis paper presents Q-OPT, a system for automatically tuning the configuration of quorum systems in strongly consistent Software Defined Storage (SDS) systems. Q-OPT is able to assign different quorum systems to different items and can be used in a large variety of settings, including systems supporting multiple tenants with different profiles, single tenant systems running applications with different requirements, or systems running a single application that exhibits non-uniform access patterns to data. Q-OPT supports automatic and dynamic reconfiguration, using a combination of complementary techniques, including top-k analysis to prioritise quorum adaptation, machine learning to determine the best quorum configuration, and a non-blocking quorum reconfiguration protocol that preserves consistency during reconfiguration. Q-OPT has been implemented as an extension to one of the most popular open-source SDS, namely Openstack's Swift. Maria Couceiro, Gayana Chandrasekara, Manuel Bravo, Matti A. Hiltunen, Paolo Romano 0002, Luís E. T. Rodrigues |
Middleware | 5 |
| 2015 | Disjoint-Access Parallelism: Impossibility, Possibility, and Cost of Transactional Memory ImplementationsabstractDisjoint-Access Parallelism (DAP) is considered one of the most desirable properties to maximize the scalability of Transactional Memory (TM). This paper investigates the possibility and inherent cost of implementing a DAP TM that ensures two properties that are regarded as important to maximize efficiency in read-dominated workloads, namely having invisible and wait-free read-only transactions. We first prove that relaxing Real-Time Order (RTO) is necessary to implement such a TM. This result motivates us to introduce Witnessable Real-Time Order (WRTO), a weaker variant of RTO that demands enforcing RTO only between directly conflicting transactions. Then we show that adopting WRTO makes it possible to design a strictly DAP TM with invisible and wait-free read-only transactions, while preserving strong progressiveness for write transactions and an isolation level known in literature as Extended Update Serializability. Finally, we shed light on the inherent inefficiency of DAP TM implementations that have invisible and wait-free read-only transactions, by establishing lower bounds on the time and space complexity of such TMs. Sebastiano Peluso, Roberto Palmieri, Paolo Romano 0002, Binoy Ravindran, Francesco Quaglia |
PODC | 3 |
| 2015 | Seer: Probabilistic Scheduling for Hardware Transactional MemoryabstractScheduling concurrent transactions to minimize contention is a well known technique in the Transactional Memory (TM) literature, which was largely investigated in the context of software TMs. However, the recent advent of Hardware Transactional Memory (HTM), and its inherently restricted nature, pose new technical challenges that prevent the adoption of existing schedulers: unlike software implementations of TM, existing HTMs provide no information on which data item or contending transaction caused abort. We propose Seer, a scheduler that addresses precisely this restriction of HTM by leveraging on an on-line probabilistic inference technique that identifies the most likely conflict relations, and establishes a dynamic locking scheme to serialize transactions in a fine-grained manner. Our evaluation shows that Seer improves the performance of the Intel TSX HTM by up to 2.5x, and by 62% on average, in TM benchmarks with 8 threads. These performance gains are not only a consequence of the reduced aborts, but also of the reduced activation of the HTM's pessimistic fall-back path. Nuno Diegues, Paolo Romano 0002, Stoyan Garbatov |
SPAA | 2 |
| 2015 | Enhancing Performance Prediction Robustness by Combining Analytical Modeling and Machine LearningabstractClassical approaches to performance prediction rely on two, typically antithetic, techniques: Machine Learning (ML) and Analytical Modeling (AM). ML takes a black box approach, whose accuracy strongly depends on the representativeness of the dataset used during the initial training phase. Specifically, it can achieve very good accuracy in areas of the features' space that have been sufficiently explored during the training process. Conversely, AM techniques require no or minimal training, hence exhibiting the potential for supporting prompt instantiation of the performance model of the target system. However, in order to ensure their tractability, they typically rely on a set of simplifying assumptions. Consequently, AM's accuracy can be seriously challenged in scenarios (e.g., workload conditions) in which such assumptions are not matched. Diego Didona, Francesco Quaglia, Paolo Romano 0002, Ennio Torre |
ICPE | 3 |
| 2015 | Hybrid Machine Learning/Analytical Models for Performance Prediction: A TutorialabstractClassical approaches to performance prediction of computer systems rely on two, typically antithetic, techniques: Machine Learning (ML) and Analytical Modeling (AM). Diego Didona, Paolo Romano 0002 |
ICPE | 2 |
| 2015 | Bumper: Sheltering distributed transactions from conflicts
Nuno Diegues, Paolo Romano 0002 |
Future Gener. Comput. Syst. | 2 |
| 2015 | Self-tuning Intel Restricted Transactional Memory
Nuno Diegues, Paolo Romano 0002 |
Parallel Comput. | 2 |
| 2015 | Chasing the Optimum in Replicated In-Memory Transactional Platforms via Protocol AdaptationabstractReplication plays an essential role for in-memory distributed transactional platforms, given that it represents the primary means to ensure data durability. Unfortunately, no single replication technique can ensure optimal performance across a wide range of workloads and system configurations. This paper tackles this problem by presenting MORPHR, a framework that allows to automatically adapt the replication protocol of in-memory transactional platforms according to the current operational conditions. MORPHR presents two key innovative aspects. On one hand, it allows to plug in, in a modular fashion, specialized algorithms to regulate the switching between arbitrary replication protocols. On the other hand, MORPHR relies on state of the art machine learning techniques to autonomously determine the best replication in face of varying workloads. We integrated MORPHR in an open-source in-memory NoSQL data grid, and evaluated it by means of an extensive experimental study. The results highlight that MORPHR is accurate in identifying the best replication strategy in presence of complex realistic workloads, and does so with minimal overhead. Maria Couceiro, Pedro Ruivo 0002, Paolo Romano 0002, Luís E. T. Rodrigues |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2014 | Virtues and limitations of commodity hardware transactional memoryabstractOver the last years Transactional Memory (TM) gained growing popularity as a simpler, attractive alternative to classic lock-based synchronization schemes. Recently, the TM landscape has been profoundly changed by the integration of Hardware TM (HTM) in Intel commodity processors, raising a number of questions on the future of TM. Nuno Diegues, Paolo Romano 0002, Luís E. T. Rodrigues |
PACT | 2 |
| 2014 | Enhancing Locality via Caching in the GMU ProtocolabstractGMU is a recently proposed genuine partial replication protocol for transactional systems that relies on an innovative, fully-decentralized multiversioning scheme to maximize efficiency and scalability. In this paper we tackle one of the key issues that affect the efficiency of GMU-based applications: enhancing their data locality, i.e. the ability to serve transactional reads locally, avoiding remote inter-node communication. To this end we introduce a lightweight caching mechanism that allows for safely accessing asynchronously replicated copies of remote data items while preserving GMU's original consistency criterion and its scalability. We assess the efficiency and effectiveness of the presented solution by means of an extensive experimental analysis, evaluating deployments on large scale public and private cloud infrastructures, and using well known benchmarks for transactional platforms. The results show striking speedups of up to 14 times in read dominated workloads, and a reduction of the network bandwidth by up to one order of magnitude. Hugo Pimentel, Paolo Romano 0002, Sebastiano Peluso, Pedro Ruivo 0002 |
CCGRID | 2 |
| 2014 | REAP: Reporting Errors Using Alternative Paths
João Garcia 0001, Paolo Romano 0002 |
ESOP | 3 |
| 2014 | Automatic Tuning of the Parallelism Degree in Hardware Transactional Memory
Diego Rughetti, Paolo Romano 0002, Francesco Quaglia, Bruno Ciciani |
Euro-Par | 2 |
| 2014 | STI-BT: A Scalable Transactional IndexabstractIn this article we present STI-BT, a highly scalable, transactional index for Distributed Key-Value (DKV) stores. STI-BT is organized as a distributed B+Tree and adopts an innovative design that allows to achieve high efficiency in large-scale, elastic DKV stores. We have implemented STI-BT on top of a mainstream open-source DKV store and deployed it on a public cloud infrastructure. Our extensive experimental study reveals the efficiency of our solution with demonstrable scalability in a cluster of 100 commodity machines, and speed ups with respect to state of the art solutions of up to 5.4x. Nuno Diegues, Paolo Romano 0002 |
ICDCS | 2 |
| 2014 | Performance Modelling of Partially Replicated In-Memory Transactional StoresabstractThis paper presents PROMPT, a PeRfOrmance Model for Partially replicated in-memory Transactional cloud stores. PROMPT combines white box Analytical Modelling and Machine Learning techniques, with the goal of achieving the best of the two methodologies: low training times, high extrapolation power, and portability across heterogeneous cloud infrastructures. We validate PROMPT via an extensive experimental study based on a popular open-source transactional in-memory data store (Red Hat's Infinispan), industry-standard benchmarks, and deployments on both public and private cloud infrastructures. Diego Didona, Paolo Romano 0002 |
MASCOTS | 2 |
| 2014 | Time-warp: lightweight abort minimization in transactional memoryabstractThe notion of permissiveness in Transactional Memory (TM) translates to only aborting a transaction when it cannot be accepted in any history that guarantees correctness criterion. This property is neglected by most TMs, which, in order to maximize implementation's efficiency, resort to aborting transactions under overly conservative conditions. In this paper we seek to identify a sweet spot between permissiveness and efficiency by introducing the Time-Warp Multi-version algorithm (TWM). TWM is based on the key idea of allowing an update transaction that has performed stale reads (i.e., missed the writes of concurrently committed transactions) to be serialized by committing it in the past, which we call a time-warp commit. At its core, TWM uses a novel, lightweight validation mechanism with little computational overheads. TWM also guarantees that read-only transactions can never be aborted. Further, TWM guarantees Virtual World Consistency, a safety property that is deemed as particularly relevant in the context of TM. We demonstrate the practicality of this approach through an extensive experimental study, where we compare TWM with four other TMs, and show an average performance improvement of 65% in high concurrency scenarios. Nuno Diegues, Paolo Romano 0002 |
PPoPP | 2 |
| 2014 | On the energy and performance of commodity hardware transactional memoryabstractThe advent of multi-core architectures has brought concurrent programming to the forefront of software development. In this context, Transactional Memory (TM) has gained increasing popularity as a simpler, attractive alternative to traditional lock-based synchronization. The recent integration of Hardware TM (HTM) in the last generation of Intel commodity processors turned TM into a mainstream technology, raising a number of questions on its future and that of concurrent programming. Nuno Diegues, Paolo Romano 0002, Luís E. T. Rodrigues |
SIGMETRICS | 2 |
| 2014 | Breaching the Wall of Impossibility Results on Disjoint-Access Parallel TM
Sebastiano Peluso, Roberto Palmieri, Paolo Romano 0002, Binoy Ravindran, Francesco Quaglia |
DISC | 3 |
| 2014 | On speculative replication of transactional systems
Paolo Romano 0002, Roberto Palmieri, Francesco Quaglia, Nuno Carvalho, Luís E. T. Rodrigues |
J. Comput. Syst. Sci. | 1 |
| 2014 | Transactional Auto Scaler: Elastic Scaling of Replicated In-Memory Transactional Data GridsabstractIn this article, we introduce TAS (Transactional Auto Scaler), a system for automating the elastic scaling of replicated in-memory transactional data grids, such as NoSQL data stores or Distributed Transactional Memories. Applications of TAS range from online self-optimization of in-production applications to the automatic generation of QoS/cost-driven elastic scaling policies, as well as to support for what-if analysis on the scalability of transactional applications. In this article, we present the key innovation at the core of TAS, namely, a novel performance forecasting methodology that relies on the joint usage of analytical modeling and machine learning. By exploiting these two classically competing approaches in a synergic fashion, TAS achieves the best of the two worlds, namely, high extrapolation power and good accuracy, even when faced with complex workloads deployed over public cloud infrastructures. We demonstrate the accuracy and feasibility of TAS’s performance forecasting methodology via an extensive experimental study based on a fully fledged prototype implementation integrated with a popular open-source in-memory transactional data grid (Red Hat’s Infinispan) and industry-standard benchmarks generating a breadth of heterogeneous workloads. Diego Didona, Paolo Romano 0002, Sebastiano Peluso, Francesco Quaglia |
ACM Trans. Auton. Adapt. Syst. | 2 |
| 2014 | AutoPlacer: Scalable Self-Tuning Data Placement in Distributed Key-Value StoresabstractThis article addresses the problem of self-tuning the data placement in replicated key-value stores. The goal is to automatically optimize replica placement in a way that leverages locality patterns in data accesses, such that internode communication is minimized. To do this efficiently is extremely challenging, as one needs not only to find lightweight and scalable ways to identify the right assignment of data replicas to nodes but also to preserve fast data lookup. The article introduces new techniques that address these challenges. The first challenge is addressed by optimizing, in a decentralized way, the placement of the objects generating the largest number of remote operations for each node. The second challenge is addressed by combining the usage of consistent hashing with a novel data structure, which provides efficient probabilistic data placement. These techniques have been integrated in a popular open-source key-value store. The performance results show that the throughput of the optimized system can be six times better than a baseline system employing the widely used static placement based on consistent hashing. João Paiva, Pedro Ruivo 0002, Paolo Romano 0002, Luís E. T. Rodrigues |
ACM Trans. Auton. Adapt. Syst. | 3 |
| 2014 | Design and Evaluation of a Parallel Invocation Protocol for Transactional Applications over the WebabstractEdge computing is a powerful tool to face the challenging performance requirements of modern Internet applications. By replicating applications' data and logic across a large number of geographically distributed servers, edge computing platforms allow to achieve significant enhancements of the proximity between clients and contents, and of the system scalability. These platforms reveal highly effective when handling requests entailing read-only access to the application data, as these requests can be autonomously served by some edge server typically located closer to the client than the origin site. However, in contexts where end users can trigger transactional manipulations of the application state (e.g., e-Commerce, auctions or financial applications), the corresponding update requests typically need to be redirected to the origin transactional data sources, thus, nullifying any performance benefit arising from data replication and client proximity. To cope with this issue, in this paper, we present a parallel invocation protocol, which exploits the path-diversity along the end-to-end interaction toward the origin sites by concurrently routing transactional requests toward multiple-edge servers. Request processing is finally carried out by a single-edge server, adaptively selected as the most responsive one depending on current system conditions. The proposed edge server selection scheme does not require coordination among (geographically distributed) edge server instances, thus, being very light and scalable. The benefits from our protocol in terms of both reduced and more predictable end-to-end latency are quantified via an extended simulation study. Paolo Romano 0002, Francesco Quaglia |
IEEE Trans. Computers | 1 |
| 2013 | Chasing the optimum in replicated in-memory transactional platforms via protocol adaptationabstractReplication plays an essential role for in-memory distributed transactional platforms, such as NoSQL data grids, given that it represents the primary mean to ensure data durability. Unfortunately, no single replication technique can ensure optimal performance across a wide range of workloads and system configurations. This paper tackles this problem by presenting MORPHR, a framework that allows to automatically adapt the replication protocol of in-memory transactional platforms according to the current operational conditions. MORPHR presents two key innovative aspects. On one hand, it allows to plug in, in a modular fashion, specialized algorithms to regulate the switching between arbitrary replication protocols. On the other hand, MORPHR relies on state of the art machine learning techniques to autonomously determine the optimal replication in face of varying workloads. We integrated MORPHR in a popular open-source in-memory NoSQL data grid, and evaluated it by means of an extensive experimental study. The results highlight that MORPHR is accurate in identifying the optimal replication strategy in presence of complex, realistic workloads, and does so with minimal overhead. Maria Couceiro, Pedro Ruivo 0002, Paolo Romano 0002, Luís E. T. Rodrigues |
DSN | 3 |
| 2013 | Bumper: Sheltering Transactions from ConflictsabstractThis paper addresses the issue of maximizing the efficiency and scalability of distributed transactional platforms, by introducing Bumper, a set of innovative techniques to minimize aborts of transactions in high-contention scenarios. At its core, Bumper relies on two key ideas: (1) sparing update transactions from spurious aborts when they access concurrently updated data, by attempting to serialize them in the past via a novel distributed concurrency control scheme that we call Distributed Time-Warping (DTW), and (2) avoiding aborts due to contention hot spots (that cannot be tackled by DTW) via a novel programming abstraction, called delayed actions, which allows to efficiently serialize, in an abort-free fashion, the execution of conflict-prone data manipulations. The techniques used in Bumper can be applied to a wide variety of transactional replication protocols to enhance their performance in contention intensive workloads. In this paper we show how they can be integrated with SCORe, a recent, highly-scalable genuine partial replication protocol. By means of an extensive evaluation using well-known benchmarks and a cluster of 160 nodes, we show that Bumper can boost performance up to 3x in conflict-intensive workloads, while imposing negligible (2.5%) overheads in uncontended scenarios. Nuno Diegues, Paolo Romano 0002 |
SRDS | 2 |
| 2013 | Exploiting Locality in Lease-Based Replicated Transactional Memory via Task Migration
Danny Hendler, Alex Naiman, Sebastiano Peluso, Francesco Quaglia, Paolo Romano 0002, Adi Suissa |
DISC | 5 |
| 2012 | Lightweight cooperative logging for fault replication in concurrent programsabstractThis paper presents CoopREP, a system that provides support for fault replication of concurrent programs, based on cooperative recording and partial log combination. CoopREP employs partial recording to reduce the amount of information that a given program instance is required to store in order to support deterministic replay. This allows to substantially reduce the overhead imposed by the instrumentation of the code, but raises the problem of finding the combination of logs capable of replaying the fault. CoopREP tackles this issue by introducing several innovative statistical analysis techniques aimed at guiding the search of partial logs to be combined and used during the replay phase. CoopREP has been evaluated using both standard benchmarks for multi-threaded applications and a real-world application. The results highlight that CoopREP can successfully replay concurrency bugs involving tens of thousands of memory accesses, reducing logging overhead with respect to state of the art non-cooperative logging schemes by up to 50 times in computationally intensive applications. Nuno Machado, Paolo Romano 0002, Luís E. T. Rodrigues |
DSN | 2 |
| 2012 | Topic 6: Grid, Cluster and Cloud Computing
Erik Elmroth, Paraskevi Fragopoulou, Artur Andrzejak 0001, Ivona Brandic, Karim Djemame, Paolo Romano 0002 |
Euro-Par | 6 |
| 2012 | When Scalability Meets Consistency: Genuine Multiversion Update-Serializable Partial Data ReplicationabstractIn this article we introduce GMU, a genuine partial replication protocol for transactional systems, which exploits an innovative, highly scalable, distributed multiversioning scheme. Unlike existing multiversion-based solutions, GMU does not rely on a global logical clock, which represents a contention point and can limit system scalability. Also, GMU never aborts read-only transactions and spares them from distributed validation schemes. This makes GMU particularly efficient in presence of read-intensive workloads, as typical of a wide range of real-world applications. GMU guarantees the Extended Update Serializability (EUS) isolation level. This consistency criterion is particularly attractive as it is sufficiently strong to ensure correctness even for very demanding applications (such as TPC-C), but is also weak enough to allow efficient and scalable implementations, such as GMU. Further, unlike several relaxed consistency models proposed in literature, EUS has simple and intuitive semantics, thus being an attractive, scalable consistency model for ordinary programmers. We integrated the GMU protocol in a popular open source in-memory transactional data grid, namely Infinispan. On the basis of a large scale experimental study performed on heterogeneous experimental platforms and using industry standard benchmarks (namely TPC-C and YCSB), we show that GMU achieves linear scalability and that it introduces negligible overheads (less than 10%), with respect to solutions ensuring non-serializable semantics, in a wide range of workloads. Sebastiano Peluso, Pedro Ruivo 0002, Paolo Romano 0002, Francesco Quaglia, Luís E. T. Rodrigues |
ICDCS | 3 |
| 2012 | SCORe: A Scalable One-Copy Serializable Partial Replication Protocol
Sebastiano Peluso, Paolo Romano 0002, Francesco Quaglia |
Middleware | 2 |
| 2012 | ASAP: An Aggressive SpeculAtive Protocol for Actively Replicated Transactional SystemsabstractRecent advances in the field of replicated, fault tolerant transactional systems make systematic use of Optimistic Atomic Broadcast (OAB) group communication primitives in order to coordinate the replicas. According to this scheme, the replicas gain information on the existence of transactional requests before a final and global agreement is reached on the transaction serialization order. Hence, speculative processing schemes can be exploited in order to maximize the overlap between local computation and distributed coordination activities. In this article we present ASAP, an innovative Aggressive SpeculAtive Protocol, which exhibits the following two peculiarities: (A) it allows speculating along different transaction serialization orders, thus increasing the likelihood of successful overlap between local processing and coordination in case of mismatches between the optimistic and the final delivery sequence of incoming requests, (B) it speculates along chains of conflicting transactions, tracking data dependencies among transactions via an innovative concurrency control mechanism, which allows determining in a timely fashion the alternative serialization orders to be speculatively explored. Via a simulation study in the context of Software Transactional Memory systems we show ASAP can achieve robust performance independently of the likelihood of reorder between optimistic and final deliveries, providing remarkable performance improvements (enhancing the maximum sustainable throughput up to a 2x factor) with respect to state of the art speculative replication protocols. Roberto Palmieri, Francesco Quaglia, Paolo Romano 0002 |
NCA | 3 |
| 2012 | SPECULA: Speculative Replication of Software Transactional MemoryabstractThis paper introduces SPECULA, a novel replication protocol for Software Transactional Memory (STM) systems that seeks maximum overlap between transaction execution and replica synchronization phases via speculative processing techniques. By removing the replica synchronization phase from the critical path of execution of transactions, SPECULA allows threads to speculatively pipeline the execution of both transactional and/or non-transactional code. The core of SPECULA is a multi-version concurrency control algorithm that supports speculative transaction processing while ensuring the strong consistency criteria that are desirable in non-sand-boxed environments like STMs. Via an experimental study, based on a fully-fledged prototype and on both synthetic and standard STM benchmarks, we demonstrate that SPECULA can achieve speedups of up to one order of magnitude with respect to state-of-the-art non-speculative replication techniques. Sebastiano Peluso, Joao Fernandes, Paolo Romano 0002, Francesco Quaglia, Luís E. T. Rodrigues |
SRDS | 3 |
| 2012 | On the analytical modeling of concurrency control algorithms for Software Transactional Memories: The case of Commit-Time-Locking
Pierangelo di Sanzo, Bruno Ciciani, Roberto Palmieri, Francesco Quaglia, Paolo Romano 0002 |
Perform. Evaluation | 5 |
| 2011 | PolyCert: Polymorphic Self-optimizing Replication for In-Memory Transactional Grids
Maria Couceiro, Paolo Romano 0002, Luís E. T. Rodrigues |
Middleware | 2 |
| 2011 | A Generic Framework for Replicated Software Transactional MemoriesabstractSoftware Transactional Memory (STM) has emerged a powerful abstraction for managing access to shared data. Therefore, it is no surprise that a handful of different STM replication schemes have been proposed in the last recent years. In this context, we propose an architecture that facilitates the integration and execution of multiple replication techniques in a single, coherent, middleware infrastructure. This paves the way towards the development of autonomic mechanisms, able to select in runtime the most appropriate replication technique for the workload at hand. Nuno Carvalho, Paolo Romano 0002, Luís E. T. Rodrigues |
NCA | 2 |
| 2011 | Exploiting Total Order Multicast in Weakly Consistent Transactional CachesabstractNowadays, distributed in-memory caches are increasingly used as a way to improve the performance of applications that require frequent access to large amounts of data. In order to maximize performance and scalability, these platforms typically rely on weakly consistent partial replication mechanisms. These schemes partition the data across the nodes and ensure a predefined (and typically very small) replication degree, thus maximizing the global memory capacity of the platform and ensuring that the cost to ensure replica consistency remains constant as the scale of the platform grows. Moreover, even though several of these platforms provide transactional support, they typically sacrifice consistency, ensuring guarantees that are weaker than classic 1-copy serializability, but that allow for more efficient implementations. This paper proposes and evaluates two partial replication techniques, providing different (weak) consistency guarantees, but having in common the reliance on total order multicast primitives to serialize transactions without incurring in distributed deadlocks, a main source of inefficiency of classical two-phase commit (2PC) based replication mechanisms. We integrate the proposed replication schemes into Infinispan, a prominent open-source distributed in-memory cache, which represents the reference clustering solution for the well-known JBoss AS platform. Our performance evaluation highlights speed-ups of up to 40× when using the proposed algorithms with respect to the native Infinispan replication mechanism, which relies on classic 2PC-based replication. Pedro Ruivo 0002, Maria Couceiro, Paolo Romano 0002, Luís E. T. Rodrigues |
PRDC | 3 |
| 2011 | OSARE: Opportunistic Speculation in Actively REplicated Transactional SystemsabstractIn this work we present OSARE, an active replication protocol for transactional systems that combines the usage of Optimistic Atomic Broadcast with a speculative concurrency control mechanism in order to overlap transaction processing and replica synchronization. OSARE biases the speculative serialization of transactions towards an order aligned with the optimistic message delivery order. However, due to the lock-free nature of its concurrency control algorithm, at high concurrency levels, namely when the probability of mismatches between optimistic and final deliveries is higher, OSARE explores additional alternative transaction serialization orders in a lightweight and opportunistic fashion. A simulation study we carried out in the context of Software Transactional Memory systems shows that OSARE achieves robust performance also in scenarios characterized by non-minimal likelihood of reorder between optimistic and final deliveries, providing remarkable speed-up with respect to state of the art speculative replication protocols. Roberto Palmieri, Francesco Quaglia, Paolo Romano 0002 |
SRDS | 3 |
| 2011 | SCert: Speculative certification in replicated software transactional memoriesabstractBeing much simpler to compose and verify than classical lock based synchronization schemes, Software Transactional Memories (STMs) have emerged as an attractive paradigm for supporting concurrent access to in-memory storage systems. This paper is focused on the issue of how to replicate STMs to enhance both their performance and dependability. This is an extremely challenging problem, since the communication/processing ratio in STMs is typically several orders of magnitude higher than in conventional database systems, thus amplifying the relative cost of replication. Nuno Carvalho, Paolo Romano 0002, Luís E. T. Rodrigues |
SYSTOR | 2 |
| 2011 | Providing e-Transaction Guarantees in Asynchronous Systems with No Assumptions on the Accuracy of Failure DetectionabstractIn this paper, we address reliability issues in three-tier systems with stateless application servers. For these systems, a framework called e-Transaction has been recently proposed, which specifies a set of desirable end-to-end reliability guarantees. In this article, we propose an innovative distributed protocol providing e-Transaction guarantees in the general case of multiple, autonomous back-end databases (typical of scenarios with multiple parties involved within a same business process). Differently from existing proposals coping with the e-Transaction framework, our protocol does not rely on any assumption on the accuracy of failure detection. Hence, it reveals suited for a wider class of distributed systems. To achieve such a target, our protocol exploits an innovative scheme for distributed transaction management (based on ad hoc demarcation and concurrency control mechanisms), which we introduce in this paper. Beyond providing the proof of protocol correctness, we also discuss hints on the protocol integration with conventional systems (e.g., database systems) and show the minimal overhead imposed by the protocol. Paolo Romano 0002, Francesco Quaglia |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2010 | An Optimal Speculative Transactional Replication ProtocolabstractIn this paper we investigate the problem of speculative processing in a replicated transactional system layered on top of an optimistic atomic broadcast service. We consider a realistic model in which transactions' read/write sets are not known a-priori, and transactions' data access patterns may vary depending on the observed snapshot. We formalize a set of correctness and optimality properties aimed at ensuring that transactions are not activated on inconsistent snapshots, as well as the minimality and completeness of the set of explored serialization orders. Finally, an optimal speculative transaction replication protocol is presented. Paolo Romano 0002, Roberto Palmieri, Francesco Quaglia, Nuno Carvalho, Luís E. T. Rodrigues |
ISPA | 1 |
| 2010 | Asynchronous Lease-Based Replication of Software Transactional Memory
Nuno Carvalho, Paolo Romano 0002, Luís E. T. Rodrigues |
Middleware | 2 |
| 2010 | AGGRO: Boosting STM Replication via Aggressively Optimistic Transaction ProcessingabstractSoftware Transactional Memories (STMs) are emerging as a potentially disruptive programming model. In this paper we are address the issue of how to enhance dependability of STM systems via replication. In particular we present AGGRO, an innovative Optimistic Atomic Broadcast-based (OAB) active replication protocol that aims at maximizing the overlap between communication and processing through a novel AGGRessively Optimistic concurrency control scheme. The key idea underlying AGGRO is to propagate dependencies across uncommitted transactions in a controlled manner, namely according to a serialization order compliant with the optimistic message delivery order provided by the OAB service. Another relevant distinguishing feature of AGGRO is of not requiring a-priori knowledge about read/write sets of transactions, but rather to detect and handle conflicts dynamically, i.e. as soon (and only if) they materialize. Based on a detailed simulation study we show the striking performance gains achievable by AGGRO (up to 6x increase of the maximum sustainable throughput, and 75% response time reduction) compared to literature approaches for active replication of transactional systems. Roberto Palmieri, Francesco Quaglia, Paolo Romano 0002 |
NCA | 3 |
| 2010 | Brief announcement: on speculative replication of transactional systemsabstractWe define the problem of speculative processing in a replicated transactional system layered on top of an optimistic atomic broadcast service. A realistic model is considered in which transactions' read and write sets are not a priori known and transactions' data access patterns may vary depending on the observed snapshot. We formalize a set of correctness and optimality properties ensuring the minimality and completeness of the set of explored serialization orders within the replicated transactional system. Paolo Romano 0002, Roberto Palmieri, Francesco Quaglia, Nuno Carvalho, Luís E. T. Rodrigues |
SPAA | 1 |
| 2009 | APART+: Boosting APART performance via optimistic pipelining of output eventsabstractAPART (A Posteriori Active ReplicaTion) is a recently proposed active replication protocol specifically tailored for multi-tier data acquisition systems. It ensures consistency of middle-tier sink replicas by means of an a-posteriori synchronization phase based on reconciliation, which is activated only in case replicas react to an input message from the sensors by generating an output event destined to the back-end tier. This paper enhances APART via a novel non-blocking synchronization scheme which prevents replicas from stalling while waiting for the outcome of an on-going synchronization phase. Contrarily, replicas are allowed to optimistically process data from the sensors, and to immediately propagate any output event towards the back-end tier. The removal of the blocking synchronization phase from the critical path gives rise to striking performance gains via an effective overlapping of event processing and synchronization. On the other hand, system consistency is ensured by enhancing the back-end tier synchronization logic in order to filter out optimistically produced output events that are incompatible with the reconciled state trajectory. Paolo Romano 0002, Francesco Quaglia, Bruno Ciciani |
IPDPS | 1 |
| 2009 | The Weak Mutual Exclusion problemabstractIn this paper we define the weak mutual exclusion (WME) problem. Analogously to classical distributed mutual exclusion (DME), WME serializes the accesses to a shared resource. Differently from DME, however, the WME abstraction regulates the access to a replicated shared resource, whose copies are locally maintained by every participating process. Also, in WME, processes suspected to have crashed are possibly ejected from the critical section. We prove that, unlike DME, WME is solvable in a partially synchronous model, i.e. a system where the bounds on communication latency and on relative process speeds are not known in advance, or are known but only hold after an unknown time. Finally, we demonstrate that diam P is the weakest failure detector for solving WME, and present an algorithm that solves WME using diam P with a majority of correct processes. Paolo Romano 0002, Luís E. T. Rodrigues, Nuno Carvalho |
IPDPS | 1 |
| 2009 | An Efficient Weak Mutual Exclusion AlgorithmabstractThe Weak Mutual Exclusion (WME) is a recently proposed abstraction which, analogously to classical Distributed Mutual Exclusion (DME), permits to serialize concurrent accesses to a shared resource. Unlike DME, however, the WME abstraction regulates the access to a replicated shared resource and is solvable in the presence of less restrictive synchrony assumptions, i.e. in an asynchronous system augmented with an eventually perfect failure detector. This paper presents an efficient WME algorithm which outperforms previous solutions in terms of both communication latency and message complexity, while relying on minimal synchrony assumptions. Paolo Romano 0002, Luís E. T. Rodrigues |
ISPDC | 1 |
| 2009 | D2STM: Dependable Distributed Software Transactional MemoryabstractAt current date the problem of how to build distributed and replicated software transactional memory (STM) to enhance both dependability and performance is still largely unexplored. This paper fills this gap by presenting D2STM, a replicated STM whose consistency is ensured in a transparent manner, even in the presence of failures. Strong consistency is enforced at transaction commit time by a non-blocking distributed certification scheme, which we name BFC (bloom filter certification). BFC exploits a novel bloom filter-based encoding mechanism that permits to significantly reduce the overheads of replica coordination at the cost of a user tunable increase in the probability of transaction abort. Through an extensive experimental study based on standard STM benchmarks we show that the BFC scheme permits to achieve remarkable performance gains even for negligible (e.g. 1%) increases of the transaction abort rate. Maria Couceiro, Paolo Romano 0002, Nuno Carvalho, Luís E. T. Rodrigues |
PRDC | 2 |
| 2008 | Integration and evaluation of Multi-Instance-Precommit schemes within postgreSQLabstractMulti-instance-precommit (MIP) has been recently presented as an innovative transaction management scheme in support of reliability for Atomic Transactions in multitier (e.g. Web-based) systems. With this scheme, fail-over of a previously activated transaction can be supported via simple retry logics, which do not require knowledge about whether, and on which sites, the original transaction was precommitted. Mutual deadlock between the original and the retried transaction are prevented via MIP facilities, which also support reconciliation mechanisms for at-most-once transaction execution semantic. In this article we present an extension of the open source PostgreSQL database system in order to support MIP. The extension is based on the exploitation of PostgreSQL native multiversion concurrency control scheme. We also present an experimental evaluation based on the TPC-W benchmark, aimed at quantifying the relative overhead of MIP facilities on transaction execution latency, system throughput and storage usage. Paolo Romano 0002, Francesco Quaglia |
DSN | 1 |
| 2008 | Accuracy vs efficiency of hyper-exponential approximations of the response time distribution of MMPP/M/1 queuesabstractThe Markov modulated Poisson process (MMPP) has been shown to well describe the flow of incoming traffic in networked systems, such as the Grid and the WWW. This makes the MMPP/M/1 queue a valuable instrument to evaluate and predict the service level of networked servers. In a recent work we have provided an approximate solution for the response time distribution of the MMPP/M/1 queue, which is based on a weighted superposition of M/M/l queues (i.e. a hyper-exponential process). In this article we address the tradeoff between the accuracy of this approximation and its computational cost. By jointly considering both accuracy and cost, we identify the scenarios where such approximate solution could be effectively used in support of network servers (dynamic) configuration and evaluation strategies, aimed at ensuring the agreed dependability levels in case of, e.g., request redirection due to faults. Paolo Romano 0002, Bruno Ciciani, Andrea Santoro, Francesco Quaglia |
IPDPS | 1 |
| 2008 | A Performance Model of Multi-Version Concurrency Control
Pierangelo di Sanzo, Bruno Ciciani, Francesco Quaglia, Paolo Romano 0002 |
MASCOTS | 4 |
| 2008 | APART: Low Cost Active Replication for Multi-tier Data Acquisition SystemsabstractThis paper proposes APART (a posteriori active replication), a novel active replication protocol specifically tailored for multi-tier data acquisition systems. Unlike existing active replication solutions, APART does not rely on a-priori coordination schemes determining a same schedule of events across all the replicas, but it ensures replicas consistency by means of an a-posteriori reconciliation phase. The latter is triggered only in case the replicated servers externalize their state by producing an output event towards a different tier. On one hand, this allows coping with non-deterministic replicas, unlike existing active replication approaches. On the other hand, it allows attaining striking performance gains in the case of silent replicated servers, which only sporadically, yet unpredictably, produce output events in response to the receipt of a (possibly large) volume of input messages. This is a common scenario in data acquisition systems, where sink processes, which filter and/or correlate incoming sensor data, produce output messages only if some application relevant event is detected. Further, the APART replica reconciliation scheme is extremely lightweight as it exploits the cross-tier communication pattern spontaneously induced by the application logic to avoid explicit replicas coordination messages. Paolo Romano 0002, Diego Rughetti, Francesco Quaglia, Bruno Ciciani |
NCA | 1 |
| 2007 | A Lightweight Heuristic-based Mechanism for Collecting Committed Consistent Global States in Optimistic SimulationabstractIn this paper we study how to reuse checkpoints taken in an uncorrelated manner during the forward execution phase in an optimistic simulation system in order to construct global consistent snapshots which are also committed (i.e. the logical time they refer to is lower than the current GVT value). This is done by introducing a heuristic-based mechanism relying on update operations applied to local committed checkpoints of the involved logical processes so to eliminate mutual dependencies among the final achieved state values. The mechanism is lightweight since it does not require any form of (distributed) coordination to determine which are the checkpoint update operations to be performed. At the same time it is likely to reduce the amount of checkpoint update operations required to realign the consistent global state exactly to the current GVT value, taken as the reference time for the snapshot. Our proposal can support, in a performance effective manner, termination detection schemes based on global predicates evaluated on a committed and consistent global snapshot, which represent an alternative as relevant as classical termination check only relying on the current GVT value. Another application concerns interactive simulation environments, where (aggregate) output information about committed and consistent snapshots needs to be frequently provided, hence requiring lightweight mechanisms for the construction of the snapshots. Diego Cucuzzo, Stefano D'Alessio, Francesco Quaglia, Paolo Romano 0002 |
DS-RT | 4 |
| 2007 | Approximate Analytical Models for Networked Servers Subject to MMPP Arrival ProcessesabstractInput characterization to describe the flow of incoming traffic in network systems, such as the GRID and the WWW, is often performed by using Markov modulated Poisson processes (MMPP). Therefore, to enact capacity planning and quality-of-service (QoS) oriented design, the model of the hosts that receive the incoming traffic is often described as a MMPP/M/1 queue. The drawback of this model is that no closed form for its solution has been derived. This means that evaluating even the simplest output statistics of the model, such as the average response times of the queue, is a computationally intensive task and its usage in the above contexts is often unadvisable. In this paper we discuss the possibility to approximate the behavior of a MMPP/M/1 queue with a computational effective analytical approximation, thus saving the large amount of calculations required to evaluate the same data by other means. The employed method consists in approximating the MMPP/M/1 queue as a weighted superposition of different M/M/1 queues. The analysis is validated by comparing the results of a discrete event simulator with those obtained from the proposed approximations, in the context of a real case study involving a GRID networked server. Bruno Ciciani, Andrea Santoro, Paolo Romano 0002 |
NCA | 3 |
| 2007 | Ensuring e-Transaction with Asynchronous and Uncoordinated Application Server ReplicasabstractA recently proposed abstraction, called e-transaction (exactly-once transaction), specifies a set of properties capturing end-to-end reliability aspects for three-tier Web-based systems. In this paper we propose a distributed protocol ensuring the e-transaction properties for the general case of multiple, autonomous back-end databases. The key idea underlying our proposal consists in distributing, across the back-end tier, some recovery information reflecting the transaction processing state. This information is manipulated at low cost via local operations at the database side, with no need for any form of coordination among asynchronous replicas of the application server within the middle-tier. Compared to existing solutions, our protocol has therefore the distinguishing features of being both very light and highly scalable. The latter aspect makes our proposal particularly attractive for the case of very high degree of replication of the application access point, with distribution of the replicas within infrastructures geographically spread on public networks over the Internet (e.g., application delivery networks), namely, a configuration that also provides the advantages of reduced user perceived latency and increased system availability Francesco Quaglia, Paolo Romano 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2006 | A simulation study of the effects of multi-path approaches in e-commerce applicationsabstractResponse time is a key factor of any e-commerce application, and a set of solutions have been proposed to provide low response time despite network congestions or failures. Being them mostly based on caching of Web objects and replication of DBMS managed data at the edges, or at intermediate points, of the Web infrastructure, they reveal effective when handling client requests only performing read access to application data. However, any update request typically needs to be redirected to the origin DBMSs, hence not taking advantage from data replication and related client proximity. In order to alleviate the effects of network congestions or failures, we have proposed a multi-path protocol that increases the likelihood for the update request to be processed along a responsive (e.g. failure free) network path in between the client location and the origin DBMS sites. In this paper we present an extensive simulation study of the effects of such a multi-path approach on the client perceived response time. The study relies on both Brite generated network topologies and the NLANR graph. Also, well known realistic TCP models are used to capture the effects of network delays during both normal and anomalous (i.e. packet loss affected) operation mode Paolo Romano 0002, Francesco Quaglia, Bruno Ciciani |
IPDPS | 1 |
| 2006 | Providing e-Transaction Guarantees in Asynchronous Systems with Inaccurate Failure DetectionabstractIn this paper we address reliability issues in Web-based transactional systems. We are interested in the category of systems characterized by stateless application servers. For these systems, a framework called e-Transaction has been recently proposed, which specifies a set of desirable end-to-end reliability guarantees. Within this framework we propose an innovative distributed protocol providing those reliability guarantees in the general case of multiple, autonomous back-end databases (typical of scenarios with multiple parties involved within a same business process). Compared to existing proposals coping with the e-Transaction framework, our protocol adopts a weaker approach to failure detection, i.e. it does not rely on any assumption on the accuracy of failure detection. Hence it reveals suited for a wider class of distributed systems, including those systems where the level of asynchrony makes stronger approaches to failure detection not feasible in practice. To achieve such a target, our protocol exploits an innovative scheme for distributed transaction management (based on ad-hoc demarcation and concurrency control mechanisms), which we introduce in this paper. We also provide hints on the protocol integration with conventional systems (e.g. database systems) Paolo Romano 0002, Francesco Quaglia |
NCA | 1 |
| 2005 | A Lightweight and Scalable e-Transaction Protocol for Three-Tier Systems with Centralized Back-End DatabaseabstractThe e-transaction abstraction is a recent formalization of end-to-end reliability properties for three-tier systems. In this work, we present a protocol ensuring the e-transaction guarantees in case the back-end tier consists of a centralized database. Our proposal addresses the case of stateless application servers, and is both simple and effective since 1) it does not employ any distributed commit protocol and 2) does not require coordination among the replicas of the application server. Paolo Romano 0002, Francesco Quaglia, Bruno Ciciani |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2004 | Ensuring E-Transaction Through a Lightweight Protocol for Centralized Back-End Database
Paolo Romano 0002, Francesco Quaglia, Bruno Ciciani |
ISPA | 1 |
| 2004 | A Protocol for Improved User Perceived QoS in Web Transactional ApplicationsabstractQuality-of-service (QoS) provisioning in the Internet has been a topic of active research in the last few years. However, due to both financial and technical reasons, the proposed solutions are not commonly employed in practice. As a consequence, the Internet architecture is still mainly oriented to a best effort delivery model, which does not provide any guarantee neither on the message delivery latency, nor on the probability that a service residing at some host becomes temporarily unreachable due to network congestion. We address this issue by presenting an innovative, application level protocol tailored for Web transactional applications, which attempts to reduce the impact of network congestion on the latency experienced by the end-users. The intuition underlying our proposal is to exploit the intrinsic potential of parallelism commonly exhibited by application service provider (ASP) infrastructures, where the application access point is replicated over a large number of geographically distributed edge servers. At this purpose, we allow privileged classes of users to concurrently contact multiple, replicated access points so to increase the probability to timely reach at least one of them and to promptly activate the application business logic for the interaction with a back-end database system. We complete our proposal with an efficient mechanism that prevents multiple, undesired updates on the back-end database and, at the same time, strongly limits the additional load on the ASP infrastructure due to the increased amount of requests from the privileged users. Paolo Romano 0002, Francesco Quaglia, Bruno Ciciani |
NCA | 1 |
| 2003 | Validiation of the Sessionless Mode of the HTTPR Protocol
Paolo Romano 0002, Milton Romero, Bruno Ciciani, Francesco Quaglia |
FORTE | 1 |