EDBT 2026 Demo / reviewers in the wild / expert
Marco Paolieri
dblp:07/165
· DBLP profile ↗
35ranked-venue papers
10as first author
11since 2021 · last 2026
0000-0001-5110-203XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 7 first-author · 6 since 2021Software engineering, systems software and programming languages · 10 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Systems for AI: Predicting Performance of Machine Learning Workloads
Zhuojin Li, Marco Paolieri, Leana Golubchik |
ICPE | 2 |
| 2026 | Joint offloading and service selection via matching and auction theory for multi-task dependent computation-intensive applications
Benedetta Picano, Marco Paolieri, Laura Carnevali, Enrico Vicario |
Perform. Evaluation | 2 |
| 2025 | Analytical Characterization and Efficient Simulation of Batched Arrivals in the Kafka BrokerabstractApache Kafka is a key component in event-driven and microservice architectures relying on distributed publishsubscribe messaging for scalable and fault-tolerant streaming of real-time data. To reduce distribution overhead, messages are buffered and dispatched to the broker when either a maximum batch size N is reached or a timeout T expires, enabling control on the trade-off between high throughput and low latency. However, this trade-off has been explored only through empirical studies, referred to specific system deployments and not suited for runtime adaptation to variable workload conditions. We provide an analytical characterization of the arrival process induced by Kafka batching policy under Poisson arrivals. The analysis develops on the observation that the time for buffering a full batch and the size of a batch dispatched at expiration of the timeout follow truncated Erlang and Poisson distributions, respectively. Leveraging this insight, we derive closed forms for quantities that characterize the arrival process, and we propose a method for efficient simulation of the process embedded at dispatching times. To support practical implementation, we evaluate solutions for drawing samples from an Erlang distribution provided by NumPy, PyTorch, and R, and we also propose a novel approach based on rejection sampling with proposal function in the family of Kumaraswamy distributions with automated optimization of parameters with respect to the number of phases. Numerical experimentation shows that: (i) aggregated simulation enabled by the analytical formulation is insensitive to the value of the batch size, and it definitely outperforms fine-grained simulation of individual message arrivals; (ii) the best efficiency is obtained with the NumPy implementation of Erlang, with promising results of the novel approach based on Kumaraswamy, which achieves results comparable to PyTorch and better than $\mathbf{R}$. András Horváth, Marco Paolieri, Benedetta Picano, Enrico Vicario |
MASCOTS | 2 |
| 2025 | Approximation of cumulative distribution functions by Bernstein phase-type distributionsabstractThe inclusion of generally distributed random variables in stochastic models is often tackled by choosing a parametric family of distributions and applying fitting algorithms to find appropriate parameters. A recent paper proposed the approximation of probability density functions (PDFs) by Bernstein exponentials, which are obtained from Bernstein polynomials by a change of variable and result in a particular case of acyclic phase-type distributions. In this paper, we show that this approximation can also be applied to cumulative distribution functions (CDFs), which enjoys advantageous properties and achieves similar accuracy; by focusing on CDFs, we propose an approach to obtain stochastically ordered approximations. The use of a scaling parameter in the approximation is also presented, evaluating its effect on approximation accuracy. András Horváth, Illés Horváth, Marco Paolieri, Miklós Telek, Enrico Vicario |
Perform. Evaluation | 3 |
| 2025 | Compositional Coordinated Resource Provisioning in Workflows With Stochastic DurationsabstractIn performance engineering of composed services, coordinated provisioning can reduce the amount of resources required to meet end-to-end response time objectives. To this aim, various intertwined aspects of the application architecture need to be taken into account, notably including precedence constraints in the composition of elementary services, along with their durations and sensitivity to the scaling of provisioned resources. We address coordinated provisioning of resources for elementary services with stochastic durations with general distributions (i.e., including non-exponential distributions). We compose services in a workflow where precedence constraints define a Directed Acyclic Graph (DAG) and the distribution of the end-to-end (E2E) response time is subject to a Service Level Objective (SLO). We leverage a surrogate model of service performance, assuming a low workload of workflow requests (i.e., a single-request scenario) and service durations inversely proportional to provisioned resources. Given the total amount of resources, our approach derives the service provisioning that optimizes the workflow E2E response time distribution, by exploiting a compositional approach and by using stochastically ordered approximations to manage dependencies in non-well-nested precedence DAGs. Then, the approach scales provisioned resources up or down to determine the minimum amount of resources needed to satisfy the SLO, while leaving the remaining resources for horizontal scaling in order to manage multiple workflow requests at high workloads. Experiments consider low-workload and high-workload scenarios, different relations between elementary service durations and provisioned resources, and workflow topologies taken from benchmarks or randomly generated with controlled statistics, using elementary service durations from a dataset of the literature. Results show that the technique is feasible also for workflows with a thousand of services and that it outperforms other provisioning methods in fitting the SLO using the same resource amount and in minimizing the resource amount needed to fit the SLO. Laura Carnevali, Marco Paolieri, Riccardo Reali, Leonardo Scommegna, Enrico Vicario |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | Inference latency prediction for CNNs on heterogeneous mobile devices and ML frameworks
Zhuojin Li, Marco Paolieri, Leana Golubchik |
Perform. Evaluation | 2 |
| 2023 | Predicting Inference Latency of Neural Architectures on Mobile DevicesabstractDue to the proliferation of inference tasks on mobile devices, state-of-the-art neural architectures are typically designed using Neural Architecture Search (NAS) to achieve good tradeoffs between machine learning accuracy and inference latency. While measuring inference latency of a huge set of candidate architectures during NAS is not feasible, latency prediction for mobile devices is challenging, because of hardware heterogeneity, optimizations applied by machine learning frameworks, and diversity of neural architectures. Motivated by these challenges, we first quantitatively assess the characteristics of neural architectures and mobile devices that have significant effects on inference latency. Based on this assessment, we propose an operation-wise framework which addresses these challenges by developing operation-wise latency predictors and achieves high accuracy in end-to-end latency predictions, as shown by our comprehensive evaluations on multiple mobile devices using multicore CPUs and GPUs. To illustrate that our approach does not require expensive data collection, we also show that accurate predictions can be achieved on real-world neural architectures using only small amounts of profiling data. Zhuojin Li, Marco Paolieri, Leana Golubchik |
ICPE | 2 |
| 2022 | Performance and Revenue Analysis of Hybrid Cloud Federations with QoS RequirementsabstractHybrid cloud architectures, where private clouds or data centers forward part of their workload to public cloud providers to satisfy quality of service (QoS) requirements, are increasingly common due to the availability of on-demand cloud resources that can be provisioned automatically through programming APIs. In this paper, we analyze performance and revenue in federations of hybrid clouds, where private clouds agree to share part of their local computing resources with other members of the federation. Through resource sharing, underprovisioned members can save on public cloud costs, while overprovisioned members can put their idle resources to work. To reward all hybrid clouds for their contributions (computing resources or workload), public cloud savings due to the federation are distributed among members according to Shapley value.We model this cloud architecture with a continuous-time Markov chain and prove that, if all hybrid clouds have the same QoS requirements, their profits are maximized when they join the federation and share all resources. We also show that this result does not hold when hybrid clouds have different QoS requirements, and we provide a solution to evaluate profit for different resource sharing decisions. Finally, our experimental evaluation compares the distribution of public cloud savings according to Shapley value with alternative approaches, illustrating its ability to discourage free riders of the federation. Marco Paolieri, Leana Golubchik |
CLOUD | 2 |
| 2022 | Defending against Poisoning Backdoor Attacks on Federated Meta-learningabstractFederated learning allows multiple users to collaboratively train a shared classification model while preserving data privacy. This approach, where model updates are aggregated by a central server, was shown to be vulnerable to poisoning backdoor attacks : a malicious user can alter the shared model to arbitrarily classify specific inputs from a given class. In this article, we analyze the effects of backdoor attacks on federated meta-learning , where users train a model that can be adapted to different sets of output classes using only a few examples. While the ability to adapt could, in principle, make federated learning frameworks more robust to backdoor attacks (when new training examples are benign), we find that even one-shot attacks can be very successful and persist after additional training. To address these vulnerabilities, we propose a defense mechanism inspired by matching networks , where the class of an input is predicted from the similarity of its features with a support set of labeled examples. By removing the decision logic from the model shared with the federation, the success and persistence of backdoor attacks are greatly reduced. Chien-Lun Chen, Sara Babakniya, Marco Paolieri, Leana Golubchik |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2022 | Predicting Throughput of Distributed Stochastic Gradient DescentabstractTraining jobs of deep neural networks (DNNs) can be accelerated through distributed variants of stochastic gradient descent (SGD), where multiple nodes process training examples and exchange updates. The total throughput of the nodes depends not only on their computing power, but also on their networking speeds and coordination mechanism (synchronous or asynchronous, centralized or decentralized), since communication bottlenecks and stragglers can result in sublinear scaling when additional nodes are provisioned. In this paper, we propose two classes of performance models to predict throughput of distributed SGD:fine-grained models, representing many elementary computation/communication operations and their dependencies; andcoarse-grained models, where SGD steps at each node are represented as a sequence of high-level phases without parallelism between computation and communication. Using a PyTorch implementation, real-world DNN models and different cloud environments, our experimental evaluation illustrates that, while fine-grained models are more accurate and can be easily adapted to new variants of distributed SGD, coarse-grained models can provide similarly accurate predictions when augmented with ad hoc heuristics, and their parameters can be estimated with profiling information that is easier to collect. Zhuojin Li, Marco Paolieri, Leana Golubchik, Sung-Han Lin, Wumo Yan |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | The ORIS Tool: Quantitative Evaluation of Non-Markovian SystemsabstractWe present the next generation of ORIS, a toolbox for quantitative evaluation of concurrent models with non-Markovian timers. The tool shifts its focus from timed models to stochastic ones, it includes a new graphical user interface, new analysis methods and a Java Application Programming Interface (API). Models can be specified as Stochastic Time Petri Nets (STPNs) through the graphical editor, validated using an interactive token game, and analyzed through several techniques to compute instantaneous or cumulative rewards. STPNs can also be exported as Java code to conduct extensive parametric studies through the Java library, now distributed as open-source. A well-engineered software architecture allows the user to implement new features for STPNs, new modeling formalisms, and new analysis methods. The most distinctive features of ORIS include transient and steady-state analysis of STPNs modeling Markov Regenerative Processes (MRPs), and transient analysis of STPNs modeling generalized semi-Markov processes. ORIS also supports state-space analysis of Time Petri Nets (TPNs), simulation of STPNs, and standard analysis techniques for continuous-time Markov chains or MRPs with at most one non-exponential timer in each state. We illustrate the general workflow for the application of ORIS to the modeling and evaluation of non-functional requirements of software-intensive systems. Marco Paolieri, Marco Biagi, Laura Carnevali, Enrico Vicario |
IEEE Trans. Software Eng. | 1 |
| 2020 | Emulating and Verifying Sensing, Computation, and Communication in Distributed Remote Sensing SystemsabstractThis paper presents updates to the Virtual Constellation Engine (VCE), a simulator and emulator which enables modeling of sensing, computation, and communication of multi-platform remote sensing systems. Users can launch heterogeneous constellations, describe different on-board processing configurations, and simulate and verify autonomous operations. Updates in the past year include emulation of on-board computing utilizing cloud based processors, inclusion of the Delay Tolerant Networking (DTN) protocol in the communication modeling, and visualization and diagnostic tools which facilitate analysis of the distributed system. To illustrate the benefits, VCE was utilized in the development of on-board processing implementation of a new, high data rate sensor, Multispectral, Imaging, Detection, and Active Reflectance (MiDAR), for both an embedded Nvidia Tegra GPU and a Xilinx MPSoC FPGA achieving a 7.6x speedup over a desktop workstation and enabling this application for integration into a distributed sensing system. Matthew French, Marco Paolieri, Vivek V. Menon, Andrew G. Schmidt |
IGARSS | 2 |
| 2020 | Throughput Prediction of Asynchronous SGD in TensorFlowabstractModern machine learning frameworks can train neural networks using multiple nodes in parallel, each computing parameter updates with stochastic gradient descent (SGD) and sharing them asynchronously through a central parameter server. Due to communication overhead and bottlenecks, the total throughput of SGD updates in a cluster scales sublinearly, saturating as the number of nodes increases. In this paper, we present a solution to predicting training throughput from profiling traces collected from a single-node configuration. Our approach is able to model the interaction of multiple nodes and the scheduling of concurrent transmissions between the parameter server and each node. By accounting for the dependencies between received parts and pending computations, we predict overlaps between computation and communication and generate synthetic execution traces for configurations with multiple nodes. We validate our approach on TensorFlow training jobs for popular image classification neural networks, on AWS and on our in-house cluster, using nodes equipped with GPUs or only with CPUs. We also investigate the effects of data transmission policies used in TensorFlow and the accuracy of our approach when combined with optimizations of the transmission schedule. Zhuojin Li, Wumo Yan, Marco Paolieri, Leana Golubchik |
ICPE | 3 |
| 2019 | Constellations in the Cloud: Virtualizing Remote Sensing SystemsabstractThis work presents the Virtual Constellation Engine (VCE), a software framework and run-time system designed to facilitate exploration of different remote sensing constellation and sensor web configurations from an on-board processing perspective, using the cloud. Users can launch heterogeneous constellations, describe different on-board processing configurations, and simulate and verify autonomous operations. VCE eases development for applications leveraging complex compute resources and has been demonstrated on Amazon Web Services to provide speedups over 20,000× of state of practice systems. Andrew G. Schmidt, Vivek Venugopalan, Marco Paolieri, Matthew French |
IGARSS | 3 |
| 2019 | A Continuous-Time Model-Based Approach for Activity Recognition in Pervasive EnvironmentsabstractWe present a model-based approach to Activity Recognition (AR) in Ambient Assisted Living (AAL). The approach leverages an a priori stochastic model termed Continuous-Time Hidden Semi-Markov Model (CT-HSMM), capturing the continuous-time durations of activities and inter-event times. The model is enhanced according to the observed statistics, associating the events with an occurrence probability, and the sojourn time and the inter-event time in each activity with a continuous-time probability density function, allowing effective fitting of observed durations through non-Markovian distributions. The model is updated at run time according to a sequence of time-stamped observations, exploiting the method of stochastic state classes to perform transient analysis and derive a measure of likelihood that an activity is currently performed. The approach supports both online AR, predicting the activity performed at time t using only the events observed until that time, and offline AR, applying a forward- backward procedure that exploits all the events observed before and after time t. The approach is experimented on a real dataset of the literature, providing performance measures that can be compared with those of offline Hidden Markov Models (HMMs) and offline Hidden Semi-Markov Models (HSMMs). Marco Biagi, Laura Carnevali, Marco Paolieri, Fulvio Patara, Enrico Vicario |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2018 | A Model-Based Approach to Streamlining Distributed Training for Asynchronous SGDabstractThe success of Deep Neural Networks (DNNs) has created significant interest in the development of software tools, hardware architectures, and cloud systems to meet the huge computational demand of their training jobs. A common approach to speeding up an individual job is to distribute training data and computation among multiple nodes, periodically exchanging intermediate results. In this paper, we address two important problems for the application of this strategy to large-scale clusters and multiple, heterogeneous jobs. First, we propose and validate a queueing model to estimate the throughput of a training job as a function of the number of nodes assigned to the job; this model targets asynchronous Stochastic Gradient Descent (SGD), a popular strategy for distributed training, and requires only data from quick, two-node profiling in addition to job characteristics (number of requested training epochs, mini-batch size, size of DNN parameters, assigned bandwidth). Throughput estimations are then used to explore several classes of scheduling heuristics to reduce response time in a scenario where heterogeneous jobs are continuously submitted to a large-scale cluster. These scheduling algorithms dynamically select which jobs to run and how many nodes to assign to each job, based on different trade-offs between service time reduction and efficiency (e.g., speedup per additional node). Heuristics are evaluated through extensive simulations of realistic DNN workloads, also investigating the effects of early termination, a common scenario for DNN training jobs. Sung-Han Lin, Marco Paolieri, Cheng-Fu Chou, Leana Golubchik |
MASCOTS | 2 |
| 2017 | Performance Driven Resource Sharing Markets for the Small CloudabstractSmall-scale clouds (SCs) often suffer from resource under-provisioning during peak demand, leading to inability to satisfy service level agreements (SLAs) and consequent loss of customers. One approach to address this problem is for a set of autonomous SCs to share resources among themselves in a cost-induced cooperative fashion, thereby increasing their individual capacities (when needed) without having to significantly invest in more resources. In this context, a central problem is how to properly share resources for a price in order to achieve profitable service, while maintaining customer SLAs. To address this problem, we propose the SC-Share framework that utilizes two interacting models: (i) a stochastic performance model that estimates the achieved performance characteristics under given SLA requirements, and (ii) a market-based game-theoretic model that (as shown empirically) converges to efficient resource sharing decisions at market equilibrium. Our results include extensive evaluations that illustrate the utility of the proposed framework. Sung-Han Lin, Ranjan Pal, Marco Paolieri, Leana Golubchik |
ICDCS | 3 |
| 2017 | Guest Editorial: Special issue on formal modeling and analysis of timed systems
Marco Paolieri, Sriram Sankaranarayanan 0001, Enrico Vicario |
Real Time Syst. | 1 |
| 2016 | Performance Evaluation of Fischer's Protocol through Steady-State Analysis of Markov Regenerative ProcessesabstractFischer's protocol is a well-known timed mechanism through which a set of processes can synchronize access to a critical section without relying on atomic test-and-set operations, as might occur in a distributed environment or on a low-level computing platform. The protocol is based on a deterministic waiting time that can be defined so as to guarantee that possible interference due to concurrent accesses with random bounded delays be resolved with certainty. While protocol correctness descends from firm lower and upper bounds on waiting times and random delays, performance attained in synchronization also depends on continuous distributions of delays. Performance evaluation of a correct implementation thus requires the solution of a non-Markovian model whose underlying stochastic process falls in the class of Markov regenerative processes (MRPs) with multiple concurrent delays with non-exponential duration. Numerical solution of this class of models is to a large extent still an open problem. We provide a twofold contribution. We first introduce a novel method for the steady-state analysis of MRPs where regenerations are reached in a bounded number of discrete events, which enlarges the class amenable to numerical solution by allowing multiple concurrent timers with non-exponential distributions. The proposed technique is then applied to Fischer's protocol by characterizing the latency overhead due to synchronization, which comprises the first case where performance of the protocol is quantitatively assessed by jointly accounting for firm bounds and continuous distributions of delays. Stefano Martina, Marco Paolieri, Tommaso Papini, Enrico Vicario |
MASCOTS | 2 |
| 2016 | Probabilistic Model Checking of Regenerative Concurrent SystemsabstractWe consider the problem of verifying quantitative reachability properties in stochastic models of concurrent activities with generally distributed durations. Models are specified as stochastic time Petri nets and checked against Boolean combinations of interval until operators imposing bounds on the probability that the marking process will satisfy a goal condition at some time in the interval [α, β] after an execution that never violates a safety property. The proposed solution is based on the analysis of regeneration points in model executions: a regeneration is encountered after a discrete event if the future evolution depends only on the current marking and not on its previous history, thus satisfying the Markov property. We analyze systems in which multiple generally distributed timers can be started or stopped independently, but regeneration points are always encountered with probability 1 after a bounded number of discrete events. Leveraging the properties of regeneration points in probability spaces of execution paths, we show that the problem can be reduced to a set of Volterra integral equations, and we provide algorithms to compute their parameters through the enumeration of finite sequences of stochastic state classes encoding the joint probability density function (PDF) of generally distributed timers after each discrete event. The computation of symbolic PDFs is limited to discrete events before the first regeneration, and the repetitive structure of the stochastic process is exploited also before the lower bound α, providing crucial benefits for large time bounds. A case study is presented through the probabilistic formulation of Fischer's mutual exclusion protocol, a well-known real-time verification benchmark. Marco Paolieri, András Horváth, Enrico Vicario |
IEEE Trans. Software Eng. | 1 |
| 2014 | On How to Efficiently Implement Regular Expression Matching on FPGA-Based SystemsabstractThis work proposes a reconfigurable system able to perform - through a parallel and pipelined core, called ReCPU - regular expression matching. The system can configure on the programmable device, such as a FPGA, a set of ReCPUs, each one exploiting a single instance of the regular expression matching task on the given input string. These cores work in parallel on the same string analyzing different possible matching of the regular expression. Since the system is able to exploit dynamic partial reconfigurations, it can adapt at run-time the number of cores configured on the device, accordingly with the complexity of the regular expression. The adoption of the proposed solution makes it also possible to parallelize the regular expression matching process with a multiple cores architecture drastically reducing the time required for the completion of the task. Finally, run-time reconfiguration capabilities also allow to reduce the amount of resources required by the proposed approach. Vincenzo Rana, Francesco Bruschi, Marco Paolieri, Donatella Sciuto, Marco D. Santambrogio |
EUC | 3 |
| 2013 | Non-markovian analysis for model driven engineering of real-time softwareabstractQuantitative evaluation of models with stochastic timings can decisively support schedulability analysis and performance engineering of real-time concurrent systems. These tasks require modeling formalisms and solution techniques that can encompass stochastic temporal parameters firmly constrained within a bounded support, thus breaking the limits of Markovian approaches. The problem is further exacerbated by the need to represent suspension of timers, which results from common patterns of real-time programming. This poses relevant challenges both in the theoretical development of non-Markovian solution techniques and in their practical integration within a viable tailoring of industrial processes. Laura Carnevali, Marco Paolieri, Alessandro Santoni, Enrico Vicario |
ICPE | 2 |
| 2013 | A hard real-time capable multi-core SMT processorabstractHard real-time applications in safety critical domains require high performance and time analyzability. Multi-core processors are an answer to these demands, however task interferences make multi-cores more difficult to analyze from a worst-case execution time point of view than single-core processors. We propose a multi-core SMT processor that ensures a bounded maximum delay a task can suffer due to inter-task interferences. Multiple hard real-time tasks can be executed on different cores together with additional non real-time tasks. Our evaluation shows that the proposed MERASA multi-core provides predictability for hard real-time tasks and also high performance for non hard real-time tasks. Marco Paolieri, Jörg Mische, Stefan Metzlaff, Mike Gerdes 0001, Eduardo Quiñones, Sascha Uhrig, Theo Ungerer, Francisco J. Cazorla |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2013 | Timing effects of DDR memory systems in hard real-time multicore architectures: Issues and solutionsabstractMulticore processors are an effective solution to cope with the performance requirements of real-time embedded systems due to their good performance-per-watt ratio and high performance capabilities. Unfortunately, their use in integrated architectures such as IMA or AUTOSAR is limited by the fact that multicores do not guarantee a time composable behavior for the applications: the WCET of a task depends on inter-task interferences introduced by other tasks running simultaneously. This article focuses on the off-chip memory system: the hardware shared resource with the highest impact on the WCET and hence the main impediment for the use of multicores in integrated architectures. We present an analytical model that computes the worst-case delay, also known as Upper Bound Delay (UBD), that a memory request can suffer due to memory interferences generated by other co-running tasks. By considering the UBD in the WCET analysis, the resulting WCET estimation is independent from the other tasks, hence ensuring the time composability property and enabling the use of multicores in integrated architectures. We propose a memory controller for hard real-time multicores compliant with the analytical model that implements extra hardware features to deal with refresh operations and interferences generated by co-running non hard real-time tasks. Marco Paolieri, Eduardo Quiñones, Francisco J. Cazorla |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2012 | Transient analysis of non-Markovian models using stochastic state classes
András Horváth, Marco Paolieri, Lorenzo Ridi, Enrico Vicario |
Perform. Evaluation | 2 |
| 2011 | Towards functional-safe timing-dependable real-time architecturesabstractIn the near future the automotive systems will include microcontrollers hosting homogeneous or heterogeneous multi-core architectures, in which two or more CPU cores are combined to satisfy the high performance requirements. For those devices, time dependability issues represent a key challenge. In addition to that, they shall satisfy standards like ISO 26262 for functional safety and AUTOSAR for software architectures. This paper focuses on the study of problems and solutions related to functional-safe timing-dependable real-time architectures; in particular we identify critical failures related to timing issues and we propose a functional-safety aware methodology combining HW and SW measures to handle such kind of failures. Marco Paolieri, Riccardo Mariani |
IOLTS | 1 |
| 2011 | A Software-Pipelined Approach to Multicore Execution of Timing Predictable Multi-threaded Hard Real-Time TasksabstractMulticore processors can deliver higher performance than single-core processors by exploiting thread level parallelism (TLP): applications are split into independent threads, each of which is mapped into a different core, reducing the execution time and potentially its worst-case execution time (WCET). Unfortunately, inter-thread interferences generated by simultaneous accesses to shared resources from different threads may completely destroy the performance benefits brought by TLP. This paper proposes a software/hardware cache partitioning approach that reduces the inter-thread memory interferences generated in hard real-time software-pipelined parallel applications. Our results show that our approach effectively reduces memory interferences, while still guaranteeing a predictable timing behaviour, achieving a WCET estimation reduction of 28% for a software pipelined version of the LU decomposition application with respect to the single-threaded version. Marco Paolieri, Eduardo Quiñones, Francisco J. Cazorla, Julian Wolf 0002, Theo Ungerer, Sascha Uhrig, Zlatko Petrov |
ISORC | 1 |
| 2011 | IA^3: An Interference Aware Allocation Algorithm for Multicore Hard Real-Time SystemsabstractIn multicore processors, the execution environment is defined as the environment in which tasks run and it is determined by the hardware resources they get and the workload with which they are executed. Thus, different execution environments lead to different inter-task interferences accessing shared hardware resources due to conflicts with the other corunning tasks, making the WCET estimation of a task dependent on the execution environment in which it runs. Despite such dependency, current partitioned scheduling approaches use a single WCET estimation per task: typically the highest for all execution environments in which a task runs. In this paper we introduce IA3: an interference-aware allocation algorithm that considers not a single WCET estimation but a set of WCET estimations per task. IA3 is based on two novel concepts: the WCET-matrix and the WCET-sensitivity. The former associates every WCET estimation with its corresponding execution environment. The latter measures the impact of changing the execution environment on the WCET estimation. This allows IA3 to reduce the number of resources required to schedule a given taskset. In particular, our results show that in a four-core processor considering tasksets with a total utilization of 2.9, IA3 is able to schedule 70% of the tasksets using 3-cores while a classical partitioned approach with a First-Fit Decreasing heuristic is able to schedule only 5% of the tasksets using 3-cores. Marco Paolieri, Eduardo Quiñones, Francisco J. Cazorla, Robert I. Davis 0001, Mateo Valero |
IEEE Real-Time and Embedded Technology and Applications Symposium | 1 |
| 2010 | A Reconfigurable System Based on a Parallel and Pipelined Solution for Regular Expression MatchingabstractWe present a reconfigurable architecture that can perform highly parallel regular expression matching. The system can be configured on programmable devices such as FPGAs as a set of instances of a predefined core called REMA. Each core addresses one of the subtasks into which the regular expression matching problem can be partitioned. These cores work in parallel on the same string analyzing different possible matchings. Since the system can exploit dynamic partial reconfigurations, it can adapt the number of cores configured on the device at run-time, according to the complexity of the regular expression, which is proportional to the number of different ways in which the reference string can be mapped on pattern. Making it possible to parallelize the matching process with a multiple core architecture drastically improves temporal efficiency (up to one order of magnitude with respect to software solutions and up to a speedup factor of 25 with respect to hardware solutions). Most important, the run-time reconfigurability feature allows to implement a just in time logic usage strategy, thus reducing the mean amount of resources required. Francesco Bruschi, Marco Paolieri, Vincenzo Rana |
FPL | 2 |
| 2009 | Hardware support for WCET analysis of hard real-time multicore systemsabstractThe increasing demand for new functionalities in current and future hard real-time embedded systems like automotive, avionics and space industries is driving an increase in the performance required in embedded processors. Multicore processors represent a good design solution for such systems due to their high performance, low cost and power consumption characteristics. However, hard real-time embedded systems require time analyzability and current multicore processors are less analyzable than single-core processors due to the interferences between different tasks when accessing shared hardware resources. In this paper we propose a multicore architecture with shared resources that allows the execution of applications with hard real-time and non hard real-time constraints at the same time, providing time analizability for the hard real-time tasks so that they can meet their deadlines. Moreover our architecture proposal provides high-performance for the non hard real-time tasks. Marco Paolieri, Eduardo Quiñones, Francisco J. Cazorla, Guillem Bernat, Mateo Valero |
ISCA | 1 |
| 2008 | An adaptable FPGA-based System for Regular Expression MatchingabstractIn many applications string pattern matching is one of the most intensive tasks in terms of computation time and memory accesses. Network Intrusion Detection Systems and DNA Sequence Matching are two examples. Since software solutions are not able to satisfy the performance requirements, specialized hardware architectures are required. In this paper we propose a complete framework for regular expression matching, both in its architecture and compiler. This special-purpose processor is programmed using regular expressions as programming language. With the parallelism exploited in the design it is possible to achieve a throughput greater than one character per clock cycle, requiring O(n) memory space. The VHDL description of the proposed architecture is fully configurable. A design space exploration to find the optimal architecture based on area and performance cost-function is presented. Ivano Bonesana, Marco Paolieri, Marco D. Santambrogio |
DATE | 2 |
| 2007 | Power Modeling and Power Analysis for IEEE 802.15.4: a Concurrent State Machine Approachabstractis a recent low-rate/low-power standard for wireless personal area and sensor networks. Its simple infrastructure, intermediate range and good power performance make it a candidate for applications that require a reasonably low throughput but a very high device lifetime and power efficiency. An experimental power analysis of an 802.15.4 implementation is carried out, providing a detailed power model of the protocol based on concurrent state machines; resulting power model is then used to generate a customized simulator. The model has been validated through a set of experiments and provides good accuracy; results are discussed, considering in particular use of the model as a basis for subsequent optimizations on 802.15.4 networks. Marcello Mura, Marco Paolieri, Fabio Fabbri, Luca Negri, Mariagiovanna Sami |
CCNC | 2 |
| 2007 | SC2 StateCharts to SystemC: Automatic Executable Models Generation
Marcello Mura, Marco Paolieri |
FDL | 2 |
| 2007 | StateCharts to systemc: a high level hardware simulation approachabstractIn this paper we present a tool that converts specifications written with a subset of StateCharts into SystemC behavioral models. The main advantages of such an approachare rapidity of use, simplicity and reusability. Various systems can be modeled at different levels of abstraction andaccuracy through StateCharts and different peculiar aspects (e.g. energy, performances) can be taken into consideration. Moreover different parts of the design can be identified at different detail levels. The kernel of the simulator is fully discussed together with its mapping to the semantics of our StateCharts diagrams. As a case study we present here a model of the IBM PowerPC 750 Cache system and the respective SystemC simulator automatically generated by our tool. Marcello Mura, Marco Paolieri, Luca Negri, Mariagiovanna Sami |
ACM Great Lakes Symposium on VLSI | 2 |
| 2007 | ReCPU: A parallel and pipelined architecture for regular expression matchingabstractText pattern matching is one of the main and most computation intensive parts of systems such as Network Intrusion Detection Systems and DNA Sequencing Matching. Software solutions to this are available but often they do not satisfy the requirements in terms of performance. This paper presents a new hardware approach for regular expression matching: ReCPU. The proposed solution is a parallel and pipelined architecture able to deal with the common regular expression semantics. This implementation based on several parallel units achieves a throughput of more than one character per clock cycle (maximum performance of state of the art solutions) requiring just O(n) memory locations (where n is the length of the regular expression). Performance has been evaluated synthesizing the VHDL description. Area and time constraints have been analyzed. Experimental results are obtained simulating the architecture. Marco Paolieri, Ivano Bonesana, Marco D. Santambrogio |
VLSI-SoC | 1 |