EDBT 2026 Demo / reviewers in the wild / expert
Cosimo Antonio Prete
dblp:76/1399
· DBLP profile ↗
29ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-8467-8198ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 2 first-authorSoftware engineering, systems software and programming languages · 3Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
1 paper |
Geometric modeling and processing · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Memory systems · 39% Parallel and multicore computing · 19% Performance modeling and evaluation · 18% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Geometric modeling and processing › computer-aided design › computer-aided geometric design
NURBS interpolation |
0.1 | 1 | 2012 | A real-time configurable NURBS interpolator with bounded acceleration, jerk and chord error · Comput. Aided Des. 2012 |
Geometric modeling and processing
real-time interpolation |
0.1 | 1 | 2012 | A real-time configurable NURBS interpolator with bounded acceleration, jerk and chord error · Comput. Aided Des. 2012 |
Memory systems
cache coherence |
0.0 | 2 | 1999 | PSCR: A Coherence Protocol for Eliminating Passive Sharing in Shared-Bus Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 1999 A Trace-Driven Simulator for Performance Evaluation of Cache-Based Multiprocessor Systems · IEEE Trans. Parallel Distributed Syst. 1995 |
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor |
0.0 | 2 | 1999 | PSCR: A Coherence Protocol for Eliminating Passive Sharing in Shared-Bus Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 1999 A Trace-Driven Simulator for Performance Evaluation of Cache-Based Multiprocessor Systems · IEEE Trans. Parallel Distributed Syst. 1995 |
Processor architecture and microarchitecture › multiprocessor architecture
bus-based multiprocessor |
0.0 | 1 | 1999 | PSCR: A Coherence Protocol for Eliminating Passive Sharing in Shared-Bus Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 1999 |
Memory systems › cache coherence
cache coherence protocol |
0.0 | 1 | 1999 | PSCR: A Coherence Protocol for Eliminating Passive Sharing in Shared-Bus Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 1999 |
Performance modeling and evaluation
simulation |
0.0 | 1 | 1995 | A Trace-Driven Simulator for Performance Evaluation of Cache-Based Multiprocessor Systems · IEEE Trans. Parallel Distributed Syst. 1995 |
Performance modeling and evaluation › simulation › discrete-event simulation
trace-driven simulation |
0.0 | 1 | 1995 | A Trace-Driven Simulator for Performance Evaluation of Cache-Based Multiprocessor Systems · IEEE Trans. Parallel Distributed Syst. 1995 |
Operating systems › resource management › process management
process migration |
0.0 | 1 | 1999 | PSCR: A Coherence Protocol for Eliminating Passive Sharing in Shared-Bus Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 1999 |
Distributed systems › distributed database
commit protocol |
0.0 | 1 | 1990 | A Distributed Commit Protocol for a Multicomputer System · IEEE Trans. Computers 1990 |
Distributed systems › distributed database
distributed transactions |
0.0 | 1 | 1990 | A Distributed Commit Protocol for a Multicomputer System · IEEE Trans. Computers 1990 |
Parallel and multicore computing
multicomputer |
0.0 | 1 | 1990 | A Distributed Commit Protocol for a Multicomputer System · IEEE Trans. Computers 1990 |
Methods — techniques the papers use, named apart from their topics
trace-driven simulation · 0.1compiler-provided information · 0.0synthetic trace generation · 0.0distributed protocol design · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Explainable ensemble learning for structural damage prediction under seismic eventsabstractThis paper presents an explainable ensemble learning framework using Bootstrap Aggregating to predict structural damage in masonry buildings during seismic events.It estimates the peak ground acceleration (PGA) leading to the damage control limit state (significant damage) based on structural parameters.The model achieves high accuracy (R 2 =0.9536,MAE=0.0057) and interpretability through SHAP, aligning with structural engineering principles.Compared to finite element analyses, it offers faster computations (milliseconds) and scalability, enabling rapid intervention planning after earthquakes.Developed under the MEDEA project (EU Grant n. 10101236), it supports disaster response and enhances seismic resilience. Michele Baldassini, Pierfrancesco Foglia, Beatrice Lazzerini, Francesco Pistolesi, Cosimo Antonio Prete |
ESANN | 5 |
| 2022 | in-Car Entertainment via Group-wise Temporary Mobile Social NetworkingabstractNext generation cars will increase the passengers’ time for fun and relax, as well as the number of unknown passengers traveling together. A key functionality to improve the users’ experience is that of Temporary Mobile Social Networking (TMSN): where passengers form, for a limited-time, a mobile social group with common interests and activities, using their already available social network accounts. The goal of TMSN is to automatically redesign the users’ profiles and interfaces into a group-wise passengers’ profile and a common interface, by reducing isolation and enabling socialization. In this paper, a TMSN-inspired music selection is proposed and developed via the Spotify music streaming service. Early results are promising and encourage further developments towards the concept of in-car entertainment. Mario G. C. A. Cimino, Antonio Di Tecco, Pierfrancesco Foglia, Raffaele Giannessi, Jacopo Malvatani, Cosimo Antonio Prete, Giulio Rossolini |
VEHITS | 6 |
| 2021 | A pre-processing technique to decrease inspection time in glass tube production linesabstractAbstract In case of glass tube for pharmaceutical applications, high‐quality defect detection is made via inspection systems based on computer vision. The processing must guarantee real‐time inspection and meet increasing rate and quality requirements. Defect detection in glass tubes is complicated by aspects that hamper the efficiency of state‐of‐the‐art techniques. This paper presents a pre‐processing algorithm which excludes portions of the image where defects are surely absent. The approach decreases the time for defect detection and classification phases (any detection algorithm can be applied), as they are applied only in high‐probability candidate sub‐image. We derive a methodology to get robust values of algorithm's parameters during production. The algorithm relies on detrended standard deviation and double threshold hysteresis, which solve issues related to the misalignment between illuminator and acquisition camera, and enable a robust detection despite rotation, vibration, and irregularities of tubes. We consider Canny, MAGDDA, and Niblack algorithms. The solution keeps the detection quality of such algorithms and reaches a 4.69× throughput gain. It represents a methodology to obtain defect detection in time‐constrained environments through a software‐only approach, and can be exploited in parallel/accelerated solutions and in contexts where a linear camera is utilized on both flat and uneven surfaces. Gabriele Antonio De Vitis, Pierfrancesco Foglia, Cosimo Antonio Prete |
IET Image Process. | 3 |
| 2020 | Row-level algorithm to improve real-time performance of glass tube defect detection in the production phaseabstractIn the case of the glass tube for pharmaceutical applications, high‐quality defect detection is made via inspection systems based on image processing. Such processing must be fast enough to guarantee real‐time inspection and to meet the increasing rate and quality required by the market. Defect detection is complex due to specific problems of the production process: vibration, rotation and irregularity of the tube. All these aspects prevent the efficient use of known techniques. The authors present an algorithm that decreases the processing time of the defect detection phase. The algorithm is based on a moving average filter working at row level, that allows to minimize the effects of rotation, vibration, and irregularity of the tube. Luminosity variations due to the tube curvature are cut by the filter and a threshold algorithm can be applied. They made the evaluation considering different solutions taken from literature. The algorithm outperforms, in processing time, all these solutions with increased accuracy. Experimental measures show that the algorithm achieves a throughput gain of 2.6 times with respect to Canny. They develop also a methodology to get the best values for the algorithm parameters directly at the factory, during the change of production batches. Gabriele Antonio De Vitis, Pierfrancesco Foglia, Cosimo Antonio Prete |
IET Image Process. | 3 |
| 2018 | Exploring the relationship between architectures and management policies in the design of NUCA-based chip multicore systems
Sandro Bartolini, Pierfrancesco Foglia, Cosimo Antonio Prete |
Future Gener. Comput. Syst. | 3 |
| 2014 | Social and Q&A interfaces for app download
Gianluca Dini, Pierfrancesco Foglia, Cosimo Antonio Prete, Michele Zanda |
Inf. Process. Manag. | 3 |
| 2014 | Evaluation of Leakage Reduction Alternatives for Deep Submicron Dynamic Nonuniform Cache Architecture CachesabstractWire delays and leakage energy consumption are both growing problems in designing large on-chip caches. Nonuniform cache architecture (NUCA) is a wire-delay aware design paradigm based on the sub-banking of a cache, which allows the banks closer to the controller to be accessed with reduced latencies with respect to the other banks. This feature is leveraged by dynamic NUCA (D-NUCA) caches via a migration mechanism which speeds up frequently used data access, further reducing the effect wire delays have on performance. To reduce leakage power consumption of static random access memory caches, various micro-architectural techniques have been proposed. In this brief, we compare the benefits and limits of the application of some of these techniques to a D-NUCA cache memory, and propose a novel hybrid scheme based on the Drowsy and Way Adaptable techniques. Such a scheme allows further improvement in leakage reduction and limits the impact of process variation on the effectiveness of the Drowsy technique. Alessandro Bardine, Manuel Comparetti, Pierfrancesco Foglia, Cosimo Antonio Prete |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2012 | Integration of existing IEC 61131-3 systems in an IEC 61499 distributed solutionabstractThe IEC 61499 standard allows to model and design new generation control systems, providing innovative concepts of software engineering (such as abstraction, encapsulation, reuse) to the world of control engineering. The industrial reception of the standard, however, is still in an early stage, also because its introduction results in the adoption of a programming paradigm profoundly different than the widespread IEC 61131-3. This paper presents a method for the integration of the two standards, that allows to exploit the benefits of both. The proposed architecture is based on the parallel execution of both environments that interact with each other through some specific interfaces. A test implementation of the architecture is also presented to demonstrate the feasibility of the proposed solution. Stefano Campanelli, Pierfrancesco Foglia, Cosimo Antonio Prete |
ETFA | 3 |
| 2012 | A real-time configurable NURBS interpolator with bounded acceleration, jerk and chord error
Massimiliano Annoni, Alessandro Bardine, Stefano Campanelli, Pierfrancesco Foglia, Cosimo Antonio Prete |
Comput. Aided Des. | 5 |
| 2010 | Re-NUCA: Boosting CMP Performance Through Block ReplicationabstractChip Multiprocessor (CMP) systems have become the reference architecture for designing micro-processors, thanks to the improvements in semiconductor nanotechnology that have continuously provided a crescent number of faster and smaller per-chip transistors. The interests for CMPs grew up since classical techniques for boosting performance, e.g. the increase of clock frequency and the amount of work performed at each clock cycle, can no longer deliver to significant improvement due to energy constrains and wire delay effects. CMP systems generally adopt a large last-level-cache (LLC) (typically, L2 or L3) shared among all cores, and private L1 caches. As the miss resolution time for private caches depends on the response time of the LLC, which is wire-delay dominated, performance are affected by wire delay. NUCA caches have been proposed for single and multi core systems as a mechanism for tolerating wire-delay effects on the overall performance. In this paper, we introduce a novel NUCA architecture, called Re-NUCA, specifically suited for (but not limited to) CMPs in which cores are placed at different sides of the shared cache. The idea is to allow shared blocks to be replicated inside the shared cache, in order to avoid the limitations to performance improvements that arise in classical D-NUCA caches due to the conflict hit problem. Our results show that Re-NUCA outperforms D-NUCA of more then 5% on average, but for those applications that strongly suffer from the conflict hit problem we observe performance improvements up to 15%. Pierfrancesco Foglia, Cosimo Antonio Prete, Marco Solinas, Giovanna Monni |
DSD | 2 |
| 2010 | Feedback-Driven Restructuring of Multi-threaded Applications for NUCA Cache Performance in CMPsabstractThis paper addresses feedback-directed restructuring techniques tuned to Non Uniform Cache Architectures (NUCA) in CMPs running multi-threaded applications. Access time to NUCA caches depends on the location of the referred block, so the locality and cache mapping of the application influence the overall performance. We show techniques for altering the distribution of applications into the cache space as to achieve improved average memory access time. In CMPs running multi-threaded applications, the aggregated accesses (and locality) of the processors form the actual cache load and pose specific issues. We consider a number of Splash-2 and Parsec benchmarks on an 8 processor system and we show that a relatively simple remapping algorithm is able to improve the average Static-NUCA (SNUCA) cache access time by 5.5% and allows an SNUCA cache to surpass the performance of a more complex dynamic-NUCA (DNUCA) for most benchmarks. Then, we present a more sophisticated remapping algorithm, relying on cache geometry information and on the access distribution statistics from individual processors, that reduces the average cache access time by 10.2% and is very stable across all benchmarks. Sandro Bartolini, Pierfrancesco Foglia, Marco Solinas, Cosimo Antonio Prete |
SBAC-PAD | 4 |
| 2009 | A power-efficient migration mechanism for D-NUCA cachesabstractD-NUCA L2 caches are able to tolerate the increasing wire delay effects due to technology scaling thanks to their banked organization, broadcast line search and data promotion/demotion mechanism. Data promotion mechanism aims at moving frequently accessed data near the core, but causes additional accesses on cache banks, hence increasing dynamic energy consumption. We shown how, in some cases, this migration mechanism is not successful in reducing data access latency and can be selectively and dynamically inhibited, thus reducing dynamic energy consumption without affecting performances. Alessandro Bardine, Manuel Comparetti, Pierfrancesco Foglia, Giacomo Gabrielli, Cosimo Antonio Prete |
DATE | 5 |
| 2009 | An Evaluation of Behaviors of S-NUCA CMPs Running Scientific WorkloadabstractModern systems are able to put two or more processors on the same die (Chip Multiprocessors, CMP), each with its private caches, while the last level caches can be either private or shared. As these systems are affected by the wire delay problem, NUCA caches have been proposed to hide the effects of such delay in order to increase performance. A CMP system that adopt a NUCA as its shared last level cache has to be able to maintain coherence among the lowest, private levels of the cache hierarchy. As NUCA caches typically adopt a NoC as the communication infrastructure (in which the communication paradigm is message-passing), the coherence protocol has to be directory based, similar to the ones proposed for classical DSM systems. Previous works focusing on NUCA-based CMP systems adopt a fixed topology (i.e. physical position of cores and NUCA banks, and the communication infrastructure) each adopting different coherence strategies. In this paper, we present an evaluation of an 8-cpu CMP system with two levels of cache, in which the Lis are private of each core, while the L2 is a Static-NUCA shared among all cores. We considered two different system topologies (the first with the eight cpus connected to the NUCA at the same side, the second with half of the cpus on one side and the others at the opposite side), and for all the topologies we considered MES1 and MOES1. The results indicate that processor topology has much more effect on performance and NOC bandwidth utilization than the coherence protocol, as a consequence of data mapping and accesses' distribution to the L2 cache that is not uniformly distributed to all the cache banks. Pierfrancesco Foglia, Francesco Panicucci, Cosimo Antonio Prete, Marco Solinas |
DSD | 3 |
| 2009 | Analysis of Performance Dependencies in NUCA-Based CMP SystemsabstractImprovements in semiconductor nanotechnology have continuously provided a crescent number of faster and smaller per-chip transistors. Consequent classical techniques for boosting performance, such as the increase of clock frequency and the amount of work performed at each clock cycle, can no longer deliver to significant improvement due to energy constrains and wire delay effects. As a consequence, designers interests have shifted toward the implementation of systems with multiple cores per chip (Chip Multiprocessors, CMP). CMP systems typically adopt a large last-level-cache (LLC) shared among all cores, and private L1 caches. As the miss resolution time for private caches depends on the response time of the LLC, which is wire-delay dominated, performance are affected by wire delay. NUCA caches have been proposed for single and multi core systems as a mechanism for such tolerating wire-delay effects on the overall performance. In this paper, we introduce our design for S-NUCA and D-NUCA cache memory systems, and we present an analysis of an 8-cpu CMP system with two levels of cache, in which the L1s are private, while the L2 is a NUCA shared among all cores. We considered two different system topologies (the first with the eight cpus connected to the NUCA at the same side -8p-, the second with half of the cpus on one side and the others at the opposite side -4+4p), and for all the configurations we evaluate the effectiveness of both the static and dynamic policies that have been proposed. Our results show that adopting a D-NUCA scheme with the 8p configuration is the best performing solution among all the considered configurations, and that for the 4+4p configuration the D-NUCA outperforms the S-NUCA in most of the cases. We highlight that performance are tied to both mapping strategy variations (Static and Dynamic) and topology changes. We also observe that bandwidth occupancy depends on both the NUCA policy and topology. Pierfrancesco Foglia, Francesco Panicucci, Cosimo Antonio Prete, Marco Solinas |
SBAC-PAD | 3 |
| 2008 | Leveraging Data Promotion for Low Power D-NUCA CachesabstractD-NUCA caches are cache memories that, thanks to banked organization, broadcast search and promotion/demotion mechanism, are able to tolerate the increasing wire delay effects introduced by technology scaling. As a consequence, they will outperform conventional caches (UCA, Uniform CacheArchitectures) in future generation cores. Due to the promotion/ demotion mechanism, we observed that the distribution of hits across the ways of a D-NUCA cache varies across applications as well as across different execution phases within a single application. In this work, we show how such a behavior can be leveraged to improve the D-NUCA power efficiency as well as to decrease its access latency.In particular, we propose: 1) A new micro-architectural technique to reduce the static power consumption of a D-NUCA cache by dynamically adapting the number of active (i.e. powered-on) ways to the need of the running application; our evaluation shows that a strong reduction of the average number of active ways (37.1%) is achievable, without significantly affecting the IPC (-2.83%), leading to a resultant reduction of the Energy Delay Product (EDP) of 30.9%. 2) A strategy to estimate the characteristic parameters of the proposed technique. 3) An evaluation of the effectiveness of the proposed technique in the multicore environment. Alessandro Bardine, Manuel Comparetti, Pierfrancesco Foglia, Giacomo Gabrielli, Cosimo Antonio Prete, Per Stenström |
DSD | 5 |
| 2008 | Performance Sensitivity of NUCA Caches to On-Chip Network ParametersabstractNon uniform cache architectures (NUCA) are a novel design paradigm for large last-level on-chip caches that has been introduced to deliver low access latencies in wire delay dominated environments. Their structure is partitioned into sub-banks and the resulting access latency is a function of the physical position of the requested data. Typically, to connect the different sub-banks and the cache controller, NUCA caches employ a switched network, made up of links and routers with buffered queues; the characteristics of such switched network may affect the performance of the entire system. This work analyzes how different parameters for the routers, namely cut-through latency and buffering capacity, affect the overall performance of NUCA based systems for the single processor case, assuming a reference organization proposed in literature. The results indicate that the sensitivity of the system to the cut-through latency is very high and that limited buffering capacity is sufficient to achieve a good performance level. As a consequence, we propose an alternative NUCA organization that limits the average number of hops experienced by cache accesses. This organization is better performing in most of the cases and scales better as the cut-through latency increases, thus simplifying the implementation of routers. Alessandro Bardine, Manuel Comparetti, Pierfrancesco Foglia, Giacomo Gabrielli, Cosimo Antonio Prete |
SBAC-PAD | 5 |
| 2005 | An Innovative Tool to Easily Get Usable Web Sites
Cosimo Antonio Prete, Pierfrancesco Foglia, Michele Zanda |
WEBIST | 1 |
| 2005 | Reducing coherence overhead and boosting performance of high-end SMP multiprocessors running a DSS workloadabstractIn this work, we characterized the memory performance—and in particular the impact of coherence overhead and process migration—of a shared-bus shared-memory multiprocessor running a DSS workload. When the number of processors is increased in order to achieve higher computational power, the bus becomes a major bottleneck of such architecture. We evaluated solutions that can greatly reduce that bottleneck. An area where this kind of optimization is important regards data base systems. For this reason, we considered a DSS workload and we setup the experiments following TPC-D specifications on the PostgreSQL DBMS in order to explore different optimizations on same kind of workloads as evaluated in the literature. In this scenario, we compare possible solutions to boost performance and we show the impact of process migration on coherence overhead. We found that the consequences of coherence overhead and process migration on performance are very important in machines with 16 or more processors. In this case, even little sharing, as in DSS applications, can become crucial for system performance. Another important result of our analysis regards the interaction between the coherence protocol and the scheduler. The basic cache affinity scheduling is useful in reducing migration, but it is not effective in every load condition. Specific coherence protocols can help reduce the effects of process migration, especially in situations when the scheduler cannot apply the affinity requirement. In these conditions, the use of a write-update protocol with a selective invalidation strategy for private data improves performance (and scalability) of about 20% with respect to a classical MESI-based solution. This advantage is about 50% in the case of high cache-to-cache transfer. Pierfrancesco Foglia, Roberto Giorgi, Cosimo Antonio Prete |
J. Parallel Distributed Comput. | 3 |
| 2005 | Optimizing instruction cache performance of embedded systemsabstractIn the embedded domain, the gap between memory and processor performance and the increase in application complexity need to be supported without wasting precious system resources: die size, power, etc. For these reasons, effective exploitation of small and simple cache memories is of the utmost importance. However, programs running on such caches can experience serious inefficiencies due to cache conflicts.We present a new Cache-Aware Code Allocation Technique (CAT), which transforms the structure of programs so that their behavior toward memory can meet the locality features the cache is able to exploit. The proposed approach uses detailed information of program execution to place program areas into memory and employs the new idea of “look-forward estimation” that helps to seek better global layouts during the placement of each area. CAT-optimized programs outperform the original ones achieving the same miss rate on two times, and sometimes four times, smaller caches. Moreover, CAT improves the instruction miss rate by more than 40% if compared to the best procedure-reordering algorithm. CAT performances derive from the increased number of cache lines that support the execution of optimized applications and from a more balanced load on them. Sandro Bartolini, Cosimo Antonio Prete |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2002 | A cache-aware program transformation technique suitable for embedded systems
Sandro Bartolini, Cosimo Antonio Prete |
Inf. Softw. Technol. | 2 |
| 2002 | Performance-steered design of software architectures for embedded multicore systemsabstractAbstract Many software applications demanding a considerable computing power are moving towards the field of embedded systems (and, in particular, hand‐held devices). A possible way to increase the computing power of this kind of platform, so that both cost and power consumption are kept low, is the employment of multiple CPU cores on the same chipset. Consequently, it is essential to design applications that meet performance requirements leveraging the underlying parallel platform. As embedded applications are usually built using different components (whose source code is often not available) from different companies, the designer can mostly only operate at the architectural level. So far, methodologies for designing software architectures have mainly addressed general‐purpose systems, often relying on hardware platforms with a high degree of parallelism. In this paper, we present our experience in architectural design of parallel embedded applications; as a result, we propose a possible methodology for the application design at the architectural level, targeted to embedded systems built upon multicore chipsets with a low degree of parallelism. It makes use of performance predictions, obtained by simulations. Such a methodology can be employed both for retargeting existing sequential applications to parallel processing platforms and for designing complete applications from scratch. We show the application of the proposed methodology to an embedded digital cartographic system. Starting with a software description using UML diagrams, candidate software architectures (utilizing different parallel solutions) are first defined and then evaluated, to end with the selection of the one yielding the highest performance gain. Copyright © 2002 John Wiley & Sons, Ltd. Alessio Bechini, Cosimo Antonio Prete |
Softw. Pract. Exp. | 2 |
| 2001 | Behavior investigation of concurrent Java programs: an approach based on source-code instrumentation
Alessio Bechini, Cosimo Antonio Prete |
Future Gener. Comput. Syst. | 2 |
| 1999 | Process Migration Effects on Memory Performance of Multiprocessor
Pierfrancesco Foglia, Roberto Giorgi, Cosimo Antonio Prete |
HiPC | 3 |
| 1999 | PSCR: A Coherence Protocol for Eliminating Passive Sharing in Shared-Bus Shared-Memory MultiprocessorsabstractIn high-performance general-purpose workstations and servers, the workload can be typically constituted of both sequential and parallel applications. Shared-bus shared-memory multiprocessor can be used to speed-up the execution of such workload. In this environment, the scheduler takes care of the load balancing by allocating a ready process on the first available processor, thus producing process migration. Process migration and the persistence of private data into different caches produce an undesired sharing, named passive sharing. The copies due to passive sharing produce useless coherence traffic on the bus and coping with such a problem may represent a challenging design problem for these machines. Many protocols use smart solutions to limit the overhead to maintain coherence among shared copies. None of these studies treats passive-sharing directly, although some indirect effect is present while dealing with the other kinds of sharing. Affinity scheduling can alleviate this problem, but this technique does not adapt to all load conditions, especially when the effects of migration are massive. We present a simple coherence protocol that eliminates passive sharing using information from the compiler that is normally available in operating system kernels. We evaluate the performance of this protocol and compare it against other solutions proposed in the literature by means of enhanced trace-driven simulation. We evaluate the complexity in terms of the number of protocol states, additional bus lines, and required software support. Our protocol further limits the coherence-maintaining overhead by using information about access patterns to shared data exhibited in parallel applications. Roberto Giorgi, Cosimo Antonio Prete |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1995 | A Trace-Driven Simulator for Performance Evaluation of Cache-Based Multiprocessor SystemsabstractWe describe a simulator which emulates the activity of a shared memory, common bus multiprocessor system with private caches. Both kernel and user program activities are considered, thus allowing an accurate analysis and evaluation of coherence protocol performance. The simulator can generate synthetic traces, based on a wide set of input parameters which specify processor, kernel and workload features. Other parameters allow us to detail the multiprocessor architecture for which the analysis has to be carried out. An actual-trace-driven simulation is possible, too, in order to evaluate the performance of a specific multiprocessor with respect to a given workload, if traces concerning this workload are available. In a separate section, we describe how actual traces can also be used to extract a set of input parameters for synthetic trace generation. Finally, we show how the simulator may be successfully employed to carry out a detailed performance analysis of a specific coherence protocol.> Cosimo Antonio Prete, Gianpaolo Prina, Luigi M. Ricciardi |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1990 | A new solution of coherence protocol for tightly coupled multiprocessor systems
Cosimo Antonio Prete |
Microprocessing and Microprogramming | 1 |
| 1990 | A Distributed Commit Protocol for a Multicomputer SystemabstractA distributed commit protocol suitable for multicomputer systems is described. A general programming environment which provides transactions as a programming tool is discussed. This environment is expected to be more dynamic than a database management system; in particular, it is not known how many and which processes will participate in a specific transaction. Therefore, a model of completely distributed transactions, without any hierarchical structure among the participant processes or any centralized locus of control, is proposed.> Paolo Ancilotti, Beatrice Lazzerini, Cosimo Antonio Prete, Maurizio Sacchi |
IEEE Trans. Computers | 3 |
| 1989 | A protocol for resource locking and deadlock detection in a multi-user environment
Andrea Domenici, Beatrice Lazzerini, Cosimo Antonio Prete |
Microprocessing and Microprogramming | 3 |
| 1986 | DISDEB: An interactive high-level debugging system for a multi-microprocessor system
Beatrice Lazzerini, Cosimo Antonio Prete |
Microprocessing and Microprogramming | 2 |