EDBT 2026 Demo / reviewers in the wild / expert
Mats Brorsson
dblp:30/2217
· DBLP profile ↗
33ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-9637-2065ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 4Software engineering, systems software and programming languages · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Performance and Usability Implications of Multiplatform and WebAssembly Containersabstractpeer reviewed Sangeeta Kakati, Mats Brorsson |
CLOSER | 2 |
| 2025 | The European master for HPC curriculumabstractInternational audience Pascal Bouvry, Mats Brorsson, Ramon Canal, Aryan Eftekhari, Siegfried Höfinger, Didier Smets, Harald Köstler, Tomás Kozubek, Ezhilmathi Krishnasamy, Josep Llosa, Alexandra Lukas-Rother, Xavier Martorell, Dirk Pleiter, Ana Proykova, Maria-Ribera Sancho, Olaf Schenk, Cristina Silvano |
J. Parallel Distributed Comput. | 2 |
| 2024 | A Cross-Architecture Evaluation of WebAssembly in the Cloud-Edge ContinuumabstractAs cloud-to-edge computing becomes increasingly prevalent, the need for an application framework capable of dynamically utilizing the entire spectrum of resources has grown. As developers increasingly seek to deploy applications across heterogeneous computing environments, this research aims to introduce how WebAssembly(Wasm) emerges as a versatile ally, seamlessly executing on different architectures, allowing developers to craft applications without being bothered by the underlying hardware platform.WebAssembly binaries need a runtime to execute. In this paper, we present a thorough performance analysis of the two most prominent WebAssembly runtimes employing an extensive array of instrumented benchmarks to ensure precise and reliable results. We focus on investigating WebAssembly’s performance characteristics while considering important metrics like execution speed and startup time. Unprecedentedly, we extended the evaluation to four diverse sets of architectures: two server-class architectures (X86_64, ARM64) and two embedded boards (Nvidia Jetson Nano with ARM64, StarFive VisionFive2 with RISCV64), marking the first-ever cross-architecture analysis of WebAssembly runtimes. This novel evaluation empowers us to offer valuable insights into the performance traits and considerations of WebAssembly. By scrutinizing architecture-specific results, we shed light on Wasm’s potential to address the requirements of a cross-architecture cloud-to-edge application framework and reshape the landscape of modern application frameworks. Sangeeta Kakati, Mats Brorsson |
CCGrid | 2 |
| 2024 | Cost-aware Service Placement and Scheduling in the Edge-Cloud ContinuumabstractThe edge to data center computing continuum is the aggregation of computing resources located anywhere between the network edge (e.g., close to 5G antennas), and servers in traditional data centers. Kubernetes is the de facto standard for the orchestration of services in data center environments, where it is very efficient. It, however, fails to give the same performance when including edge resources. At the edge, resources are more limited, and networking conditions are changing over time. In this article, we present a methodology that lowers the costs of running applications in the edge-to-cloud computing continuum. This methodology can adapt to changing environments, e.g., moving end-users. We are also monitoring some Key Performance Indicators of the applications to ensure that cost optimizations do not negatively impact their Quality of Service. In addition, to ensure that performances are optimal even when users are moving, we introduce a background process that periodically checks if a better location is available for the service and, if so, moves the service. To demonstrate the performance of our scheduling approach, we evaluate it using a vehicle cooperative perception use case, a representative 5G application. With this use case, we can demonstrate that our scheduling approach can robustly lower the cost in different scenarios, while other approaches that are already available fail in either being adaptive to changing environments or will have poor cost-effectiveness in some scenarios. Samuel Rac, Mats Brorsson |
ACM Trans. Archit. Code Optim. | 2 |
| 2023 | DIPPM: A Deep Learning Inference Performance Predictive Model Using Graph Neural NetworksabstractAbstract Deep Learning (DL) has developed to become a corner-stone in many everyday applications that we are now relying on. However, making sure that the DL model uses the underlying hardware efficiently takes a lot of effort. Knowledge about inference characteristics can help to find the right match so that enough resources are given to the model, but not too much. We have developed a DL Inference Performance Predictive Model (DIPPM) that predicts the inference latency, energy, and memory usage of a given input DL model on the NVIDIA A100 GPU. We also devised an algorithm to suggest the appropriate A100 Multi-Instance GPU profile from the output of DIPPM. We developed a methodology to convert DL models expressed in multiple frameworks to a generalized graph structure that is used in DIPPM. It means DIPPM can parse input DL models from various frameworks. Our DIPPM can be used not only helps to find suitable hardware configurations but also helps to perform rapid design-space exploration for the inference performance of a model. We constructed a graph multi-regression dataset consisting of 10,508 different DL models to train and evaluate the performance of DIPPM, and reached a resulting Mean Absolute Percentage Error (MAPE) as low as 1.9%. Karthick Panner Selvam, Mats Brorsson |
Euro-Par | 2 |
| 2023 | Cost-Effective Scheduling for Kubernetes in the Edge-to-Cloud ContinuumabstractThe edge to data center computing continuum is the aggregation of computing resources located anywhere between the network edge (e.g. close to 5G antennas), and servers in traditional data centers. Kubernetes is the de facto standard for container orchestration. It is very efficient in a data center environment, but it fails to give the same performance when adding edge resources. At the edge, resources are more limited, and networking conditions are changing over time.In this paper, we present a methodology that lowers the costs of running applications in the edge-to-cloud computing continuum. A cost-aware scheduler enables this optimization. We are also monitoring the Key Performance Indicators of the applications to ensure that cost optimizations do not impact negatively their Quality of Service. In addition, to ensure that performances are optimal even when users are moving, we introduce a background process that periodically checks if a better location is available for the application. To demonstrate the performance of our scheduling approach, we evaluate it on a vehicle cooperative perception use case, a representative 5G application. Samuel Rac, Mats Brorsson |
IC2E | 2 |
| 2023 | Performance Analysis and Benchmarking of a Temperature Downscaling Deep Learning ModelabstractWe are presenting here a detailed analysis and performance characterization of a statistical temperature downscaling application used in the MAELSTROM EuroHPC project. This application uses a deep learning methodology to convert low-resolution atmospheric temperature states into high-resolution. We have performed in-depth profiling and roofline analysis at different levels (Operators, Training, Distributed Training, Inference) of the downscaling model on different hardware architectures (Nvidia V100 & A100 GPUs). Finally, we compare the training and inference cost of the downscaling model with various cloud providers. Our results identify the model bottlenecks which can be used to enhance the model architecture and determine hardware configuration for efficiently utilizing the HPC. Furthermore, we provide a comprehensive methodology for in-depth profiling and benchmarking of the deep learning models. Karthick Panner Selvam, Mats Brorsson |
PDP | 2 |
| 2022 | Performance Modeling of Weather Forecast Machine Learning for Efficient HPCabstractHigh-performance computing is a prime area for many applications. Majorly, weather and climate forecast applications use the HPC system because it needs to give a good result with low latency. In recent years machine learning and deep learning models have been widely used to forecast the weather. However, to the best of the author’s knowledge, many applications do not effectively utilise the HPC system for training, testing, validation, and inference of weather data. Our experiment is to conduct performance modeling and benchmark analysis of weather and climate forecast machine learning models and determine the characteristics between the application, model and the underlying HPC system. Our results will help the researchers improvise and optimise the weather forecast system and use the HPC system efficiently. Karthick Panner Selvam, Mats Brorsson |
ICDCS | 2 |
| 2020 | Message from Program Co-Chairs: PDP 2020abstractParallel, Distributed, and Network-Based Processing has undergone impressive change over recent years. New architectures and applications have rapidly become the central focus of the discipline. These changes are often a result of cross-fertilization of parallel and distributed technologies with other rapidly evolving technologies. This is the reason why the PDP conference continues to have a distinctive composition: a main track invites papers over a broad range of topics, and ten Special Sessions focus each on a particular sub-domain related to the Parallel, Distributed and Network-based Computing research fields. Each Special Session has its own Chair(s) and Program Committee and invites and selects its own papers, all under the umbrella of the overall conference structure. The growing number of interesting and significant research papers submitted to PDP demonstrates that the conference is becoming an ever more important international event in the field of parallel and distributed computing research. In particular, the Program Committee of this edition received 120 submissions from 31 countries. On average each paper received 3.5 reviews, with no paper receiving fewer than three reviews. Masoud Daneshtalab, Mats Brorsson |
PDP | 2 |
| 2019 | time series modelling of market price in real-time bidding
Manxing Du, Christian A. Hammerschmidt, Georgios Varisteas, Radu State, Mats Brorsson |
ESANN | 5 |
| 2017 | Improving Real-Time Bidding Using a Constrained Markov Decision Process
Manxing Du, Redouane Sassioui, Georgios Varisteas, Radu State, Mats Brorsson, Omar Cherkaoui |
ADMA | 5 |
| 2016 | Node architecture implications for in-memory data analytics on scale-in clustersabstractWhile cluster computing frameworks are continuously evolving to provide real-time data analysis capabilities, Apache Spark has managed to be at the forefront of big data analytics. Recent studies propose scale-in clusters with in-storage processing devices to process big data analytics with Spark However the proposal is based solely on the memory bandwidth characterization of in-memory data analytics and also does not shed light on the specification of host CPU and memory. Through empirical evaluation of in-memory data analytics with Apache Spark on an Ivy Bridge dual socket server, we have found that (i) simultaneous multi-threading is effective up to 6 cores (ii) data locality on NUMA nodes can improve the performance by 10% on average, (iii) disabling next-line L1-D prefetchers can reduce the execution time by up to 14%, (iv) DDR3 operating at 1333 MT/s is sufficient and (v) multiple small executors can provide up to 36% speedup over single large executor. Ahsan Javed Awan, Vladimir Vlassov, Mats Brorsson, Eduard Ayguadé |
BDCAT | 3 |
| 2016 | Behavior profiling for mobile advertisingabstractBehavioral and targeted profiling of users is an important task in marketing and in the advertising industry. Being able to match a given user profile to an advertising that leads to effective purchases is challenging because of a very tiny proportion of users willing to purchase goods and thus monetize the advertising. With such proportions being less than one percent of the overall user population, efficient feature extraction and modeling techniques are required in order to capture and recognize the potential consumers. This paper proposes a new approach for modeling the observed behavior in a mobile advertising platform, where time related features are correlated with additional system level and campaign related performance statistics. We capture the temporal behavior with Hawkes processes and use the estimated parameters as additional features for predicting if a given user profile will be a revenue generating customer. Manxing Du, Radu State, Mats Brorsson, Tigran Avanesov |
BDCAT | 3 |
| 2016 | Grain graphs: OpenMP performance analysis made easyabstractAverage programmers struggle to solve performance problems in OpenMP programs with tasks and parallel for-loops. Existing performance analysis tools visualize OpenMP task performance from the runtime system's perspective where task execution is interleaved with other tasks in an unpredictable order. Problems with OpenMP parallel for-loops are similarly difficult to resolve since tools only visualize aggregate thread-level statistics such as load imbalance without zooming into a per-chunk granularity. The runtime system/threads oriented visualization provides poor support for understanding problems with task and chunk execution time, parallelism, and memory hierarchy utilization, forcing average programmers to rely on experts or use tedious trial-and-error tuning methods for performance. We present grain graphs, a new OpenMP performance analysis method that visualizes grains -- computation performed by a task or a parallel for-loop chunk instance -- and highlights problems such as low parallelism, work inflation and poor parallelization benefit at the grain level. We demonstrate that grain graphs can quickly reveal performance problems that are difficult to detect and characterize in fine detail using existing visualizations in standard OpenMP programs, simplifying OpenMP performance analysis. This enables average programmers to make portable optimizations for poor performing OpenMP programs, reducing pressure on experts and removing the need for tedious trial-and-error tuning. Ananya Muddukrishna, Peter A. Jonsson, Artur Podobas, Mats Brorsson |
PPoPP | 4 |
| 2016 | Palirria: accurate on-line parallelism estimation for adaptive work-stealingabstractSummary We present Palirria, a self‐adapting work‐stealing scheduling method for nested fork/join parallelism that can be used to estimate the number of utilizable workers and self‐adapt accordingly. The estimation mechanism is optimized for accuracy, minimizing the requested resources without degrading performance. We implemented Palirria for both the Linux and Barrelfish operating systems and evaluated it on two platforms: a 48‐core Non‐Uniform Memory Access (NUMA) multiprocessor and a simulated 32‐core system. Compared with state‐of‐the‐art, we observed higher accuracy in estimating resource requirements. This leads to improved resource utilization and performance on par or better to executing with fixed resource allotments. Copyright © 2015 John Wiley & Sons, Ltd. Georgios Varisteas, Mats Brorsson |
Concurr. Comput. Pract. Exp. | 2 |
| 2015 | Green-CM: Energy Efficient Contention Management for Transactional MemoryabstractTransactional memory (TM) is emerging as an attractive synchronization mechanism for concurrent computing. In this work we aim at filling a relevant gap in the TM literature, by investigating the issue of energy efficiency for one crucial building block of TM systems: contention management. Green-CM, the solution proposed in this paper, is the first contention management scheme explicitly designed to jointly optimize both performance and energy consumption. To this end Green-TM combines three key mechanisms: i) it leverages on a novel asymmetric design, which combines different back-off policies in order to take advantage of dynamic frequency and voltage scaling, ii) it introduces an energy efficient design of the back-off mechanism, which combines spin-based and sleep-based implementations, iii) it makes extensive use of self-tuning mechanisms to pursue optimal efficiency across highly heterogeneous workloads. We evaluate Green-CM from both the energy and performance perspectives, and show that it can achieve enhanced efficiency by up to 2.35 times with respect to state of the art contention managers, with an average gain of more than 60% when using 64 threads. Shady Issa, Paolo Romano 0002, Mats Brorsson |
ICPP | 3 |
| 2015 | A comparative performance study of common and popular task-centric programming frameworksabstractSUMMARY Programmers today face a bewildering array of parallel programming models and tools, making it difficult to choose an appropriate one for each application. An increasingly popular programming model supporting structured parallel programming patterns in a portable and composable manner is the task‐centric programming model. In this study, we compare several popular task‐centric programming frameworks, including Cilk Plus, Threading Building Blocks, and various implementations of OpenMP 3.0. We have analyzed their performance on the Barcelona OpenMP Tasking Suite benchmark suite both on a 48‐core AMD Opteron 6172 server and a 64‐core TILEPro64 embedded many‐core processor. Our results show that the OpenMP offers the highest flexibility for programmers, and this flexibility comes to a cost. Frameworks supporting only a specific and more restrictive model, such as Cilk Plus and Threading Building Blocks, are generally more efficient both in terms of performance and energy consumption. However, Intel's implementation of OpenMP tasks performs the best and closest to the specialized run‐time systems. Copyright © 2013 John Wiley & Sons, Ltd. Artur Podobas, Mats Brorsson, Karl-Filip Faxén |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Noodle: A Heuristic Algorithm for Task Scheduling in MPSoC ArchitecturesabstractTask scheduling is crucial for the performance of parallel applications. Given dependence constraints between tasks, their arbitrary sizes, and bounded resources available for execution, optimal task scheduling is considered as an NP-hard problem. Therefore, proposed scheduling algorithms are based on heuristics. This paper1 presents a novel heuristic algorithm, called the Noodle heuristic, which differs from the existing list scheduling techniques in the way it assigns task priorities. We conduct an extensive experimental to validate Noodle for task graphs taken from Standard Task Graph (STG). Results show that Noodle produces schedules that are within a maximum of 12% (in worst-case) of the optimal schedule for 2, 4, and 8 core systems. We also compare Noodle with existing scheduling heuristics and perform comparative analysis of its performance. Muhammad Khurram Bhatti, Isil Öz, Ananya Muddukrishna, Konstantin Popov, Mats Brorsson |
DSD | 5 |
| 2009 | Two-Level Dictionary Code Compression: A New Scheme to Improve Instruction Code Density of Embedded ApplicationsabstractDictionary code compression is a technique which has been studied as a method to reduce the energy consumed in the instruction fetch path of processors. Instructions or instruction sequences in the code are replaced with short code words. These code words are later used to index a dictionary which contains the original uncompressed instruction or an entire sequence. In this paper, we present a new method which improves on code density compared to previously published dictionary methods. It uses a two-level dictionary design and is capable of handling compression of both individual instructions and code sequences of 2-16 instructions. The two dictionaries are in separate pipeline stages and work together to decompress sequences and instructions. The impact on storage size for the dictionaries is rather small as the sequences in the dictionary are stored as individually compressed instructions, instead of normal instructions. Compared to previous dictionary code compression methods we achieve improved dynamic compression rate, potential for better performance with reasonable static compression rate and with still small dictionary size suitable for context switching. Mikael Collin, Mats Brorsson |
CGO | 2 |
| 2006 | Adaptive and flexible dictionary code compression for embedded applicationsabstractDictionary code compression is a technique where long instructions in the memory are replaced with shorter code words used as index in a table to look up the original instructions. We present a new view of dictionary code compression for moderately high-performance processors for embedded applications. Previous work with dictionary code compression has shown decent performance and energy savings results which we verify with our own measurement that are more thorough than previously published. We also augment previous work with a more thorough analysis on the effects of cache and line size changes. In addition, we introduce the concept of aggregated profiling to allow for two or more programs to share the same dictionary contents. Finally, we also introduce dynamic dictionaries where the dictionary contents is considered to be part of the context of a process and show that the performance overhead of reloading the dictionary contents on a context switch is negligible while on the same time we can save considerable energy with a more specialized dictionary contents. Mats Brorsson, Mikael Collin |
CASES | 1 |
| 2004 | A Low Power Strategy for Future Mobile TerminalsabstractIn this paper, we have investigated the efficiency of two power-saving strategies that reduces both static and dynamic power consumption when applied to a chip-multiprocessor (CMP). They are evaluated under two workload scenarios and compared against a conventional uni-processor architecture and a CMP without any power-aware scheduling. The results show that energy due to static and dynamic power consumption can be reduced by up to 78% and that further 8% energy can be saved at the expense of response-time of non-critical applications. Furthermore, a small study on the potential impact of system-level events showed that system calls can contribute significantly to the total energy consumed. Mladen Nikitovic, Mats Brorsson |
DATE | 2 |
| 2004 | A Multiprogrammed Workload Model for Energy and Performance Estimation of Adaptive Chip-MultiprocessorsabstractSummary form only given. Today, there is a trend towards steadily increasing functionality in mobile terminals. This trend in turn increases the performance demand on the architecture that is supposed to do all the work. It is likely that more traditional architectures like multiprocessors are used in future mobile terminals. They are attractive because they can now be integrated on a single chip and can provide the desired performance efficiently if intelligently managed. Choosing the most efficient architecture configuration is however a complex issue and depends on multiple factors. We believe that the way the behavior of the workload is modeled is of paramount importance when estimating the efficiency of any proposed architecture for future mobile terminals. Therefore, a deterministic and simple workload description is needed. In this paper, we show how such a multiprogrammed workload is created and used for energy and performance estimation of an adaptive chip-multiprocessor (CMP) architecture. Mladen Nikitovic, Mats Brorsson |
IPDPS | 2 |
| 2002 | An adaptive chip-multiprocessor architecture for future mobile terminalsabstractPower consumption has become an increasingly important factor in the field of computer architecture. It affects issues such as heat dissipation and packaging cost, which in turn affects the design and cost of a mobile terminal. Today, a lot of effort is put into the design of architectures and software implementation to increase performance. However, little is done on a system level to minimize power consumption, which is crucial in mobile systems.We propose an adaptive chip-multiprocessor (CMP) architecture, where the number of active processors is dynamically adjusted to the current workload need in order to save energy while preserving performance. The architecture is suitable in future mobile terminals where we anticipate a bursty and performance demanding workload.We have carried out an evaluation of the performance and power consumption of the proposed architecture using previously validated high-level simulation models. Our experiments show that orders of magnitude in power consumption can be saved compared to a conventional architecture to a negligable performance cost. The method used is complementary to other power saving techniques such as voltage and frequency scaling. Mladen Nikitovic, Mats Brorsson |
CASES | 2 |
| 2002 | A Fully Compliant OpenMP Implementationon Software Distributed Shared Memory
Sven Karlsson, Sung-Woo Lee, Mats Brorsson |
HiPC | 3 |
| 2001 | Priority Based Messaging for Software Distributed Shared MemoryabstractSoftware Distributed Shared Memory (DSM) systems can be used to provide a coherent shared address space on multicomputers and other parallel systems without support for shared memory in hardware. The coherency software automatically translates shared memory accesses to explicit messages exchanged among the nodes in the system. Many applications exhibit a good performance on such systems but it has been shown that, for some applications, performance critical messages can be delayed behind less important messages because of the enqueuing behavior in the communication libraries used in current systems. We present in this paper a new portable communication library that supports priorities to remedy this situation. We describe an implementation of the communication library and a quantitative model that is used to estimate the performance impact of priorities for a typical situation. Using the model, we show that the use of high-priority communication reduces the latency of performance critical messages substantially over a wide range of network design parameters. The latency is reduced with up to 10–25% for each delaying low priority message in the queue ahead. Sven Karlsson, Mats Brorsson |
IPDPS | 2 |
| 2000 | Special Issue: EWOMP'99 - First European Workshop on OpenMP
Mats Brorsson, Barbara M. Chapman |
Concurr. Pract. Exp. | 1 |
| 2000 | OdinMP/CCp - a portable implementation of OpenMP for CabstractWe describe here the design and performance of OdinMP/CCp, which is a portable compiler for C-programs using the OpenMP directives for parallel processing with shared memory. OdinMP/CCp was written in Java for portability reasons and takes a C-program with OpenMP directives and produces a C-program for POSIX threads. We describe some of the ideas behind the design of OdinMP/CCp and show some performanceresults achieved on an SGI Origin 2000 and a Sun E10000. Speedup measurements relative to a sequential version of the test programs show that OpenMP programs using OdinMP/CCp exhibit excellent performance on the Sun E10000 and reasonable performance on the Origin 2000. Copyright © 2000 John Wiley & Sons, Ltd. Christian Brunschen, Mats Brorsson |
Concurr. Pract. Exp. | 2 |
| 1999 | Programming Effort vs. Performance with a Hybrid Programming Model for Distributed Memory Parallel Architectures
Andreas Rodman, Mats Brorsson |
Euro-Par | 2 |
| 1999 | Producer-Push - A Protocol Enhancement to Page-Based Software Distributed Shared Memory SystemsabstractThis paper describes a technique called producer-push that enhances the performance of a page-based software distributed shared memory system. Shared data, in software DSM systems, must normally be requested from the node that produced the latest value. Producer-push utilizes the execution history to predict this communication so that the data is pushed to the consumer before it is requested. In contrast to previously proposed mechanisms to proactively send data to where it is needed, producer-push uses information about the source code location of communication to more accurately predict the needed communication. Producer-push requires no source code modifications of the application and it effectively reduces the latency of shared memory accesses. This is confirmed by our performance evaluation which shows that the average time to wait for memory updates is reduced by 74%. Producer-push also changes the communication pattern of an application making it more suitable for modern networks. The latter is a result of a 44% reduction of the average number of messages and an enlargement of the average message size by 65%. Sven Karlsson, Mats Brorsson |
ICPP | 2 |
| 1999 | Performance Tuning Software DSM Applications using Visualisation
Mats Brorsson, Martin Kral |
J. Supercomput. | 1 |
| 1996 | Characterising and Modelling Shared Memory Accesses in Multiprocessor Programs
Mats Brorsson, Per Stenström |
Parallel Comput. | 1 |
| 1995 | SM-prof: A Tool to Visualise and Find Cache Coherence Performance Bottlenecks in Multiprocessor ProgramsabstractCache misses due to coherence actions are often the major source for performance degradation in cache coherent multiprocessors. It is often difficult for the programmer to take cache coherence into account when writing the program since the resulting access pattern is not apparent until the program is executed.SM-prof is a performance analysis tool that addresses this problem by visualising the shared data access pattern in a diagram with links to the source code lines causing performance degrading access patterns. The execution of a program is divided into time slots and each data block is classified based on the accesses made to the block during a time slot. This enables the programmer to follow the execution over time and it is possible to track the exact position responsible for accesses causing many cache misses related to coherence actions.Matrix multiplication and the MP3D application from SPLASH are used to illustrate the use of SM-prof. For MP3D, SM-prof revealed performance limitations that resulted in a performance improvement of over 75%.The current implementation is based on program-driven simulation in order to achieve non-intrusive profiling. If a small perturbation of the program execution is acceptable, it is also possible to use software tracing techniques given that a data address can be related to the originating instruction. Mats Brorsson |
SIGMETRICS | 1 |
| 1993 | An Adaptive Cache Coherence Protocol Optimized for Migratory SharingabstractParallel programs that use critical sections and are executed on a shared-memory multiprocessor with a write-invalidate protocol result in invalidation actions that could be eliminated. For this type of sharing, called migratory sharing, each processor typically causes a cache miss followed by an invalidation request which could be merged with the preceding cache-miss request. Per Stenström, Mats Brorsson, Lars Sandberg |
ISCA | 2 |