EDBT 2026 Demo / reviewers in the wild / expert
Pierre Michaud
dblp:84/2225
· DBLP profile ↗
20ranked-venue papers
13as first author
1since 2021 · last 2022
0000-0001-7037-4014ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 11 first-author · 1 since 2021Software engineering, systems software and programming languages · 7 · 3 first-authorDatabases, data management, data science and information retrieval · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
13 papers |
Processor architecture and microarchitecture · 66% Memory systems · 16% Performance modeling and evaluation · 6% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 30 heaviest of 34, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
superscalar processor |
0.8 | 2 | 2022 | HAIR: Halving the Area of the Integer Register File with Odd/Even Banking · ACM Trans. Archit. Code Optim. 2022 Revisiting Clustered Microarchitecture for Future Superscalar Cores: A Case for Wide Issue Clusters · ACM Trans. Archit. Code Optim. 2015 |
Processor architecture and microarchitecture › register file
banked register file |
0.6 | 1 | 2022 | HAIR: Halving the Area of the Integer Register File with Odd/Even Banking · ACM Trans. Archit. Code Optim. 2022 |
Processor architecture and microarchitecture
register file |
0.6 | 1 | 2022 | HAIR: Halving the Area of the Integer Register File with Odd/Even Banking · ACM Trans. Archit. Code Optim. 2022 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.4 | 2 | 2015 | Revisiting Clustered Microarchitecture for Future Superscalar Cores: A Case for Wide Issue Clusters · ACM Trans. Archit. Code Optim. 2015 Long term parking (LTP): criticality-aware resource allocation in OOO processors · MICRO 2015 |
Processor architecture and microarchitecture
out-of-order execution |
0.4 | 2 | 2015 | Long term parking (LTP): criticality-aware resource allocation in OOO processors · MICRO 2015 Cost-effective speculative scheduling in high performance processors · ISCA 2015 |
Processor architecture and microarchitecture
branch prediction |
0.4 | 3 | 2018 | An Alternative TAGE-like Conditional Branch Predictor · ACM Trans. Archit. Code Optim. 2018 Trading Conflict and Capacity Aliasing in Conditional Branch Predictors · ISCA 1997 Multiple-Block Ahead Branch Predictors · ASPLOS 1996 |
Processor architecture and microarchitecture › branch prediction
conditional branch predictor |
0.3 | 1 | 2018 | An Alternative TAGE-like Conditional Branch Predictor · ACM Trans. Archit. Code Optim. 2018 |
Processor architecture and microarchitecture › branch prediction
TAGE predictor |
0.3 | 1 | 2018 | An Alternative TAGE-like Conditional Branch Predictor · ACM Trans. Archit. Code Optim. 2018 |
Memory systems
cache |
0.3 | 3 | 2016 | Best-offset hardware prefetching · HPCA 2016 Cost-effective speculative scheduling in high performance processors · ISCA 2015 Trading Conflict and Capacity Aliasing in Conditional Branch Predictors · ISCA 1997 |
Memory systems › cache
cache behavior |
0.2 | 1 | 2016 | Some Mathematical Facts About Optimal Cache Replacement · ACM Trans. Archit. Code Optim. 2016 |
Memory systems › cache management
cache replacement |
0.2 | 1 | 2016 | Some Mathematical Facts About Optimal Cache Replacement · ACM Trans. Archit. Code Optim. 2016 |
Memory systems › cache › prefetching
hardware prefetching |
0.2 | 1 | 2016 | Best-offset hardware prefetching · HPCA 2016 |
Processor architecture and microarchitecture › clustered architecture
clustered microarchitecture |
0.2 | 1 | 2015 | Revisiting Clustered Microarchitecture for Future Superscalar Cores: A Case for Wide Issue Clusters · ACM Trans. Archit. Code Optim. 2015 |
Processor architecture and microarchitecture
instruction scheduling |
0.2 | 1 | 2015 | Cost-effective speculative scheduling in high performance processors · ISCA 2015 |
Embedded and real-time systems › real-time scheduling
mixed-criticality scheduling |
0.2 | 1 | 2015 | Long term parking (LTP): criticality-aware resource allocation in OOO processors · MICRO 2015 |
Cloud and datacenter computing
resource allocation |
0.2 | 1 | 2015 | Long term parking (LTP): criticality-aware resource allocation in OOO processors · MICRO 2015 |
Processor architecture and microarchitecture › instruction scheduling
speculative scheduling |
0.2 | 1 | 2015 | Cost-effective speculative scheduling in high performance processors · ISCA 2015 |
Performance modeling and evaluation
benchmarking |
0.2 | 1 | 2014 | Multiprogram Throughput Metrics: A Systematic Approach · ACM Trans. Archit. Code Optim. 2014 |
Energy-efficient computing
power management |
0.2 | 1 | 2022 | HAIR: Halving the Area of the Integer Register File with Odd/Even Banking · ACM Trans. Archit. Code Optim. 2022 |
Performance modeling and evaluation › benchmarking
benchmark evaluation |
0.1 | 1 | 2016 | Best-offset hardware prefetching · HPCA 2016 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 1 | 2007 | A study of thread migration in temperature-constrained multicores · ACM Trans. Archit. Code Optim. 2007 |
Energy-efficient computing
thermal management |
0.1 | 1 | 2007 | A study of thread migration in temperature-constrained multicores · ACM Trans. Archit. Code Optim. 2007 |
Parallel and multicore computing › parallel programming runtimes › thread management
thread migration |
0.1 | 1 | 2007 | A study of thread migration in temperature-constrained multicores · ACM Trans. Archit. Code Optim. 2007 |
Memory systems › memory hierarchy › cache hierarchy
l1 cache |
0.1 | 1 | 2015 | Cost-effective speculative scheduling in high performance processors · ISCA 2015 |
Memory systems › memory access optimization
memory-level parallelism |
0.1 | 1 | 2015 | Long term parking (LTP): criticality-aware resource allocation in OOO processors · MICRO 2015 |
Distributed systems › remote execution
computation migration |
0.0 | 1 | 2004 | Exploiting the Cache Capacity of a Single-Chip Multi-Core Processor with Execution Migration · HPCA 2004 |
Processor architecture and microarchitecture › out-of-order execution
instruction window |
0.0 | 1 | 2001 | Data-Flow Prescheduling for Large Instruction Windows in Out-of-Order Processors · HPCA 2001 |
Processor architecture and microarchitecture › out-of-order execution
out-of-order processor |
0.0 | 1 | 2001 | Data-Flow Prescheduling for Large Instruction Windows in Out-of-Order Processors · HPCA 2001 |
Parallel and multicore computing › parallel scheduling
thread scheduling |
0.0 | 1 | 2007 | A study of thread migration in temperature-constrained multicores · ACM Trans. Archit. Code Optim. 2007 |
Compilers and program optimization
instruction scheduling |
0.0 | 1 | 2001 | Data-Flow Prescheduling for Large Instruction Windows in Out-of-Order Processors · HPCA 2001 |
Methods — techniques the papers use, named apart from their topics
microarchitectural simulation · 0.6statistical correction · 0.3sandbox method · 0.2mathematical proof · 0.2OPT tokens algorithm · 0.2cycle-accurate simulation · 0.2criticality prediction · 0.2weighted speedup · 0.2harmonic mean of speedups · 0.2performance simulation · 0.1simulation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | HAIR: Halving the Area of the Integer Register File with Odd/Even BankingabstractThis article proposes a new microarchitectural scheme for reducing the hardware complexity of the integer register file of a superscalar processor. The register file is split into two banks holding even-numbered and odd-numbered physical registers, respectively. Each bank provides one read port to each two-input integer execution unit. This way, each bank has half the total number of read ports, and the register file area is roughly halved, which reduces the energy dissipated per register access and the register access time. However, a bank conflict occurs when both inputs of a two-input micro-operation lie in the same bank. Bank conflicts hurt performance, and we propose a simple solution to remove most bank conflicts, thus recovering most of the lost performance. Pierre Michaud, Anis Peysieux |
ACM Trans. Archit. Code Optim. | 1 |
| 2018 | An Alternative TAGE-like Conditional Branch PredictorabstractTAGE is one of the most accurate conditional branch predictors known today. However, TAGE does not exploit its input information perfectly, as it is possible to obtain significant prediction accuracy improvements by complementing TAGE with a statistical corrector using the same input information. This article proposes an alternative TAGE-like predictor making statistical correction practically superfluous. Pierre Michaud |
ACM Trans. Archit. Code Optim. | 1 |
| 2016 | Best-offset hardware prefetchingabstractHardware prefetching is an important feature of modern high-performance processors. When the application working set is too large to fit in on-chip caches, disabling hardware pre-fetchers may result in severe performance reduction. A new prefetcher was recently introduced, the Sandbox prefetcher, that tries to find dynamically the best prefetch offset using the sandbox method. The Sandbox prefetcher uses simple hardware and was shown to be quite effective. However, the sandbox method does not take into account prefetch timeliness. We propose an offset prefetcher with a new method for selecting the prefetch offset that takes into account prefetch timeliness. We show that our Best-Offset prefetcher outperforms the Sandbox prefetcher on the SPEC CPU2006 benchmarks, with equally simple hardware. Pierre Michaud |
HPCA | 1 |
| 2016 | A simple proof of optimality for the MIN cache replacement policy
Mun-Kyu Lee, Pierre Michaud, Jeong Seop Sim, DaeHun Nyang |
Inf. Process. Lett. | 2 |
| 2016 | Some Mathematical Facts About Optimal Cache ReplacementabstractThis article exposes and proves some mathematical facts about optimal cache replacement that were previously unknown or not proved rigorously. An explicit formula is obtained, giving OPT hits and misses as a function of past references. Several mathematical facts are derived from this formula, including a proof that OPT miss curves are always convex, and a new algorithm called OPT tokens , for reasoning about optimal replacement. Pierre Michaud |
ACM Trans. Archit. Code Optim. | 1 |
| 2015 | Cost-effective speculative scheduling in high performance processorsabstractTo maximize performance, out-of-order execution processors sometimes issue instructions without having the guarantee that operands will be available in time; e.g. loads are typically assumed to hit in the L1 cache and dependent instructions are issued accordingly. This form of speculation -- that we refer to as speculative scheduling -- has been used for two decades in real processors, but has received little attention from the research community. Arthur Perais, André Seznec, Pierre Michaud, Andreas Sembrant, Erik Hagersten |
ISCA | 3 |
| 2015 | Revisiting symbiotic job schedulingabstractSymbiotic job scheduling exploits the fact that in a system with shared resources, the performance of jobs is impacted by the behavior of other co-running jobs. By coscheduling combinations of jobs that have low interference, the performance of a system can be increased. In this paper, we investigate the impact of using symbiotic job scheduling for increasing throughput. We find that even for a theoretically optimal scheduler, this impact is very low, despite the substantial sensitivity of per job performance to which other jobs are coscheduled: for example, our experiments on a 4-thread SMT processor show that, on average, the job IPC varies by 37% depending on coscheduled jobs, the per-coschedule throughput varies by 69%, and yet the average throughput gain brought by optimal symbiotic scheduling is only 3%. This small margin of improvement can be explained by the observation that all the jobs need to be eventually executed, restricting the job combinations a symbiotic job scheduler can select to optimize throughput. We explain why previous work reported a substantial gain from symbiotic job scheduling, and we find that (only) reporting turnaround time can lead to misleading conclusions. Furthermore, we show how the impact of scheduling can be evaluated in microarchitectural studies, without having to implement a scheduler. Stijn Eyerman, Pierre Michaud, Wouter Rogiest |
ISPASS | 2 |
| 2015 | Long term parking (LTP): criticality-aware resource allocation in OOO processorsabstractModern processors employ large structures (IQ, LSQ, register file, etc.) to expose instruction-level parallelism (ILP) and memory-level parallelism (MLP). These resources are typically allocated to instructions in program order. This wastes resources by allocating resources to instructions that are not yet ready to be executed and by eagerly allocating resources to instructions that are not part of the application's critical path. Andreas Sembrant, Trevor E. Carlson, Erik Hagersten, David Black-Schaffer, Arthur Perais, André Seznec, Pierre Michaud |
MICRO | 7 |
| 2015 | Revisiting Clustered Microarchitecture for Future Superscalar Cores: A Case for Wide Issue ClustersabstractDuring the past 10 years, the clock frequency of high-end superscalar processors has not increased. Performance keeps growing mainly by integrating more cores on the same chip and by introducing new instruction set extensions. However, this benefits only some applications and requires rewriting and/or recompiling these applications. A more general way to accelerate applications is to increase the IPC, the number of instructions executed per cycle. Although the focus of academic microarchitecture research moved away from IPC techniques, the IPC of commercial processors was continuously improved during these years. We argue that some of the benefits of technology scaling should be used to raise the IPC of future superscalar cores further. Starting from microarchitecture parameters similar to recent commercial high-end cores, we show that an effective way to increase the IPC is to allow the out-of-order engine to issue more micro-ops per cycle. But this must be done without impacting the clock cycle. We propose combining two techniques: clustering and register write specialization. Past research on clustered microarchitectures focused on narrow issue clusters, as the emphasis at that time was on allowing high clock frequencies. Instead, in this study, we consider wide issue clusters, with the goal of increasing the IPC under a constant clock frequency. We show that on a wide issue dual cluster, a very simple steering policy that sends 64 consecutive instructions to the same cluster, the next 64 instructions to the other cluster, and so forth, permits tolerating an intercluster delay of three cycles. We also propose a method for decreasing the energy cost of sending results from one cluster to the other cluster. Pierre Michaud, Andrea Mondelli, André Seznec |
ACM Trans. Archit. Code Optim. | 1 |
| 2014 | Multiprogram Throughput Metrics: A Systematic ApproachabstractRunning multiple programs on a processor aims at increasing the throughput of that processor. However, defining meaningful throughput metrics in a simulation environment is not as straightforward as reporting execution time. This has led to an ongoing debate on what forms a meaningful throughput metric for multiprogram workloads. We present a method to construct throughput metrics in a systematic way: we start by expressing assumptions on job size, job distribution, scheduling, and so forth that together define a theoretical throughput experiment. The throughput metric is then the average throughput of this experiment. Different assumptions lead to different metrics, so one should be aware of these assumptions when making conclusions based on results using a specific metric. Throughput metrics should always be defined from explicit assumptions, because this leads to a better understanding of the implications and limits of the results obtained with that metric. We elaborate multiple metrics based on different assumptions. In particular, we identify the assumptions that lead to the commonly used weighted speedup and harmonic mean of speedups. Our study clarifies that they are actual throughput metrics, which was recently questioned. We also propose some new throughput metrics, which cannot always be expressed as a closed formula. We use real experimental data to characterize metrics and show how they relate to each other. Stijn Eyerman, Pierre Michaud, Wouter Rogiest |
ACM Trans. Archit. Code Optim. | 2 |
| 2013 | Selecting benchmark combinations for the evaluation of multicore throughputabstractMost high-performance processors today are able to execute multiple threads of execution simultaneously. Threads share processor resources, like the last-level cache, which may decrease throughput in a non obvious way, depending on threads' characteristics. Computer architects usually study multiprogrammed workloads by considering a set of benchmarks and some combinations of these benchmarks. Because detailed microarchitecture simulators are slow, we want a subset of combinations that is as small as possible, yet representative. However, there is no standard method for selecting such sample, and different authors have used different methods. It is not clear how the choice of a particular sample impacts the conclusions of a study. We propose and compare different sampling methods for defining multiprogrammed workloads for computer architecture studies. We evaluate their effectiveness with a case study, the comparison of several multicore last-level cache replacement policies. We show that random sampling, the simplest method, is a possible way to define a representative workload sample, provided the sample is large enough. We propose a method for estimating the required sample size based on fast approximate simulation. We propose a new method, workload stratification, which is very effective at reducing the sample size in situations where random sampling would require large samples. Ricardo A. Velásquez 0001, Pierre Michaud, André Seznec |
ISPASS | 2 |
| 2011 | Replacement policies for shared caches on symmetric multicores: a programmer-centric point of viewabstractThe presence of shared caches in current multicore processors may generate a lot of performance variability in multi-programmed environments. For applications with quality-of-service requirements, this performance variability may lead the programmer to be overly pessimistic about performance and reduce the application features and/or spend a lot of effort optimizing the algorithms. To solve this problem, there must be a way for the programmer to define a reasonable performance target and a guarantee that the actual performance is very unlikely to be below the targeted performance. We propose that the performance target be defined as the performance measured when each core runs a copy of the application, which we call self-performance. This study characterizes self-performance and explains how the shared-cache replacement policy can be modified for self-performance to be meaningful. Pierre Michaud |
HiPEAC | 1 |
| 2009 | Online compression of cache-filtered address tracesabstractTrace-driven simulation is potentially much faster than cycle-accurate simulation. However, one drawback is the large amount of storage that may be necessary to store traces. Trace compression techniques are useful for decreasing the storage space requirement. But the compression ratio of existing trace compressors is limited because they implement lossless compression. We propose two new methods for compressing cache filtered address traces. The first method, byte sort, is a lossless compression method that achieves high compression ratios on cache-filtered address traces. The second method is a lossy one, based on the concept of phase. We have combined these two methods in a trace compressor called ATC. Our experimental results show that ATC gives high compression ratio while keeping the memory-locality characteristics of the original trace. Pierre Michaud |
ISPASS | 1 |
| 2007 | A study of thread migration in temperature-constrained multicoresabstractTemperature has become an important constraint in high-performance processors, especially multicores. Thread migration will be essential to exploit the full potential of future thermally constrained multicores. We propose and study a thread migration method that maximizes performance under a temperature constraint, while minimizing the number of migrations and ensuring fairness between threads. We show that thread migration brings important performance gains and that it is most effective during the first tens of seconds following a decrease of the number of running threads. Pierre Michaud, André Seznec, Damien Fetis, Yiannakis Sazeides, Theofanis Constantinou |
ACM Trans. Archit. Code Optim. | 1 |
| 2004 | Exploiting the Cache Capacity of a Single-Chip Multi-Core Processor with Execution MigrationabstractWe propose to modify a conventional single-chip multicore so that a sequential program can migrate from one core to another automatically during execution. The goal of execution migration is to take advantage of the overall on-chip cache capacity. We introduce the affinity algorithm, a method for distributing cache lines automatically on several caches. We show that on working-sets exhibiting a property called "splittability", it is possible to trade cache misses for migrations. Our experimental results indicate that the proposed method has a potential for improving the performance of certain sequential programs, without degrading significantly the performance of others. Pierre Michaud |
HPCA | 1 |
| 2003 | A statistical model of skewed-associativityabstractThis paper presents a statistical model for explaining why skewed-associativity removes conflicts better than set-associativity. We show that, with a high probability, 2-way skewed associativity emulates full associativity for working-sets up to half the cache size, and we show that 3-way skewed-associativity is almost equivalent to full associativity. Pierre Michaud |
ISPASS | 1 |
| 2001 | Data-Flow Prescheduling for Large Instruction Windows in Out-of-Order ProcessorsabstractThe performance of out-of-order processors increases with the instruction window size, In conventional processors, the effective instruction window cannot be larger than the issue buffer. Determining which instructions from the issue buffer can be launched to the execution units is a time-critical operation which complexity increases with the issue buffer size. We propose to relieve the issue stage by reordering instructions before they enter the issue buffer. This study introduces the general principle of data flow prescheduling. Then we describe a possible implementation. Our preliminary results show that data-flow prescheduling makes it possible to enlarge the effective instruction window while keeping the issue buffer small. Pierre Michaud, André Seznec |
HPCA | 1 |
| 1997 | Trading Conflict and Capacity Aliasing in Conditional Branch PredictorsabstractAs modern microprocessors employ deeper pipelines and issue multiple instructions per cycle, they are becoming increasingly dependent on accurate branch prediction. Because hardware resources for branch-predictor tables are invariably limited, it is not possible to hold all relevant branch history for all active branches at the same time, especially for large workloads consisting of multiple processes and operating-system code. The problem that results, commonly referred to as aliasing in the branch-predictor tables, is in many ways similar to the misses that occur in finite-sized hardware caches.In this paper we propose a new classification for branch aliasing based on the three-Cs model for caches, and show that conflict aliasing is a significant source of mispredictions. Unfortunately, the obvious method for removing conflicts --- adding tags and associativity to the predictor tables --- is not a cost-effective solution.To address this problem, we propose the skewed branch predictor, a multi-bank, tag-less branch predictor, designed specifically to reduce the impact of conflict aliasing. Through both analytical and simulation models, we show that the skewed branch predictor removes a substantial portion of conflict aliasing by introducing redundancy to the branch-predictor tables. Although this redundancy increases capacity aliasing compared to a standard one-bank structure of comparable size, our simulations show that the reduction in conflict aliasing overcomes this effect to yield a gain in prediction accuracy. Alternatively, we show that a skewed organization can achieve the same prediction accuracy as a standard one-bank organization but with half the storage requirements. Pierre Michaud, André Seznec, Richard Uhlig |
ISCA | 1 |
| 1997 | Clustering techniques
Pierre Michaud |
Future Gener. Comput. Syst. | 1 |
| 1996 | Multiple-Block Ahead Branch PredictorsabstractA basic rule in computer architecture is that a processor cannot execute an application faster than it fetches its instructions. This paper presents a novel cost-effective mechanism called the two-block ahead branch predictor. Information from the current instruction block is not used for predicting the address of the next instruction block, but rather for predicting the block following the next instruction block.This approach overcomes the instruction fetch bottle-neck exhibited by wide-dispatch "brainiac" processors by enabling them to efficiently predict addresses of two instruction blocks in a single cycle. Furthermore, pipelining the branch prediction process can also be done by means of our predictor for "speed demon" processors to achieve higher clock rate or to improve the prediction accuracy by means of bigger prediction structures.Moreover, and unlike the previously-proposed multiple predictor schemes, multiple-block ahead branch predictors can use any of the branch prediction schemes to perform the very accurate predictions required to achieve high-performance on superscalar processors. André Seznec, Stéphan Jourdan, Pascal Sainrat, Pierre Michaud |
ASPLOS | 4 |