EDBT 2026 Demo / reviewers in the wild / expert
Marius Evers
dblp:64/5625
· DBLP profile ↗
7ranked-venue papers
3as first author
0since 2021 · last 2007
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Processor architecture and microarchitecture · 95% Performance modeling and evaluation · 4% Memory systems · 1% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
branch prediction |
0.1 | 5 | 2001 | Understanding branches and designing branch predictors for high-performance microprocessors · Proc. IEEE 2001 Improving Trace Cache Effectiveness with Branch Promotion and Trace Packing · ISCA 1998 An Analysis of Correlation and Predictability: What Makes Two-Level Branch Predictors Work · ISCA 1998 |
Processor architecture and microarchitecture
instruction fetch |
0.1 | 3 | 2001 | Understanding branches and designing branch predictors for high-performance microprocessors · Proc. IEEE 2001 Improving Trace Cache Effectiveness with Branch Promotion and Trace Packing · ISCA 1998 Increasing the Instruction Fetch Rate via Block-structured Instruction Set Architectures · MICRO 1996 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 2 | 1998 | Increasing the Instruction Fetch Rate via Block-structured Instruction Set Architectures · MICRO 1996 Variable Length Path Branch Prediction · ASPLOS 1998 |
Processor architecture and microarchitecture › branch prediction
path-based prediction |
0.0 | 1 | 1998 | Variable Length Path Branch Prediction · ASPLOS 1998 |
Processor architecture and microarchitecture › instruction fetch
trace cache |
0.0 | 1 | 1998 | Improving Trace Cache Effectiveness with Branch Promotion and Trace Packing · ISCA 1998 |
Processor architecture and microarchitecture › branch prediction
two-level adaptive branch prediction |
0.0 | 1 | 1998 | An Analysis of Correlation and Predictability: What Makes Two-Level Branch Predictors Work · ISCA 1998 |
Processor architecture and microarchitecture › branch prediction
hybrid branch predictor |
0.0 | 1 | 1996 | Using Hybrid Branch Predictors to Improve Branch Prediction Accuracy in the Presence of Context Switches · ISCA 1996 |
Processor architecture and microarchitecture › superscalar processor
wide-issue processor |
0.0 | 1 | 1996 | Increasing the Instruction Fetch Rate via Block-structured Instruction Set Architectures · MICRO 1996 |
Processor architecture and microarchitecture
pipelining |
0.0 | 1 | 2001 | Understanding branches and designing branch predictors for high-performance microprocessors · Proc. IEEE 2001 |
Performance modeling and evaluation › workload characterization › program behavior
branch behavior |
0.0 | 1 | 1998 | An Analysis of Correlation and Predictability: What Makes Two-Level Branch Predictors Work · ISCA 1998 |
Processor architecture and microarchitecture
speculative execution |
0.0 | 1 | 1998 | Variable Length Path Branch Prediction · ASPLOS 1998 |
Processor architecture and microarchitecture
superscalar processor |
0.0 | 1 | 1998 | Improving Trace Cache Effectiveness with Branch Promotion and Trace Packing · ISCA 1998 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 1998 | An Analysis of Correlation and Predictability: What Makes Two-Level Branch Predictors Work · ISCA 1998 |
Memory systems
context switch effects |
0.0 | 1 | 1996 | Using Hybrid Branch Predictors to Improve Branch Prediction Accuracy in the Presence of Context Switches · ISCA 1996 |
Processor architecture and microarchitecture › pipelining
pipeline stall |
0.0 | 1 | 1996 | Using Hybrid Branch Predictors to Improve Branch Prediction Accuracy in the Presence of Context Switches · ISCA 1996 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.0trace cache simulation · 0.0profiling · 0.0dynamic branch prediction · 0.0correlation analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2007 | Circuit optimization for leakage power reduction using multi-threshold voltages for high performance microprocessorsabstractA common concern as we scale down transistor threshold voltages while migrating to new process technologies is the requirement to achieve timing closure within a given power budget over various process corners. High performance microprocessors are designed keeping in mind the various process technologies, application space and multi-site fabrication requirements. Described here is an optimization methodology and a unique topology-aware heuristic algorithm employed for high speed microprocessor designs capable of simultaneous threshold voltage selection for library cells across various technology process corners. The algorithm uses knowledge of the circuit topology rather than considering only the immediate local connectivity as is suggested in other heuristic methods and evaluates timing criticalities originating from different input and output logic cones associated with every pin of a failing path. The VTH selection is done so as to affect multiple failing paths with each low VTH cell selection, hence reducing leakage power. Two sets of algorithms are used alternately. One takes advantage of the circuit topology to address multiple failing paths simultaneously. The other performs a fine tuned optimization that has more granularity while considering a particular failing path. This flow is not limited to dual threshold VTH selection but can also support the use of multi-VTH library cells. This flow and its algorithms reduced the usage of low VTH in a particular multi-million transistor design from 35.3% to 10.7% without any loss of performance thus resulting in a 55.6% drop in leakage power. Reducing the usage of lower VTH cells results in significant power reduction. This reduction in power could also allow running the chip at a higher VDD and frequency within the original power envelope. Production results from this tool exceeded the optimization efforts of another commercially used EDA optimization tool. Jeegar Tilak Shah, Marius Evers, Jeff Trull, Alper Halbutogullari |
ISPD | 2 |
| 2001 | Understanding branches and designing branch predictors for high-performance microprocessorsabstractBranch prediction is important in high-performance processors and its importance continues to grow. In the drive for higher execution frequencies, pipelines are lengthened and memory latencies are increased. This increases the cost of branch mispredictions. In this paper we describe some behavior patterns of branches. We believe that understanding the behavior of branches is helpful when designing fetch mechanisms for high-performance microprocessors. We also examine several current branch predictors and discuss how they work. Finally, we look at some of the challenges that we are faced with when designing fetch mechanisms and predictors for future microprocessors and discuss some of the possible solutions. Marius Evers, Tse-Yu Yeh |
Proc. IEEE | 1 |
| 1998 | Variable Length Path Branch PredictionabstractAccurate branch prediction is required to achieve high performance in deeply pipelined, wide-issue processors. Recent studies have shown that conditional and indirect (or computed) branch targets can be accuratelypredicted by recording the path, which consists of the target addresses of recent branches, leading up to the branch. In current path based branch predictors, the N most recent target addresses are hashed together to form an index into a table, where N is some fixed integer. The indexed table entry isused to make a prediction for the current branch.This paper introduces a new branch predictor in which the value of N is allowed to vary. By constructing the index into the table using the last N target addresses, and using profiling information to select the proper value of N for each branch, extremely accurate branch prediction is achieved. For the SPECint95 gee benchmark, this new predictor has a conditional branch misprediction rate of 4.3% given a 4K byte hardware budget. For comparison, the gshare predictor, a predictor known for its high accuracy, has a conditional branch misprediction rate of 8.8% given the same hardware budget. For the indirect branches in gee, the new predictor achieves a misprediction rate of 27.7% when given a hardware budget of 512 bytes, whereas the best competingpredictor achieves a misprediction rate of 44.2% when given the same hardware budget. Jared Stark, Marius Evers, Yale N. Patt |
ASPLOS | 2 |
| 1998 | An Analysis of Correlation and Predictability: What Makes Two-Level Branch Predictors WorkabstractPipeline flushes due to branch mispredictions is one of the most serious problems facing the designer of a deeply pipelined, superscalar processor. Many branch predictors have been proposed to help alleviate this problem, including two-level adaptive branch predictors and hybrid branch predictors. Numerous studies have shown which predictors and configurations best predict the branches in a given set of benchmarks. Some studies have also investigated effects, such as pattern history table interference, that can be detrimental to the performance of these predictors. However, little research has been done on which characteristics of branch behavior make predictors perform well. In this paper we investigate and quantify reasons why branches are predictable. We show that some of this predictability is not captured by the two-level adaptive branch predictors. An understanding of the predictability of branches may lead to insights ultimately resulting in better or less complex predictors. We also investigate and quantify what function of the branches in each benchmark is predictable using each of the methods described in this paper. Marius Evers, Sanjay J. Patel, Robert Chappell, Yale N. Patt |
ISCA | 1 |
| 1998 | Improving Trace Cache Effectiveness with Branch Promotion and Trace PackingabstractThe increasing widths of superscalar processors are placing greater demands upon the fetch mechanism. The trace cache meets these demands by placing logically contiguous instructions in physically contiguous storage. As a result, the trace cache delivers instructions at a high rate by supplying multiple fetch blocks each cycle. In this paper we examine two techniques to improve the number of instructions delivered each cycle by the trace cache. The first technique, branch promotion, dynamically converts strongly biased branches into branches with static predictions. Because these promoted branches require no dynamic prediction, the branch predictor suffers less from the negative effects of interference. Branch promotion unlocks the potential of the second technique: trace packing. With trace packing, trace segments are packed with as many instructions as will fit, without regard to naturally occurring fetch block boundaries. With both techniques, the effective fetch rate of the trace cache jumps up 17% over a trace cache which implements neither on a machine where the execution engine has a very aggressive memory disambiguator; the performance of a machine using branch promotion and trace packing is on average 11% higher than a machine using neither technique. Sanjay J. Patel, Marius Evers, Yale N. Patt |
ISCA | 2 |
| 1996 | Using Hybrid Branch Predictors to Improve Branch Prediction Accuracy in the Presence of Context SwitchesabstractPipeline stalls due to conditional branches represent one of the most significant impediments to realizing the performance potential of deeply pipelined, superscalar processors. Many branch predictors have been proposed to help alleviate this problem, including the Two-Level Adaptive Branch Predictor, and more recently, two-component hybrid branch predictors.In a less idealized environment, such as a time-shared system, code of interest involves context switches. Context switches, even at fairly large intervals, can seriously degrade the performance of many of the most accurate branch prediction schemes. In this paper, we introduce a new hybrid branch predictor and show that it is more accurate (for a given cost) than any previously published scheme, especially if the branch histories are periodically flushed due to the presence of context switches. Marius Evers, Po-Yung Chang, Yale N. Patt |
ISCA | 1 |
| 1996 | Increasing the Instruction Fetch Rate via Block-structured Instruction Set ArchitecturesabstractTo exploit larger amounts of instruction level parallelism, processors are being built with wider issue widths and larger numbers of functional units. Instruction fetch rate must also be increased in order to effectively exploit the performance potential of such processors. Block-structured ISAs provide an effective means of increasing the instruction fetch rare. We define an optimization, called block enlargement, that can be applied to a block-structured ISA to increase the instruction fetch rate of a processor that implements that ISA. We have constructed a compiler that generates block-structured ISA code, and a simulator that models the execution of that code on a block-structured ISA processor. We show that for the SPECint95 benchmarks, the block-structured ISA processor executing enlarged atomic blocks outperforms a conventional ISA processor by 12% while using simpler microarchitectural mechanisms to support wide-issue and dynamic scheduling. Eric Hao, Po-Yung Chang, Marius Evers, Yale N. Patt |
MICRO | 3 |