EDBT 2026 Demo / reviewers in the wild / expert
Robert H. Bell Jr.
dblp:72/442
· DBLP profile ↗
7ranked-venue papers
3as first author
0since 2021 · last 2011
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 5 · 2 first-authorSystems, architecture and hardware · 3 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Performance modeling and evaluation · 64% Electronic design automation · 22% Processor architecture and microarchitecture · 11% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2008 | Distilling the essence of proprietary workloads into miniature benchmarks · ACM Trans. Archit. Code Optim. 2008 |
Performance modeling and evaluation › benchmarking › benchmark design
synthetic benchmark generation |
0.1 | 1 | 2008 | Distilling the essence of proprietary workloads into miniature benchmarks · ACM Trans. Archit. Code Optim. 2008 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2008 | Distilling the essence of proprietary workloads into miniature benchmarks · ACM Trans. Archit. Code Optim. 2008 |
Electronic design automation
design space exploration |
0.0 | 1 | 2004 | Control Flow Modeling in Statistical Simulation for Accurate and Efficient Processor Design Studies · ISCA 2004 |
Electronic design automation › circuit simulation › probabilistic simulation
statistical simulation |
0.0 | 1 | 2004 | Control Flow Modeling in Statistical Simulation for Accurate and Efficient Processor Design Studies · ISCA 2004 |
Processor architecture and microarchitecture
superscalar processor |
0.0 | 1 | 2004 | Control Flow Modeling in Statistical Simulation for Accurate and Efficient Processor Design Studies · ISCA 2004 |
Performance modeling and evaluation
simulation |
0.0 | 1 | 2008 | Distilling the essence of proprietary workloads into miniature benchmarks · ACM Trans. Archit. Code Optim. 2008 |
Energy-efficient computing
power-performance modeling |
0.0 | 1 | 2004 | Control Flow Modeling in Statistical Simulation for Accurate and Efficient Processor Design Studies · ISCA 2004 |
Methods — techniques the papers use, named apart from their topics
synthetic benchmark cloning · 0.1behavioral attribute distillation · 0.1trace generation · 0.0statistical simulation · 0.0branch predictor modeling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2011 | Automatic performance model synthesis from hardware verification modelsabstractPerformance models are typically written by hand for a new model or assembled piece-meal from the prior simulation code of an old model. In either case, many man-months of work may be required to write the new model and validate design details against a prior or current design. In reality, the majority of information about the performance of the design already exists in the design structure of either the old hardware model or the new model or both. To harvest this information and eliminate the significant duplicate coding and validation efforts, we propose that a performance model be automatically synthesized from a prior or current hardware design using a bottom-up, design-oriented approach. We demarcate the performance-critical boundaries of the design and perform backward-trace cone analysis to identify logic to include in the performance model. We then abstract specific components for design changes and expend modeling effort only on the few functions relevant to a particular design study. Engineering effort then becomes focused on workload selection and quality, defining and projecting new designs, and assessing design tradeoffs and sensitivities - the small set of tasks with the highest potential to improve design performance. Robert H. Bell Jr., Mátyás A. Sustik, David W. Cummings, Jonathan R. Jackson |
ICPE | 1 |
| 2011 | Power and energy-aware processor schedulingabstractPower consumption is a critical consideration in high computing systems. We propose a novel job scheduler that optimizes power and energy consumed by clusters when running parallel benchmarks with minimal impact on performance. We construct accurate models for estimating power consumption. These models are based on measurements of power consumption on benchmarks with different characteristics and on systems with processors using different micro-architectures. We show the power estimation models achieve less than 2% error versus actual measurements. We show a job scheduler can be enhanced to make it 'power-aware' and to optimize power consumption of jobs with similar performance characteristics. The enhanced scheduler can estimate the power consumed by a particular job using the power estimation model, configure the nodes in the cluster via suitably adjusting processor frequency on each of the nodes to maximize performance, minimize power, or minimize energy with a predictable impact on power, energy and performance. Luigi Brochard, Raj Panda, Don DeSota, Francois Thomas, Robert H. Bell Jr. |
ICPE | 5 |
| 2008 | Distilling the essence of proprietary workloads into miniature benchmarksabstractBenchmarks set standards for innovation in computer architecture research and industry product development. Consequently, it is of paramount importance that these workloads are representative of real-world applications. However, composing such representative workloads poses practical challenges to application analysis teams and benchmark developers (1) real-world workloads are intellectual property and vendors hesitate to share these proprietary applications; and (2) porting and reducing these applications to benchmarks that can be simulated in a tractable amount of time is a nontrivial task. In this paper, we address this problem by proposing a technique that automatically distills key inherent behavioral attributes of a proprietary workload and captures them into a miniature synthetic benchmark clone. The advantage of the benchmark clone is that it hides the functional meaning of the code but exhibits similar performance characteristics as the target application. Moreover, the dynamic instruction count of the synthetic benchmark clone is substantially shorter than the proprietary application, greatly reducing overall simulation time for SPEC CPU, the simulation time reduction is over five orders of magnitude compared to entire benchmark execution. Using a set of benchmarks representative of general-purpose, scientific, and embedded applications, we demonstrate that the power and performance characteristics of the synthetic benchmark clone correlate well with those of the original application across a wide range of microarchitecture configurations. Ajay Joshi, Lieven Eeckhout, Robert H. Bell Jr., Lizy Kurian John |
ACM Trans. Archit. Code Optim. | 3 |
| 2006 | Automatic testcase synthesis and performance model validation for high performance PowerPC processorsabstractThe latest high-performance IBM PowerPC microprocessor, the POWERS chip, poses challenges for performance model validation. The current state-of-the-art is to use simple hand-coded bandwidth and latency testcases, but these are not comprehensive for processors as complex as the POWER5 chip. Applications and benchmark suites such as SPEC CPU are difficult to set up or take too long to execute on functional models or even on detailed performance models. We present an automatic testcase synthesis methodology to address these concerns. By basing testcase synthesis on the workload characteristics of an application, source code is created that largely represents the performance of the application, but which executes in a fraction of the runtime. We synthesize representative PowerPC versions of the SPEC2000, STREAM, TPC-C and Java benchmarks, compile and execute them, and obtain an average IPC within 2.4% of the average IPC of the original benchmarks and with many similar average workload characteristics. The synthetic testcases often execute two orders of magnitude faster than the original applications, typically in less than 300K instructions, making performance model validation for today's complex processors feasible. Robert H. Bell Jr., Rajiv R. Bhatia, Lizy Kurian John, Jeffrey Stuecheli, John Griswell, Yiliu Tu, Louis Capps, Anton Blanchard, Ravel Thai |
ISPASS | 1 |
| 2006 | Evaluating the efficacy of statistical simulation for design space explorationabstractRecent research has proposed statistical simulation as a technique for fast performance evaluation of superscalar microprocessors. The idea in statistical simulation is to measure a program's key performance characteristics, generate a synthetic trace with these characteristics, and simulate the synthetic trace. Due to the probabilistic nature of statistical simulation the performance estimate quickly converges to a solution, making it an attractive technique to efficiently cull a large microprocessor design space. In this paper, we evaluate the efficacy of statistical simulation in exploring the design space. Specifically, we characterize the following aspects of statistical simulation: (i) fidelity of performance bottlenecks, with respect to cycle-accurate simulation of the program, (ii) ability' to track design changes, and (Hi) trade-off between accuracy and complexity in statistical simulation models. In our characterization experiments, we use the Plackett & Burman (P&B) design to systematically stress statistical simulation by creating different performance bottlenecks. The key results from this paper are: (1) Synthetic traces stress at least the same 10 most significant processor performance bottlenecks as the original workload, (2) Statistical simulation can effectively track design changes to identify feasible design points in a large design space of aggressive microarchitectures, (3) Our evaluation of 4 statistical simulation models shows that although a very detailed model is needed to achieve a good absolute accuracy in performance estimation, a simple model is sufficient to achieve good relative accuracy, and (4) The P&B design technique can be used to quickly identify areas to focus on to improve the accuracy of the statistical simulation model. Ajay Joshi, Joshua J. Yi, Robert H. Bell Jr., Lieven Eeckhout, Lizy Kurian John, David J. Lilja |
ISPASS | 3 |
| 2005 | Improved automatic testcase synthesis for performance model validationabstractPerformance simulation tools must be validated during the design process as functional models and early hardware are developed, so that designers can be sure of the performance of their designs as they implement changes. The current state-of-the-art is to use simple hand-coded bandwidth and latency testcases to assess early performance and to calibrate performance models. Applications and benchmark suites such as SPEC CPU are difficult to set up or take too long to execute on functional models. Short trace snippets from applications can be executed on performance and functional simulators, but not without difficulty on hardware, and there is no guarantee that hand-coded tests and short snippets cover the performance of the original applications.We present a new automatic testcase synthesis methodology to address these concerns. By basing testcase synthesis on the workload characteristics of an application, we create source code that largely represents the performance of the application, but which executes in a fraction of the runtime. We synthesize representative versions of the SPEC2000 benchmarks, compile and execute them, and obtain an average IPC within 2.4% of the average IPC of the original benchmarks with similar average workload characteristics. In addition, the changes in IPC due to design changes are found to be proportional to the changes in IPC for the original applications. The synthetic testcases execute more than three orders of magnitude faster than the original applications, typically in less than 300K instructions, making performance model validation feasible. Robert H. Bell Jr., Lizy Kurian John |
ICS | 1 |
| 2004 | Control Flow Modeling in Statistical Simulation for Accurate and Efficient Processor Design StudiesabstractDesigning a new microprocessor is extremely time-consuming. One of the contributing reasons is that computer designers rely heavily on detailed architectural simulations, which are very time-consuming. Recent work has focused on statistical simulation to address this issue. The basic idea of statistical simulation is to measure characteristics during program execution, generate a synthetic trace with those characteristics and then simulate the synthetic trace. The statistically generated synthetic trace is orders of magnitude smaller than the original program sequence and hence results in significantly faster simulation. This paper makes the following contributions to the statistical simulation methodology. First, we propose the use of a statistical flow graph to characterize the control flow of a program execution. Second, we model delayed update of branch predictors while profiling program execution characteristics. Experimental results show that statistical simulation using this improved control flow modeling attains significantly better accuracy than the previously proposed HLS system. We evaluate both the absolute and the relative accuracy of our approach for power/performance modeling of superscalar microarchitectures. The results show that our statistical simulation framework can be used to efficiently explore processor design spaces. Lieven Eeckhout, Robert H. Bell Jr., Bastiaan Stougie, Koen De Bosschere, Lizy Kurian John |
ISCA | 2 |