Richard Uhlig

dblp:03/3431 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
0since 2021 · last 2000
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-authorSoftware engineering, systems software and programming languages · 7 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Memory systems · 52% Processor architecture and microarchitecture · 20% Performance modeling and evaluation · 17%
Software engineering, system software, and programming languages
4 papers
Operating systems · 100%

Topics — the 19 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
thread-level parallelism
0.012000
Thread Level Parallelism and Interactive Performance of Desktop Applications · ASPLOS 2000
Memory systems › memory management › virtual memory › address translation
TLB
0.021994
Design Tradeoffs for Software-Managed TLBs · ACM Trans. Comput. Syst. 1994
Design Tradeoffs for Software-Managed TLBs · ISCA 1993
Processor architecture and microarchitecture
branch prediction
0.011997
Trading Conflict and Capacity Aliasing in Conditional Branch Predictors · ISCA 1997
Performance modeling and evaluation
simulation
0.021994
Trap-driven Simulation with Tapeworm II · ASPLOS 1994
Kernel-Based Memory Simulation · SIGMETRICS 1994
Operating systems › resource management › memory management
virtual memory
0.021994
Design Tradeoffs for Software-Managed TLBs · ISCA 1993
Design Tradeoffs for Software-Managed TLBs · ACM Trans. Comput. Syst. 1994
Processor architecture and microarchitecture
instruction fetch
0.011995
Instruction Fetching: Coping with Code Bloat · ISCA 1995
Memory systems › cache
prefetching
0.011995
Instruction Fetching: Coping with Code Bloat · ISCA 1995
Memory systems › memory hierarchy
cache and TLB effects
0.011994
Trap-driven Simulation with Tapeworm II · ASPLOS 1994
Memory systems
memory referencing behavior
0.011994
Trap-driven Simulation with Tapeworm II · ASPLOS 1994
Performance modeling and evaluation › simulation › architectural simulation
memory system simulation
0.011994
Kernel-Based Memory Simulation · SIGMETRICS 1994
Memory systems › memory management › on-chip memory management
on-chip memory allocation
0.011994
Optimal Allocation of On-Chip Memory for Multiple-API Operating Systems · ISCA 1994
Memory systems › virtual memory management
software TLB
0.011994
Design Tradeoffs for Software-Managed TLBs · ACM Trans. Comput. Syst. 1994
Memory systems › memory management
virtual memory
0.011994
Design Tradeoffs for Software-Managed TLBs · ACM Trans. Comput. Syst. 1994
Memory systems › memory architecture
interleaved memory
0.011991
Using Lookahead to reduce memory bank contention for decoupled operand references · SC 1991
Memory systems › memory interference
memory bank conflicts
0.011991
Using Lookahead to reduce memory bank contention for decoupled operand references · SC 1991
Memory systems
cache
0.011997
Trading Conflict and Capacity Aliasing in Conditional Branch Predictors · ISCA 1997
Performance modeling and evaluation
benchmarking
0.011995
Instruction Fetching: Coping with Code Bloat · ISCA 1995
Memory systems › cache
cache organization
0.011995
Instruction Fetching: Coping with Code Bloat · ISCA 1995
Operating systems
kernel instrumentation
0.011994
Trap-driven Simulation with Tapeworm II · ASPLOS 1994

Methods — techniques the papers use, named apart from their topics

simulation · 0.1hardware monitoring · 0.1trace-driven simulation · 0.1trap-driven simulation · 0.0kernel-based analysis · 0.0analytical modeling · 0.0kernel-based simulation · 0.0lookahead control · 0.0
YearPublicationVenuePosition
2000 Thread Level Parallelism and Interactive Performance of Desktop Applications
Krisztián Flautner, Richard Uhlig, Steven K. Reinhardt, Trevor N. Mudge
ASPLOS2
1997 Trading Conflict and Capacity Aliasing in Conditional Branch Predictors
abstract
As modern microprocessors employ deeper pipelines and issue multiple instructions per cycle, they are becoming increasingly dependent on accurate branch prediction. Because hardware resources for branch-predictor tables are invariably limited, it is not possible to hold all relevant branch history for all active branches at the same time, especially for large workloads consisting of multiple processes and operating-system code. The problem that results, commonly referred to as aliasing in the branch-predictor tables, is in many ways similar to the misses that occur in finite-sized hardware caches.In this paper we propose a new classification for branch aliasing based on the three-Cs model for caches, and show that conflict aliasing is a significant source of mispredictions. Unfortunately, the obvious method for removing conflicts --- adding tags and associativity to the predictor tables --- is not a cost-effective solution.To address this problem, we propose the skewed branch predictor, a multi-bank, tag-less branch predictor, designed specifically to reduce the impact of conflict aliasing. Through both analytical and simulation models, we show that the skewed branch predictor removes a substantial portion of conflict aliasing by introducing redundancy to the branch-predictor tables. Although this redundancy increases capacity aliasing compared to a standard one-bank structure of comparable size, our simulations show that the reduction in conflict aliasing overcomes this effect to yield a gain in prediction accuracy. Alternatively, we show that a skewed organization can achieve the same prediction accuracy as a standard one-bank organization but with half the storage requirements.
Pierre Michaud, André Seznec, Richard Uhlig
ISCA3
1995 Instruction Fetching: Coping with Code Bloat
abstract
Previous research has shown that the SPEC benchmarks achieve low miss ratios in relatively small instruction caches. This paper presents evidence that current software-development practices produce applications that exhibit substantially higher instruction-cache miss ratios than do the SPEC benchmarks. To represent these trends, we have assembled a collection of applications, called the Instruction Benchmark Suite (IBS), that provides a better test of instruction-cache performance. We discuss the rationale behind the design of IBS and characterize its behavior relative to the SPEC benchmark suite. Our analysis is based on trace-driven and trap-driven simulations and takes into full account both the application and operating-system components of the workloads.This paper then reexamines a collection of previously-proposed hardware mechanisms for improving instruction-fetch performance in the context of the IBS workloads. We study the impact of cache organization, transfer bandwidth, prefetching, and pipelined memory systems on machines that rely on the use of relatively small primary instruction caches to facilitate increased clock rates. We find that, although of little use for SPEC, the right combination of these techniques substantially benefits IBS. Even so, under IBS, a stubborn lower bound on the instruction-fetch CPI remains as an obstacle to improving overall processor performance.
Richard Uhlig, David Nagle, Trevor N. Mudge, Stuart Sechrest, Joel S. Emer
ISCA1
1994 Trap-driven Simulation with Tapeworm II
abstract
Tapeworm II is a software-based simulation tool that evaluates the cache and TLB performance of multiple-task and operating system intensive workloads. Tapeworm resides in an OS kernel and causes a host machine's hardware to drive simulations with kernel traps instead of with address traces, as is conventionally done. This allows Tapeworm to quickly and accurately capture complete memory referencing behavior with a limited degradation in overall system performance. This paper compares trap-driven simulation, as implemented in Tapeworm, with the more common technique of trace-driven memory simulation with respect to speed, accuracy, portability and flexibility.
Richard Uhlig, David Nagle, Trevor N. Mudge, Stuart Sechrest
ASPLOS1
1994 Optimal Allocation of On-Chip Memory for Multiple-API Operating Systems
abstract
The allocation of die area to different processor components is a central issue in the design of single-chip microprocessors. Chip area is occupied by both core execution logic, such as ALU and FPU datapaths, and memory structures, such as caches, TLBs, and write buffers. The authors focus on the allocation of die area to memory structures through a cost/benefit analysis. The cost of memory structures with different sizes and associativities is estimated by using an established area model for on-chip memory. The performance benefits of selecting a given structure are measured through a collection of methods including on-the-fly hardware monitoring, trace-driven simulation and kernel-based analysis. Special consideration is given to operating systems that support multiple application programming interfaces (APIs), a software trend that substantially affects on-chip memory allocation decisions. Results: Small adjustments in cache and TLB design parameters can significantly impact overall performance. Operating systems that support multiple APIs, such as Mach 3.0, increase the relative importance of on-chip instruction caches and TLBs when compared against single-API systems such as Ultrix.>
David Nagle, Richard Uhlig, Trevor N. Mudge, Stuart Sechrest
ISCA2
1994 Kernel-Based Memory Simulation
abstract
No abstract available.
Richard Uhlig, David Nagle, Trevor N. Mudge, Stuart Sechrest
SIGMETRICS1
1994 Design Tradeoffs for Software-Managed TLBs
abstract
An increasing number of architectures provide virtual memory support through software-managed TLBs. However, software management can impose considerable penalties that are highly dependent on the operating system's structure and its use of virtual memory. This work explores software-managed TLB design tradeoffs and their interaction with a range of monolithic and microkernel operating systems. Through hardware monitoring and simulation, we explore TLB performance for benchmarks running on a MIPS R2000-based workstation running Ultrix, OSF/1, and three versions of Mach 3.0.
Richard Uhlig, David Nagle, Tim J. Stanley, Trevor N. Mudge, Stuart Sechrest, Richard B. Brown
ACM Trans. Comput. Syst.1
1993 Design Tradeoffs for Software-Managed TLBs
abstract
An increasing number of architectures provide virtual memory support through software-managed TLBs. However, software management can impose considerable penalties, which are highly dependent on the operating system's structure and its use of virtual memory. This work explores software-managed TLB design tradeoffs and their interaction with a range of operating systems including monolithic and microkernel designs. Through hardware monitoring and simulations, we explore TLB performance for benchmarks running on a MIPS R2000-based workstation running Ultrix, OSF/1, and three versions of mach 3.0.
David Nagle, Richard Uhlig, Tim J. Stanley, Stuart Sechrest, Trevor N. Mudge, Richard B. Brown
ISCA2
1991 Using Lookahead to reduce memory bank contention for decoupled operand references
abstract
We introduce a storage system control structure which employs lookahead to minimize the incidence of bank collisions for operand references in memory systems with interleaved banks.We show that by employing lookahead a memory system achieves high data throughput with only a small interleave factor.5'imulation results show that an 8 bank system with lookahead performs as well, or better, than a 32 bank system without Iookahead.A Iookahead controller simplifies the design of bank management hardware, thus reducing the complexity of high throughput memory designs.
Peter L. Bird, Richard Uhlig
SC2