Wolf-Dietrich Weber

dblp:21/2538 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 4 first-authorSoftware engineering, systems software and programming languages · 7 · 4 first-authorDatabases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Energy-efficient computing · 40% Cloud and datacenter computing · 23% Storage systems · 22%

Topics — the 21 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing
energy proportionality
0.112011
Power management of online data-intensive services · ISCA 2011
Cloud and datacenter computing › datacenter services › online service systems › internet services
online data-intensive services
0.112011
Power management of online data-intensive services · ISCA 2011
Energy-efficient computing
power management
0.112011
Power management of online data-intensive services · ISCA 2011
Energy-efficient computing
datacenter power management
0.112007
Power provisioning for a warehouse-sized computer · ISCA 2007
Storage systems › storage reliability
disk failure
0.112007
Failure Trends in a Large Disk Drive Population · FAST 2007
Storage systems › magnetic storage
hard disk drive
0.112007
Failure Trends in a Large Disk Drive Population · FAST 2007
Energy-efficient computing › datacenter power management
power provisioning
0.112007
Power provisioning for a warehouse-sized computer · ISCA 2007
Storage systems
storage reliability
0.112007
Failure Trends in a Large Disk Drive Population · FAST 2007
Cloud and datacenter computing
datacenter workloads
0.012011
Power management of online data-intensive services · ISCA 2011
Cloud and datacenter computing
latency-critical applications
0.012011
Power management of online data-intensive services · ISCA 2011
Memory systems
cache coherence
0.041992
Cache Invalidation Patterns in Shared-Memory Multiprocessors · IEEE Trans. Computers 1992
Comparative Evaluation of Latency Reducing and Tolerating Techniques · ISCA 1991
Exploring the Benefits of Multiple Hardware Contexts in a Multiprocessor Architecture: Preliminary Results · ISCA 1989
Parallel and multicore computing › multiprocessor system
multiprocessor server
0.011997
The Mercury Interconnect Architecture: A Cost-effective Infrastructure for High-performance Servers · ISCA 1997
Processor architecture and microarchitecture
memory latency tolerance
0.021991
Comparative Evaluation of Latency Reducing and Tolerating Techniques · ISCA 1991
Exploring the Benefits of Multiple Hardware Contexts in a Multiprocessor Architecture: Preliminary Results · ISCA 1989
Memory systems › cache coherence
directory-based coherence
0.021989
Exploring the Benefits of Multiple Hardware Contexts in a Multiprocessor Architecture: Preliminary Results · ISCA 1989
Analysis of Cache Invalidation Patterns in Multiprocessors · ASPLOS 1989
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor
0.021992
Cache Invalidation Patterns in Shared-Memory Multiprocessors · IEEE Trans. Computers 1992
Analysis of Cache Invalidation Patterns in Multiprocessors · ASPLOS 1989
Performance modeling and evaluation › simulation › parallel architecture simulation
multiprocessor simulation
0.011992
Cache Invalidation Patterns in Shared-Memory Multiprocessors · IEEE Trans. Computers 1992
Performance modeling and evaluation › parallel system performance
multiprocessor performance evaluation
0.011991
Comparative Evaluation of Latency Reducing and Tolerating Techniques · ISCA 1991
Performance modeling and evaluation › simulation
simulation-based evaluation
0.011991
Comparative Evaluation of Latency Reducing and Tolerating Techniques · ISCA 1991
Processor architecture and microarchitecture › multithreading
context switching
0.011989
Exploring the Benefits of Multiple Hardware Contexts in a Multiprocessor Architecture: Preliminary Results · ISCA 1989
Processor architecture and microarchitecture
multithreading
0.011989
Exploring the Benefits of Multiple Hardware Contexts in a Multiprocessor Architecture: Preliminary Results · ISCA 1989
Processor architecture and microarchitecture
multiprocessor architecture
0.011989
Exploring the Benefits of Multiple Hardware Contexts in a Multiprocessor Architecture: Preliminary Results · ISCA 1989

Methods — techniques the papers use, named apart from their topics

workload characterization · 0.1power provisioning strategy · 0.1field data analysis · 0.1RAS features · 0.0multiprocessor simulation · 0.0simulation · 0.0trace-driven simulation · 0.0trace analysis · 0.0classification scheme · 0.0
YearPublicationVenuePosition
2011 Power management of online data-intensive services
abstract
Much of the success of the Internet services model can be attributed to the popularity of a class of workloads that we call Online Data-Intensive (OLDI) services. These workloads perform significant computing over massive data sets per user request but, unlike their offline counterparts (such as MapReduce computations), they require responsiveness in the sub-second time scale at high request rates. Large search products, online advertising, and machine translation are examples of workloads in this class. Although the load in OLDI services can vary widely during the day, their energy consumption sees little variance due to the lack of energy proportionality of the underlying machinery. The scale and latency sensitivity of OLDI workloads also make them a challenging target for power management techniques.
David Meisner, Christopher M. Sadler, Luiz André Barroso, Wolf-Dietrich Weber, Thomas F. Wenisch
ISCA4
2007 Failure Trends in a Large Disk Drive Population
Eduardo Pinheiro, Wolf-Dietrich Weber, Luiz André Barroso
FAST2
2007 Power provisioning for a warehouse-sized computer
abstract
Large-scale Internet services require a computing infrastructure that can beappropriately described as a warehouse-sized computing system. The cost ofbuilding datacenter facilities capable of delivering a given power capacity tosuch a computer can rival the recurring energy consumption costs themselves.Therefore, there are strong economic incentives to operate facilities as closeas possible to maximum capacity, so that the non-recurring facility costs canbe best amortized. That is difficult to achieve in practice because ofuncertainties in equipment power ratings and because power consumption tends tovary significantly with the actual computing activity. Effective powerprovisioning strategies are needed to determine how much computing equipmentcan be safely and efficiently hosted within a given power budget.
Xiaobo Fan, Wolf-Dietrich Weber, Luiz André Barroso
ISCA2
2005 A Quality-of-Service Mechanism for Interconnection Networks in System-on-Chips
abstract
As Moore's Law continues to fuel the ability to build ever increasing complex systems-on-chips (SoCs), achieving performance goals is rising as a critical challenge to completing designs. In particular, the system interconnect must efficiently service a diverse set of data flows with widely ranging quality-of-service (QoS) requirements. However the known solutions for off-chip interconnects, such as large-scale networks, are not necessarily applicable to the on-chip environment. Latency and memory constraints for on-chip interconnects are quite different from larger-scale interconnects. The paper introduces a novel on-chip interconnect arbitration scheme. We show how this scheme can be distributed across a chip for high-speed implementation. We compare the performance of the arbitration scheme with other known interconnect arbitration schemes. Existing schemes typically focus heavily on either low latency of service for some initiators or on guaranteed bandwidth delivery for other initiators. Our scheme allows service latency on some initiators to be traded off smoothly against jitter bounds on other initiators, while still delivering bandwidth guarantees. This scheme is a subset of the QoS controls that are available in the SonicsMX/spl trade/ (SMX) product.
Wolf-Dietrich Weber, Joe Chou, Ian Swarbrick, Drew Wingard
DATE1
1997 The Mercury Interconnect Architecture: A Cost-effective Infrastructure for High-performance Servers
abstract
This paper presents HAL's Mercury Interconnect Architecture, an interconnect infrastructure designed to link commodity microprocessors, memory, and I/O components into high-performance multiprocessing servers. Both shared-memory and message-passing systems, as well as hybrid systems are supported by the interconnect. The key attributes of the Mercury Interconnect Architecture are: low latency, high bandwidth, a modular and flexible design, reliability/availability/serviceability (RAS) features, and a simplicity that enables very cost-effective implementations. The first implementation of the architecture links multiple 4-processor Pentium™ Pro based nodes. In a 4-node (16-processor) shared-memory configuration, this system achieves a remote read latency of just over 1 µs, and a maximum interconnect bandwidth of 6.4 GByte/s. Both of these parameters far outpace comparable SCI-based solutions, while utilizing much fewer hardware components.
Wolf-Dietrich Weber, Stephen Gold, Pat Helland, Takeshi Shimizu, Thomas Wicki, Winfried W. Wilcke
ISCA1
1992 Cache Invalidation Patterns in Shared-Memory Multiprocessors
abstract
The cache invalidation patterns of several parallel applications are analyzed. The results are based on multiprocessor simulations with 8, 16, and 32 processors. To provide deeper insight into the observed invalidation behavior the invalidations observed in the simulations are linked to the high-level objects causing them in the programs. To predict what the invalidation patterns would look like beyond 32 processors, a classification scheme for data objects found in parallel programs is proposed. The classification scheme provides a powerful conceptual tool to reason about the invalidation patterns of parallel applications. Results indicate that it should be possible to scale well-written parallel programs to a large number of processors without an explosion in invalidation traffic. At the same time, the invalidation patterns are such that directory-based schemes with just a few pointers per entry can be very effective. The variations in invalidation behavior with different cache line sizes are discussed. The results indicate that cache line sizes in the 32-byte range yield the lowest data and invalidation traffic.>
Anoop Gupta, Wolf-Dietrich Weber
IEEE Trans. Computers2
1991 Comparative Evaluation of Latency Reducing and Tolerating Techniques
abstract
Techniques that can cope with the large latency of memory accesses are essential for achieving high processor utilization in large-scale shared-memory multiprocessors. In this paper, we consider four architectural techniques that address the latency problem: (i) hardware coherent caches, (ii) relaxed memory consistency, (iii) softwarecontrolled prefetching, and (iv) multiple-context support. While some studies of benefits of the individual techniques have been done, no study evaluates all of the techniques within a consistent framework. This paper attempts to remedy this by providing a comprehensive evaluation of the benefits of the four techniques, both individually and in combinations, using a consistent set of architectural assumptions. The results in this paper have been obtained using detailed simulations of a large-scale shared-memory multiprocessor. Our results show that caches and relaxed consistency uniformly improve performance. The improvements due to prefetching and multiple contexts are sizeable, but are much more applicationdependent. Combinations of the various techniques generally attain better performance than each one on its own. Overall, we show that using suitable combinations of the techniques, performance can be improved by 4 to 7 times.
Anoop Gupta, John L. Hennessy, Kourosh Gharachorloo, Todd C. Mowry, Wolf-Dietrich Weber
ISCA5
1990 Reducing Memory and Traffic Requirements for Scalable Directory-Based Cache Coherence Schemes
Anoop Gupta, Wolf-Dietrich Weber, Todd C. Mowry
ICPP (1)2
1989 Analysis of Cache Invalidation Patterns in Multiprocessors
abstract
To make shared-memory multiprocessors scalable, researchers are now exploring cache coherence protocols that do not rely on broadcast, but instead send invalidation messages to individual caches that contain stale data. The feasibility of such directory-based protocols is highly sensitive to the cache invalidation patterns that parallel programs exhibit. In this paper, we analyze the cache invalidation patterns caused by several parallel applications and investigate the effect of these patterns on a directory-based protocol. Our results are based on multiprocessor traces with 4, 8 and 16 processors. To gain insight into what the invalidation patterns would look like beyond 16 processors, we propose a classification scheme for data objects found in parallel applications and link the invalidation traffic patterns observed in the traces back to these high-level objects. Our results show that synchronization objects have very different invalidation patterns from those of other data objects. A write reference to a synchronization object usually causes invalidations in many more caches. We point out situations where restructuring the application seems appropriate to reduce the invalidation traffic, and others where hardware support is more appropriate. Our results also show that it should be possible to scale “well-written” parallel programs to a large number of processors without an explosion in invalidation traffic.
Wolf-Dietrich Weber, Anoop Gupta
ASPLOS1
1989 Exploring the Benefits of Multiple Hardware Contexts in a Multiprocessor Architecture: Preliminary Results
abstract
A fundamental problem that any scalable multiprocessor must address is the ability to tolerate high latency memory operations. This paper explores the extent to which multiple hardware contexts per processor can help to mitigate the negative effects of high latency. In particular, we evaluate the performance of a directory-based cache coherent multiprocessor using memory reference traces obtained from three parallel applications. We explore the case where there are a small fixed number (2-4) of hardware contexts per processor and the context switch overhead is low. In contrast to previously proposed approaches, we also use a very simple context switch criterion, namely a cache miss or a write-hit to shared data. Our results show that the effectiveness of multiple contexts depends on the nature of the applications, the context switch overhead, and the inherent latency of the machine architecture. Given reasonably low overhead hardware context switches, we show that two or four contexts can achieve substantial performance gains over a single context. For one application, the processor utilization increased by about 46% with two contexts and by about 80% with four contexts.
Wolf-Dietrich Weber, Anoop Gupta
ISCA1