Jean-Loup Baer

dblp:b/JLBaer · DBLP profile ↗
← Back
52ranked-venue papers
19as first author
0since 2021 · last 2000
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 44 · 12 first-authorSoftware engineering, systems software and programming languages · 21 · 7 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
35 papers
Memory systems · 55% Performance modeling and evaluation · 19% Processor architecture and microarchitecture · 15%
Software engineering, system software, and programming languages
3 papers
Runtime systems and virtual machines · 96% Compilers and program optimization · 4%

Topics — the 30 heaviest of 85, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache coherence
0.191998
Optimizing Software Cache-coherent Cluster Architectures · SC 1998
A Performance Evaluation of Cluster-Based Architectures · SIGMETRICS 1997
On the Use and Performance of Explicit Communication Primitives in Cache-Coherent Multiprocessor Systems · HPCA 1997
Memory systems
cache
0.182000
Modified LRU Policies for Improving Second-Level Cache Behavior · HPCA 2000
Instruction Cache Fetch Policies for Speculative Execution · ISCA 1995
A Performance Study of Software and Hardware Data Prefetching Schemes · ISCA 1994
Performance modeling and evaluation
workload characterization
0.131999
On the Use of Trace Sampling for Architectural Studies of Desktop Applications · SIGMETRICS 1999
Execution Characteristics of Desktop Applications on Windows NT · ISCA 1998
The Structure and Performance of Interpreters · ASPLOS 1996
Memory systems › cache
prefetching
0.031995
Effective Hardware Based Data Prefetching for High-Performance Processors · IEEE Trans. Computers 1995
Reducing Memory Latency via Non-blocking and Prefetching Caches · ASPLOS 1992
An effective on-chip preloading scheme to reduce data access penalty · SC 1991
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor
0.041998
On the Use and Performance of Explicit Communication Primitives in Cache-Coherent Multiprocessor Systems · HPCA 1997
Optimizing Software Cache-coherent Cluster Architectures · SC 1998
A Performance Study of Software and Hardware Data Prefetching Schemes · ISCA 1994
Memory systems › cache › prefetching
hardware prefetching
0.031994
A Performance Study of Software and Hardware Data Prefetching Schemes · ISCA 1994
Reducing Memory Latency via Non-blocking and Prefetching Caches · ASPLOS 1992
An effective on-chip preloading scheme to reduce data access penalty · SC 1991
Memory systems › cache management
cache replacement
0.022000
Modified LRU Policies for Improving Second-Level Cache Behavior · HPCA 2000
Dynamic Improvement of Locality in Virtual Memory Systems · IEEE Trans. Software Eng. 1976
Performance modeling and evaluation
trace sampling
0.011999
On the Use of Trace Sampling for Architectural Studies of Desktop Applications · SIGMETRICS 1999
Processor architecture and microarchitecture
multicore design
0.021998
Optimizing Software Cache-coherent Cluster Architectures · SC 1998
Cache Coherence Protocols: Evaluation Using a Multiprocessor Simulation Model · ACM Trans. Comput. Syst. 1986
Processor architecture and microarchitecture › branch handling
branch architecture
0.011998
Execution Characteristics of Desktop Applications on Windows NT · ISCA 1998
Performance modeling and evaluation
analytical modeling
0.011997
A Performance Evaluation of Cluster-Based Architectures · SIGMETRICS 1997
Processor architecture and microarchitecture
chip multiprocessor
0.011997
A Performance Evaluation of Cluster-Based Architectures · SIGMETRICS 1997
Distributed systems › operating system support › interprocess communication
communication primitives
0.011997
On the Use and Performance of Explicit Communication Primitives in Cache-Coherent Multiprocessor Systems · HPCA 1997
Performance modeling and evaluation › queueing models
mean value analysis
0.011997
A Performance Evaluation of Cluster-Based Architectures · SIGMETRICS 1997
Memory systems › cache coherence
software cache coherence
0.011997
On the Use and Performance of Explicit Communication Primitives in Cache-Coherent Multiprocessor Systems · HPCA 1997
Memory systems › cache
cache behavior
0.031991
Efficient Trace-Driven Simulation Methods for Cache Performance Analysis · ACM Trans. Comput. Syst. 1991
Efficient Trace-Driven Simulation Methods for Cache Performance Analysis · SIGMETRICS 1990
Improving Quicksort Performance with a Codewort Data Structure · IEEE Trans. Software Eng. 1989
Runtime systems and virtual machines
interpreter
0.011996
The Structure and Performance of Interpreters · ASPLOS 1996
Runtime systems and virtual machines › interpreter
interpreter performance
0.011996
The Structure and Performance of Interpreters · ASPLOS 1996
Processor architecture and microarchitecture
branch prediction
0.011995
Instruction Cache Fetch Policies for Speculative Execution · ISCA 1995
Memory systems
cache design
0.011995
Two Techniques for Improving Performance on Bus-Based Multiprocessors · HPCA 1995
Memory systems › cache › CPU cache
instruction cache
0.011995
Instruction Cache Fetch Policies for Speculative Execution · ISCA 1995
Memory systems › cache › cache organization
sectored cache
0.011995
Two Techniques for Improving Performance on Bus-Based Multiprocessors · HPCA 1995
Memory systems › cache coherence › cache coherence protocol
snoopy coherence
0.011995
Two Techniques for Improving Performance on Bus-Based Multiprocessors · HPCA 1995
Processor architecture and microarchitecture
speculative execution
0.011995
Instruction Cache Fetch Policies for Speculative Execution · ISCA 1995
Memory systems › cache › prefetching
data prefetching
0.011994
A Performance Study of Software and Hardware Data Prefetching Schemes · ISCA 1994
Memory systems
software prefetching
0.011994
A Performance Study of Software and Hardware Data Prefetching Schemes · ISCA 1994
Memory systems › memory hierarchy
cache hierarchy
0.021989
Organization and Performance of a Two-Level Virtual-Real Cache Hierarchy · ISCA 1989
On the Inclusion Properties for Multi-Level Cache Hierarchies · ISCA 1988
Processor architecture and microarchitecture
latency hiding
0.011992
Reducing Memory Latency via Non-blocking and Prefetching Caches · ASPLOS 1992
Memory systems › memory consistency
memory consistency model
0.011992
A Performance Study of Memory Consistency Models · ISCA 1992
Memory systems › cache design
non-blocking cache
0.011992
Reducing Memory Latency via Non-blocking and Prefetching Caches · ASPLOS 1992

Methods — techniques the papers use, named apart from their topics

simulation · 0.1trace-driven simulation · 0.1performance measurement · 0.0temporal locality detection · 0.0profile-based trace analysis · 0.0online locality table · 0.0trace sampling · 0.0measurement · 0.0message passing · 0.0PRAM model · 0.0benchmarking · 0.0analysis · 0.0proof methodology · 0.0formal specification · 0.0petri net modeling · 0.0learning algorithm · 0.0graph reduction · 0.0branch-and-bound · 0.0
YearPublicationVenuePosition
2000 Modified LRU Policies for Improving Second-Level Cache Behavior
abstract
Main memory accesses continue to be a significant bottleneck for applications whose working sets do not fit in second-level caches. With the trend of greater associativity in second-level caches, implementing effective replacement algorithms might become more important than reducing conflict misses. After showing that an opportunity exists to close part of the gap between the OPT and the LRU algorithms, we present a replacement algorithm based on the detection of temporal locality in lines residing in the L2 cache. Rather than always replacing the LRU line, the victim is chosen by considering both its priority in the LRU stack and whether it exhibits temporal locality or not. We consider two strategies which use this replacement algorithm: a profile-based scheme where temporal locality is detected by processing a trace from a training set of the application and an on-line scheme, where temporal locality is detected with the assistance of a small locality table. Both schemes improve on the second-level cache miss rate over a pure LRU algorithm, by as much as 12% in the profiling case and 20% in the dynamic case.
Wayne A. Wong, Jean-Loup Baer
HPCA2
2000 Characterizing processor architectures for programmable network interfaces
abstract
The rapid advancements of networking technology have boosted potential bandwidth to the point that the cabling is no longer the bottleneck. Rather, the bottlenecks lie at the crossing points, the nodes of the network, where data traffic is intercepted or forwarded. As a result, there has been tremendous interest in speeding those nodes, making the equipment run faster by means of specialized chips to handle data trafficking. The Network Processor is the blanket name thrown over such chips in their varied forms. To date, no performance data exist to aid in the decision of what processor architecture to use in next generation network processor. Our goal is to remedy this situation. In this study, we characterize both the application workloads that network processors need to support as well as emerging applications that we anticipate may be supported in the future. Then, we consider the performance of three sample benchmarks drawn from these workloads on several state-of-the-art processor architectures, including: an aggressive, out-of-order, speculative super-scalar processor, a fine-grained multithreaded processor, a single chip multiprocessor, and a simultaneous multithreaded processor (SMT). The network interface environment is simulated in detail, and our results indicate that SMT is the architecture best suited to this environment.
Patrick Crowley, Marc E. Fiuczynski, Jean-Loup Baer, Brian N. Bershad
ICS3
1999 Pursuing the Performance Potential of Dynamic Cache Line Sizes
abstract
We examine the application of offline algorithms for determining the optical sequence of loads and superloads (a load of multiple consecutive cache lines) for direct-mapped caches. We evaluate potential gains in terms of miss rate and bandwidth and find that in many cases optimal superloading can noticeably reduce the miss rate without appreciably increasing bandwidth. Then we examine how this performance potential might be realized. We examine the effectiveness of a dynamic online algorithm and of static analysis (profiling) for superloading and compare these to next-line prefetching. Experimental results show improvements comparable to those of the optimal algorithm in terms of miss rates.
Peter van Vleet, Eric J. Anderson, Lindsay Brown, Jean-Loup Baer, Anna R. Karlin
ICCD4
1999 On the Use of Trace Sampling for Architectural Studies of Desktop Applications
abstract
No abstract available.
Patrick Crowley, Jean-Loup Baer
SIGMETRICS2
1998 Execution Characteristics of Desktop Applications on Windows NT
abstract
This paper examines the performance of desktop applications running on the Microsoft Windows NT operating system on Intel x86 processors, and contrasts these applications to the programs in the integer SPEC95 benchmark suite. We present measurements of basic instruction set and program characteristics, and detailed simulation results of the way these programs use the memory system and processor branch architecture. We show that the desktop applications have similar characteristics to the integer SPEC95 benchmarks for many of these metrics, However compared to the integer SPEC95 applications, desktop applications have larger instruction working sets, execute instructions in a greater number of unique functions, cross DLL boundaries frequently, and execute a greater number of indirect calls.
Dennis C. Lee, Patrick Crowley, Jean-Loup Baer, Thomas E. Anderson, Brian N. Bershad
ISCA3
1998 Optimizing Software Cache-coherent Cluster Architectures
abstract
Software cache-coherent systems using programmable protocol processors provide a flexible infrastructure to expand the systems in size and function. However this flexibility comes at a cost in performance. First, the software implementation of protocols is inherently slower than a hardware implementation. Second, when multiple processors share a protocol processor, contention may result in a substantial increase in memory latency. In this paper, we study how the overhead of a software scheme can be reduced in the context of a shared- memory system consisting of SMP clusters. We study various design choices including hardware assists such as forwarding logic in the protocol processor and software hints through explicit communication primitives. We conduct our experiments via trace-driven simulation and compare the execution of three programs from the SPLASH-2 suite. We found that small cluster sizes (up to 4 processors/node) work well for both hardware and software implementations. When the forwarding logic is incorporated with the software scheme, the performance is competitive to that of the hardware scheme. When enhanced further by explicit communication primitives, the software scheme can perform even better than a pure hardware implementation. This is particularly noticeable when the network latency is high.
Xiaohan Qin, Jean-Loup Baer
SC2
1997 On the Use and Performance of Explicit Communication Primitives in Cache-Coherent Multiprocessor Systems
abstract
Recent developments in shared-memory multiprocessor systems advocate using off-the-shelf hardware to provide basic communication mechanisms and using software to implement cache coherence policies. The exposure of communication mechanisms to software opens many opportunities for enhancing application performance. In this paper we propose a set of communication primitives implemented on a communication co-processor that introduce a flavor of message passing and permit protocol optimization. To assess the overhead of the software implementation of the primitives and protocols, we compare a PRAM model, a hardware cache coherence scheme, a software scheme implementing only the basic cache coherence protocol, and an optimized software solution supporting the additional communication primitives and running with applications annotated with those primitives. With the parameters we chose for the communication processor, the overall memory system overhead of the basic software scheme is at least 50% higher than that of the hardware implementation. With the adequate insertion of the communication primitives, the optimized software solution has a performance comparable to that of the hardware scheme.
Xiaohan Qin, Jean-Loup Baer
HPCA2
1997 A Performance Evaluation of Cluster-Based Architectures
abstract
This paper investigates the performance of shared-memory cluster-based architectures where each cluster is a shared-bus multiprocessor augmented with a protocol processor maintaining cache coherence across clusters. For a given number of processors, sixteen in this study, we evaluate the performance of various cluster configurations. We also consider the impact of adding a remote shared cache in each cluster. We use Mean Value Analysis to estimate the cache miss latencies of various types and the overall execution time. The service demands of shared resources are characterized in detail by examining the sub-requests issued in resolving cache misses. In addition to the architectural system parameters and the service demands on resources, the analytical model needs parameters pertinent to applications. The latter, in particular cache miss profiles, are obtained by trace-driven simulation of three benchmarks.Our results show that without remote caches the performance of cluster-based architectures is mixed. In some configurations, the negative effects of the longer latency of inter-cluster misses and of the contention on the protocol processor are too large to counter-balance the lower contention on the data buses. For two out of the three applications best results are obtained when the system has clusters of size 2 or 4. The cluster-based architectures with remote caches consistently outperform the single bus system for all 3 applications. We also exercise the model with parameters reflecting the current trend in technology making the processor relatively faster than the bus and memory. Under these new conditions, our results show a clear performance advantage for the cluster-based architectures, with or without remote caches, over single bus systems.
Xiaohan Qin, Jean-Loup Baer
SIGMETRICS2
1996 The Structure and Performance of Interpreters
abstract
Interpreted languages have become increasingly popular due to demands for rapid program development, ease of use, portability, and safety. Beyond the general impression that they are "slow," however, little has been documented about the performance of interpreters as a class of applications.This paper examines interpreter performance by measuring and analyzing interpreters from both software and hardware perspectives. As examples, we measure the MIPSI, Java, Perl, and Tcl interpreters running an array of micro and macro benchmarks on a DEC Alpha platform. Our measurements of these interpreters relate performance to the complexity of the interpreter's virtual machine and demonstrate that native runtime libraries can play a key role in providing good performance. From an architectural perspective, we show that interpreter performance is primarily a function of the interpreter itself and is relatively independent of the application being interpreted. We also demonstrate that high-level interpreters' demands on processor resources are comparable to those of other complex compiled programs, such as gcc. We conclude that interpreters, as a class of applications, do not currently motivate special hardware support for increased performance.
Theodore H. Romer, Dennis Lee 0001, Geoffrey M. Voelker, Alec Wolman, Wayne A. Wong, Jean-Loup Baer, Brian N. Bershad, Henry M. Levy
ASPLOS6
1995 Two Techniques for Improving Performance on Bus-Based Multiprocessors
abstract
We explore two techniques for reducing memory latency in bus-based multiprocessors. The first one, designed for sector caches, is a snoopy cache coherence protocol that uses a large transfer block to take advantage of spatial locality, while using a small coherence block (called a subblock to avoid false sharing). The second technique is read snarfing (or read broadcasting), in which all caches can acquire data transmitted in response to a read request to update invalid blocks in their own cache. We evaluated the two techniques by simulating 6 applications that exhibit a variety of reference patterns. We compared the performance of the new protocol against that of the Illinois protocol with both small and large block sizes and found that it was effective in reducing memory latency and providing more consistent, good results than the Illinois protocol with a given line size. Read snarfing also improved performance mostly for protocols that use large line sizes.>
Craig Anderson 0001, Jean-Loup Baer
HPCA2
1995 Instruction Cache Fetch Policies for Speculative Execution
abstract
Current trends in processor design are pointing to deeper and wider pipelines and superscalar architectures. The efficient use of these resources requires speculative execution, a technique whereby the processor continues executing the predicted path of a branch before the branch condition is resolved.In this paper, we investigate the implications of speculative execution on instruction cache performance. We explore policies for managing instruction cache misses ranging from aggressive policies (always fetch on the speculative path) to conservative ones (wait until branches are resolved). We test these policies and their interaction with next-line prefetching by simulating the effects on instruction caches with varying architectural parameters. Our results suggest that an aggressive policy combined with next-line prefetching is best for small latencies while more conservative policies are preferable for large latencies.
Dennis Lee 0001, Jean-Loup Baer, Brad Calder, Dirk Grunwald
ISCA2
1995 Two techniques for improving performance on bus-based multiprocessors
Craig Anderson 0001, Jean-Loup Baer
Future Gener. Comput. Syst.2
1995 Effective Hardware Based Data Prefetching for High-Performance Processors
abstract
Memory latency and bandwidth are progressing at a much slower pace than processor performance. In this paper, we describe and evaluate the performance of three variations of a hardware function unit whose goal is to assist a data cache in prefetching data accesses so that memory latency is hidden as often as possible. The basic idea of the prefetching scheme is to keep track of data access patterns in a reference prediction table (RPT) organized as an instruction cache. The three designs differ mostly on the timing of the prefetching. In the simplest scheme (basic), prefetches can be generated one iteration ahead of actual use. The lookahead variation takes advantage of a lookahead program counter that ideally stays one memory latency time ahead of the real program counter and that is used as the control mechanism to generate the prefetches. Finally the correlated scheme uses a more sophisticated design to detect patterns across loop levels. These designs are evaluated by simulating the ten SPEC benchmarks on a cycle-by-cycle basis. The results show that 1) the three hardware prefetching schemes all yield significant reductions in the data access penalty when compared with regular caches, 2) the benefits are greater when the hardware assist augments small on-chip caches, and 3) the lookahead scheme is the preferred one cost-performance wise.>
Tien-Fu Chen, Jean-Loup Baer
IEEE Trans. Computers2
1994 A Parallel Trace-driven Simulator: Implementation and Performance
abstract
The simulation of parallel architectures requires an enormous amount of CPU cycles and, in the case of trace-driven simulation, of disk storage. In this paper, we consider the evaluation of the memory hierarchy of multiprocessor systems via parallel trace-driven simulation. We refine Lin et al.[8] original algorithm, whose main characteristic is to insert the shared references from every trace in all other traces, by reducing the amount of communication between simulation processes. We have implemented our algorithm on a KSR-1. Results of our experiments on traces of four applications and three different cache coherence protocols show that parallel trace-driven simulation yields significant speedups over its sequential counter-part. The communication overhead is not substantial compared to the dominant overhead due to the processing of replicated inserted references. We also investigate filtering techniques and show how to filter in parallel private and shared references for various block sizes in one pass. Simulation of filtered traces is faster but with a lower speedup.
Xiaohan Qin, Jean-Loup Baer
ICPP (2)2
1994 A Performance Study of Software and Hardware Data Prefetching Schemes
abstract
Prefetching, i.e., exploiting the overlap of processor computations with data accesses, is one of several approaches for tolerating memory latencies. Prefetching can be either hardware-based or software-directed or a combination of both. Hardware-based prefetching, requiring some support unit connected to the cache, can dynamically handle prefetches at run-time without compiler intervention. Software-directed approaches rely on compiler technology to insert explicit prefetch instructions. Mowry et al.'s software scheme (1991,1992) and the authors' hardware approach (1991) are two representative schemes. In this paper, the authors evaluate approximations to these two schemes in the context of a shared-memory multiprocessor environment. Their qualitative comparisons indicate that both schemes are able to reduce cache misses in the domain of linear array references. When complex data access patterns are considered, the software approach has compile-time information to perform sophisticated prefetching whereas the hardware scheme has the advantage of manipulating dynamic information. The performance results from an instruction-level simulation of four benchmarks confirm these observations. Simulations show that the hardware scheme introduces more memory traffic into the network and that the software scheme introduces a non-negligible instruction execution overhead. An approach combining software and hardware schemes is proposed; it shows promise in reducing the memory latency with least overhead.>
Tien-Fu Chen, Jean-Loup Baer
ISCA2
1992 Reducing Memory Latency via Non-blocking and Prefetching Caches
abstract
Non-blocking caches and prefetehing caches are two techniques for hiding memory latency by exploiting the overlap of processor computations with data accesses.A nonblocking cache allows execution to proceed concurrently with cache misses as long as dependency constraints are observed, thus exploiting post-miss operations, A prefetching cache generates prefetch requests to bring data in the cache before it is actually needed, thus allowing overlap with premiss computations.In this paper, we evaluate the effectiveness of these two hardware-based schemes.We propose a hybrid design based on the combination of these approaches.We also consider compiler-based optimization to enhance the effectiveness of non-blocking caches.Results from instruction level simulations on the SPEC benchmarks show that the hardware prefetching caches generally outperform nonblocking caches.Also, the relative effectiveness of nonblocklng caches is more adversely affected by an increase in memory latency than that of prefetching caches,, However, the performance of non-blocking caches can be improved substantially by compiler optimizations such as instruction scheduling and register renaming.The hybrid design cm be very effective in reducing the memory latency penalty for many applications.
Tien-Fu Chen, Jean-Loup Baer
ASPLOS2
1992 A Performance Study of Memory Consistency Models
abstract
Recent advances in technology are such that the speed of processors is increasing faster than memory latency is decreasing. Therefore the relative cost of a cache miss is becoming more important. However, the full cost of a cache miss need not be paid every time in a multiprocessor. The frequency with which the processor must stall on a cache miss can be reduced by using a relaxed model of memory consistency.
Richard N. Zucker, Jean-Loup Baer
ISCA2
1992 Design and Analysis of a Scalable Cache Coherence Scheme Based on Clocks and Timestamps
abstract
A timestamp-based software-assisted cache coherence scheme that does not require any global communication to enforce the coherence of multiple private caches is proposed. It is intended for shared memory multiprocessors. The scheme is based on a compile-time marking of references and a hardware-based local incoherence detection scheme. The possible incoherence of a cache entry is detected and the associated entry is implicitly invalidated by comparing a clock (related to program flow) and a timestamp (related to the time of update in the cache). Results of a performance comparison, which is based on a trace-driven simulation using actual traces. between the proposed timestamp-based scheme and other software-assisted schemes indicate that the proposed scheme performs significantly better than previous software-assisted schemes, especially when the processors are carefully scheduled so as to maximize the reuse of cache contents. This scheme requires neither a shared resource nor global communication and is, therefore, scalable up to a large number of processors.>
Sang Lyul Min, Jean-Loup Baer
IEEE Trans. Parallel Distributed Syst.2
1991 On Synchronization Patterns in Parallel Programs
Jean-Loup Baer, Richard N. Zucker
ICPP (2)1
1991 An effective on-chip preloading scheme to reduce data access penalty
abstract
Conventional cache prefetching approaches can be either hardware-based, generally by using a one-block-Iookahead technique, or compiler-directed, with insertions of non-blocking prefetch instructions.We introduce a new hardware scheme based on the prediction of the execution of the instruction stream and associated operand references.It consists of a reference prediction table and a look-ahead program counter and its associated logic.With this scheme, data with regular access patterns is preloaded, independently of the stride size, and preloading of data with irregular access patterns is prevented.We evaluate our design through trace driven simulation by comparing it with a pure data cache approach under three different memory access models.Our experiments show that this scheme is very effective for reducing the data access penalty for scientific programs and that is has moderate success for other applications.
Jean-Loup Baer, Tien-Fu Chen
SC1
1991 Efficient Trace-Driven Simulation Methods for Cache Performance Analysis
abstract
We propose improvements to current trace-driven cache simulation methods to make them faster and more economicalWe attack the large time and space demands of cache simulation in two ways.First, we reduce the program traces to the extent that exact performance can still be obtained from the reduced traces.Second, we devise an algorithm that can produce performance results for a variety of metrics (hit ratio, write-back counts, bus traffic) for a large number of set-associative write-back caches in just a single simulation run, The trace reduction and the efficient simulation techniques are extended to parallel multiprocessor cache simulations, Our simulation results show that our approach substantially reduces the disk space needed to store the program traces and can dramatically speed up cache simulations and still produce the exact results.
Wen-Hann Wang, Jean-Loup Baer
ACM Trans. Comput. Syst.2
1990 A Performance Comparison of Directory-based and Timestamp-based Cache Coherence Schemes
Sang Lyul Min, Jean-Loup Baer
ICPP (1)2
1990 An efficient caching support for critical sections in large-scale shared-memory multiprocessors
Sang Lyul Min, Jean-Loup Baer
ICS2
1990 Efficient Trace-Driven Simulation Methods for Cache Performance Analysis
abstract
We propose improvements to current trace-driven cache simulation methods to make them faster and more economical. We attack the large time and space demands of cache simulation in two ways. First, we reduce the program traces to the extent that exact performance can still be obtained from the reduced traces. Second, we devise an algorithm that can produce performance results for a variety of metrics (hit ratio, write-back counts, bus traffic) for a large number of set-associative write-back caches in just a single simulation run. The trace reduction and the efficient simulation techniques are extended to parallel multiprocessor cache simulations. Our simulation results show that our approach substantially reduces the disk space needed to store the program traces and can dramatically speedup cache simulations and still produce the exact results.
Wen-Hann Wang, Jean-Loup Baer
SIGMETRICS2
1989 A Timestamp-based Cache Coherence Scheme
Sang Lyul Min, Jean-Loup Baer
ICPP (1)2
1989 Extending the Memory Hierarchy into Multiprocessor Interconnection Networks: A Performance Analysis
Haim E. Mizrahi, Jean-Loup Baer, Edward D. Lazowska, John Zahorjan
ICPP (1)2
1989 Introducing Memory into Switch Elements of Multiprocessor Interconnection Networks
abstract
As VLSI technology continues to improve, circuit area is gradually being replaced by pin restrictions as the limiting factor in design. Thus, it is reasonable to anticipate that on-chip memory will become increasingly inexpensive since it is a simple, regular structure than can easily take advantage of higher densities.
Haim E. Mizrahi, Jean-Loup Baer, Edward D. Lazowska, John Zahorjan
ISCA2
1989 Organization and Performance of a Two-Level Virtual-Real Cache Hierarchy
abstract
We propose and analyze a two-level cache organization that provides high memory bandwidth. The first-level cache is accessed directly by virtual addresses. It is small, fast, and, without the burden of address translation, can easily be optimized to match the processor speed. The virtually-addressed cache is backed up by a large physically-addressed cache; this second-level cache provides a high hit ratio and greatly reduces memory traffic. We show how the second-level cache can be easily extended to solve the synonym problem resulting from the use of a virtually-addressed cache at the first level. Moreover, the second-level cache can be used to shield the virtually-addressed first-level cache from irrelevant cache coherence interference. Finally, simulation results show that this organization has a performance advantage over a hierarchy of physically-addressed caches in a multiprocessor environment.
Wen-Hann Wang, Jean-Loup Baer, Henry M. Levy
ISCA2
1989 Multilevel Cache Hierarchies: Organizations, Protocols, and Performance
Jean-Loup Baer, Wen-Hann Wang
J. Parallel Distributed Comput.1
1989 Improving Quicksort Performance with a Codewort Data Structure
abstract
The problem is discussed of how the use of a new data structure, the codeword structure, can help improve the performance of quicksort when the records to be sorted are long and the keys are alphanumeric sequences of bytes. The codeword is a compact representation of a key with respect to some codeword generator. It consists of a byte for a character count of equal bytes, a byte for the first nonequal byte, and a pointer to the record. It is shown how the ordering of keys is preserved by an adequate choice of the code generator and how this can be applied to the quicksort algorithm. An analysis of the potential saving son various architectures and actual measurements shows the improvements that can be attained by using codewords rather than pointers. Architecturally independent parameters, such as the number of bytes to be compared, the number of swaps, architecture-dependent parameters such as caches and their write policies, and compiler optimizations such as in-line expansion and register allocation are considered.>
Jean-Loup Baer, Yi-Bing Lin
IEEE Trans. Software Eng.1
1988 A Notation for Describing Multiple Views of VLSI Circuits
Jean-Loup Baer, Meei-Chiueh Liem, Larry McMurchie, Rudolf Nottrott, Lawrence Snyder 0001, Wayne Winder
DAC1
1988 On the Inclusion Properties for Multi-Level Cache Hierarchies
abstract
The inclusion property is essential in reducing the cache coherence complexity for multiprocessors with multilevel cache hierarchies. Some necessary and sufficient conditions for imposing the inclusion property for fully-associative and set-associative caches, which allow different block sizes at different levels of the hierarchy, are given. Three multiprocessor structures with a two-level cache hierarchy (single cache extension, multiport second-level cache, and bus-based) are examined. The feasibility of imposing the inclusion property in these structures is discussed. This leads to the presentation of an inclusion-coherence mechanism for two-level bus-based architectures.>
Jean-Loup Baer, Wen-Hann Wang
ISCA1
1987 Architectural Choices for Multi-level Cache Hierarchies
Jean-Loup Baer, Wen-Hann Wang
ICPP1
1986 Cache Coherence Protocols: Evaluation Using a Multiprocessor Simulation Model
abstract
Using simulation, we examine the efficiency of several distributed, hardware-based solutions to the cache coherence problem in shared-bus multiprocessors. For each of the approaches, the associated protocol is outlined. The simulation model is described, and results from that model are presented. The magnitude of the potential performance difference between the various approaches indicates that the choice of coherence solution is very important in the design of an efficient shared-bus multiprocessor, since it may limit the number of processors in the system.
James K. Archibald, Jean-Loup Baer
ACM Trans. Comput. Syst.2
1985 Parallel Tag-Distribution Sort
Sai Choi Kwan, Jean-Loup Baer, G. Zick, T. Snyder
ICPP2
1985 The I/O Performance of Multiway Mergesort and Tag Sort
abstract
We develop models of secondary storage to evaluate external sorting and use them to analyze the average I/O access time of mergesort and tag sort on files with uniform key distribution. The k-way mergesort takes [logkR] merge passes to sort a file with R initial sorted runs. Choosing k as large as possible reduces the number of merge passes, but we show that under the assumptions of our models, the I/O access time of the merge phase in mergesort increases as a function of k. For large files with short keys, tag sort provides a promising alternative to mergesort. We analyze the I/O access time of tag sort with a ``sequential scan'' distribution method. We show that for large files tag sort takes asymptotically less I/O time than mergesort.
Sai Choi Kwan, Jean-Loup Baer
IEEE Trans. Computers2
1984 An Economical Solution to the Cache Coherence Problem
abstract
In this paper we review and qualitatively evaluate schemes to maintain cache coherence in tightly-coupled multiprocessor systems. This leads us to propose a more economical (hardware-wise), expandable and modular variation of the “global directory” approach. Protocols for this solution are described. Performance evaluation studies indicate the limits (number of processors, level of sharing) within which this approach is viable.
James K. Archibald, Jean-Loup Baer
ISCA2
1983 On the Performance of Interleaved Memories with Non-Uniform Access Probabilities
David Hung-Chang Du, Jean-Loup Baer
ICPP2
1983 Binary Search in a Multiprocessing Environment
abstract
In this paper we consider variations on the binary search algorithm when placed in the context of a multiprocessing environment. Several organizations are investigated covering the spectrum from total independence (or free competition for access to common resources) to cooperation as in SIMD architectures. It is assumed that the two main sources of overhead are memory interference and interprocessor synchronization. An organization combining interference-free access to memory by implicit synchronization and a small degree of cooperation yields the best results.
Jean-Loup Baer, David Hung-Chang Du, Richard E. Ladner
IEEE Trans. Computers1
1981 The Two-Step Commitment Protocol: Modeling, Specification and Proof Methodology
Jean-Loup Baer, Georges Gardarin, Claude Girault, Gérard Roucairol
ICSE1
1979 On the Minimization of the Width of the Control Memory of Microprogammed Processors
abstract
A branch and bound method to minimize the width of the control memory of microprogrammed processors is given. Although it is exponential in the worst case, it appears much more effective than previous enumerative solutions. Furthermore, it can lead quickly to near-optimal solutions representing "good engineering" reductions.
Jean-Loup Baer, Barbara Koyama
IEEE Trans. Computers1
1978 Software control and program design issues for alterable architectures
abstract
A classification of alterable architectures is given. Problem areas in the software control and program design for these architectures are delineated. This is illustrated by considering an application area, namely searching.
Jean-Loup Baer
COMPSAC1
1977 Simulation of Large Parallel Systems: Modelling of Tasks
Jean-Loup Baer, John E. Jensen
Performance1
1977 Model, Design, and Evaluation of a Compiler for a Parallel Processing Environment
abstract
The problem of designing compilers for a multiprocessing environment is approached. We show that by modeling an existing sequential compiler, we gain an understanding of the modifications necessary to transform the sequential structure into a pipeline of processes. The pipelined compiler is then evaluated through measurements and simulation. Properties of the model, a generalized Petri Net, are also discussed.
Jean-Loup Baer, Carla Schlatter Ellis
IEEE Trans. Software Eng.1
1976 A Model of Interference in a Shared Resource Multiprocessor
abstract
This paper presents a generalized model of tightly-coupled multiprocessor systems which is then simplified to form a stochastic model for the study of interference. Analysis is performed on the resource contention which is characteristic of such systems in order to find a measure of system performance. After reviewing the problem of memory interference, the analysis is extended to contention in other individual resources, then combined to form a model for the interacting effects of contention in systems where processors contend for several shared resources.
John E. Jensen, Jean-Loup Baer
ISCA2
1976 The Formal Definition of Semantics by String Automata
G. Kampen, Jean-Loup Baer
Comput. Lang.2
1976 Multiprocessing Systems
abstract
This paper surveys the state of the art in the design and evaluation of multiprocessing systems. Multiprocessor architectures of the SIMD and MIMD type are reviewed and further classified depending on their tight or loose coupling and their homogeneity. The additional complexity of the software for control, synchronization, efficient utilization, and performance monitoring of multiple processors is emphasized.
Jean-Loup Baer
IEEE Trans. Computers1
1976 Dynamic Improvement of Locality in Virtual Memory Systems
abstract
Replacement algorithms for virtual memory systems are typically based on temporal measures of locality, while predictive loading and program restructuring are based on spatial measures of locality. This paper suggests some techniques for dynamically improving the spatial locality of a program via predictive loading and virtual space restructuring, and presents the results of applying these techniques to actual programs. Bounds are derived for the performance of the methods.
Jean-Loup Baer, Gary R. Sager
IEEE Trans. Software Eng.1
1976 Correction to "Dynamic Improvement of Locality in Virtual Memory Systems"
Jean-Loup Baer, Gary R. Sager
IEEE Trans. Software Eng.1
1974 On Program Placement in a Directly Executable Hierarchy of Memories
abstract
The efficient utilization of a two-level directly executable memory system is investigated. After defining the time and space product resulting from static allocation of the most often referenced pages, from paging, and from an optimal algorithm when the amount of primary memory is constrained, we introduce a learning algorithm. Its basic feature is to monitor references in such a way that it prevents seldom accessed pages to be brought into primary memory. The additional hardware requirements are not extensive. Simulations attest to the validity of the concept, and show that results are comparable with those obtained from the static allocation (the latter being impractical since it requires the knowledge of the whole reference stream) and superior to those obtained with paging. In the case of application programs, contributions to the learning algorithm can be made at compile time. Algorithms and data stuctures necessitated in an optimizing phase of the compiler are described.
Jean-Loup Baer
IEEE Trans. Computers1
1970 Legality and Other Properties of Graph Models of Computations
abstract
Directed graphs having logical control associated with each vertex have been introduced as models of computational tasks for automatic assignment and sequencing on parallel processor systems. A brief review of their properties is given. A procedure to test the “legality” of graphs in this class is described, and leads to algorithms for counting the number of all possible executions (AND-type subgraphs), and for evaluating the probability of ever reaching a given vertex in the graph. Numerical results are given for some example graphs.
Jean-Loup Baer, Daniel P. Bovet, Gerald Estrin
J. ACM1
1969 Bounds for Maxium Parallelism in a Bilogic Graph Model of Computations
abstract
Given an acyclic directed graph where vertices represent computational tasks, arcs represent transfer of control, and two labels—called input and output logics—associated with each vertex show either the concurrency or the mutual exclusiveness of tasks, procedures are given to determine a lower and an upper bound on the number of processors required for maximum parallelism. The lower bound is obtained via a mean path length approach, while the upper bound is based on the structure of the graph. A detailed algorithm is given for the latter. First, some reduction rules are applied yielding a subset of the vertices which can be performed in parallel. Then the maximum cut in the graph is determined taking into account mutually exclusive vertices. Results are given for example graphs.
Jean-Loup Baer, Gerald Estrin
IEEE Trans. Computers1