Wolfgang Karl

dblp:52/6717 · DBLP profile ↗
← Back
37ranked-venue papers
2as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 29 · 2 first-authorSecurity and privacy · 2Software engineering, systems software and programming languages · 2Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 49% High-performance computing · 31% Hardware accelerators and domain-specific architectures · 12%

Topics — the 17 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › accelerator orchestration
accelerator selection
0.112012
Seamlessly portable applications: Managing the diversity of modern heterogeneous systems · ACM Trans. Archit. Code Optim. 2012
High-performance computing
application portability
0.112012
Seamlessly portable applications: Managing the diversity of modern heterogeneous systems · ACM Trans. Archit. Code Optim. 2012
Parallel and multicore computing
parallel programming runtimes
0.112012
Seamlessly portable applications: Managing the diversity of modern heterogeneous systems · ACM Trans. Archit. Code Optim. 2012
High-performance computing
performance optimization at scale
0.112012
What scientific applications can benefit from hardware transactional memory? · SC 2012
Parallel and multicore computing
synchronization
0.112012
What scientific applications can benefit from hardware transactional memory? · SC 2012
Parallel and multicore computing
transactional memory
0.112012
What scientific applications can benefit from hardware transactional memory? · SC 2012
Parallel and multicore computing › parallel programming models › message passing
MPI applications
0.012012
Seamlessly portable applications: Managing the diversity of modern heterogeneous systems · ACM Trans. Archit. Code Optim. 2012
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP
0.012012
What scientific applications can benefit from hardware transactional memory? · SC 2012
Parallel and multicore computing
parallel programming models
0.012012
What scientific applications can benefit from hardware transactional memory? · SC 2012
Memory systems › cache
cache optimization
0.011999
Memory Characteristics of Iterative Methods · SC 1999
Memory systems › cache
cache performance
0.011999
Memory Characteristics of Iterative Methods · SC 1999
High-performance computing
iterative methods
0.011999
Memory Characteristics of Iterative Methods · SC 1999
Performance modeling and evaluation › workload characterization
memory system behavior
0.011999
Memory Characteristics of Iterative Methods · SC 1999
High-performance computing › numerical linear algebra › linear solver › iterative linear solvers
multigrid method
0.011999
Memory Characteristics of Iterative Methods · SC 1999
High-performance computing
scientific computing systems
0.011999
Memory Characteristics of Iterative Methods · SC 1999
Memory systems › cache
cache behavior
0.011999
Memory Characteristics of Iterative Methods · SC 1999
Performance modeling and evaluation
profiling
0.011999
Memory Characteristics of Iterative Methods · SC 1999

Methods — techniques the papers use, named apart from their topics

performance characterization · 0.1online-learning history-based selection · 0.1best practices · 0.1program transformation · 0.0profiling · 0.0cache optimization · 0.0
YearPublicationVenuePosition
2016 An Energy-Efficient Middleware for Computation Offloading in Real-Time Embedded Systems
abstract
Embedded systems have limited resources, such as computation capabilities and battery life. The Dynamic Voltage and Frequency Scaling (DVFS) technique is used to save energy by running the processor of the embedded system at low voltage and frequency levels. However, this prolongs the execution time, which may cause potential deadline misses for real-time tasks. In this paper, we propose a general-purpose middleware to reduce the energy consumption in embedded systems without violating the real-time constraints. The algorithms in the middleware adopt the computation offloading concept to reduce the workload on the processor of the embedded system by sending the computation-intensive tasks to a powerful server. The algorithms are further combined with the DVFS technique to find the running frequency (or speed) such that the energy consumption is minimized and the real-time constraints are satisfied. The evaluation shows that our approach reduces the average energy consumption down to nearly 60%, compared to executing all the tasks locally at the maximum processor speed.
Anas Toma, Santiago Pagani, Jian-Jia Chen, Wolfgang Karl, Jörg Henkel
RTCSA4
2015 Automatic task mapping and heterogeneity-aware fault tolerance: The benefits for runtime optimization and application development
Mario Kicherer, Wolfgang Karl
J. Syst. Archit.2
2015 Combined hardware-software multi-parallel prefiltering on the Convey HC-1 for fast homology detection
Michael Bromberger, Fabian Nowak, Wolfgang Karl
Parallel Comput.3
2013 Topic 4: High-Performance Architectures and Compilers - (Introduction)
Denis Barthou, Wolfgang Karl, Ramón Doallo, Evelyn Duesterwald, Sami Yehia
Euro-Par2
2013 Evaluation of Two Formulations of the Conjugate Gradients Method with Transactional Memory
Martin Schindewolf, Björn Rocker, Wolfgang Karl, Vincent Heuveline
Euro-Par3
2013 Multi-parallel prefiltering on the convey HC-1 for supporting homology detection
abstract
Gene databases used in research are huge and still grow at a fast pace. Many comparisons need to be done when searching similar (homologous) sequences in these databases for a given query sequence. Therefore, highly parallel architectures and much bandwidth are required for handling processing and transferring massive amounts of data. The Convey HC-1 with four FPGAs and high memory bandwidth of up to 76.8 GB/s seems very suitable for supporting this task as other bioinformatics applications have already been greatly supported by the HC-1. We research accelerating an application for searching homologous sequences. Limited by FPGA size only, we present a design that calculates 3 prefiltering scores per FPGA concurrently, i.e. 12 calculations in total. This score calculation for database sequences against the query profile is done by a modified Smith-Waterman scheme that is internally parallelized 16*8=128 times in contrast to the SSE implementation where only 16-fold parallelism can be exploited and where memory bandwidth poses the limiting factor. Preloading the query profile, we are able to transform the memory-bound SSE implementation to a compute-bound FPGA design which is only limited by FPGA size. Despite much lower clock rates, the FPGAs outperform SSE for the calculation of the prefiltering scores by a factor of 4.46. We achieve application speedup of 1.79 against the original, unmodified state-of-the-art SSE-based implementation because the score calculation accounts for less than 63% of the application runtime.
Fabian Nowak, Michael Bromberger, Martin Schindewolf, Wolfgang Karl
EuroMPI4
2012 A Scalable Monitoring Infrastructure for Self-Organizing Many-Core Architectures
abstract
Self-organizing principles can address the growing complexity and the huge challenge of management and efficient utilization of adaptive many-core architectures. Fundamental for realizing a self-organizing behavior within such architectures is a dedicated monitoring infrastructure that provides the essential information about the system status and system behavior for realizing the basic property of self-awareness. This paper therefore proposes a flexible, hierarchical and scalable monitoring infrastructure for self-organizing, adaptive many-core architectures. The employed basic monitoring unit in the bottom monitoring layer performs data aggregation and filtering and reduces the amount of data that must be processed in higher monitoring layers. The middle layer performs first data analysis and is further responsible for hiding the heterogeneity of the underlying hardware configuration to the topmost monitoring layer. The latter is finally responsible for detecting changes in the system behavior and realizing self-awareness. The proposed monitoring infrastructure was evaluated entirely using a simulation framework. Results show that the infrastructure is able of detecting changes in the system behavior of an entire many-core system causing only a minor system disturbance. Further, the prototypical implementation of the basic monitoring unit proved that it can be realized very efficiently in hardware.
David Kramer, Wolfgang Karl
DSD2
2012 A Low-Overhead Profiling and Visualization Framework for Hybrid Transactional Memory
abstract
Multi-core prototyping presents a good opportunity for establishing low overhead and detailed profiling and visualization in order to study new research topics. In this paper, we design and implement a low execution, low area overhead profiling mechanism and a visualization tool for observing Transactional Memory behaviors on FPGA. To achieve this, we non-disruptively create and bring out events on the fly and process them offline on a host. There, our tool regenerates the execution from the collected events and produces traces for comprehensively inspecting the behavior of interacting multithreaded programs. With zero execution overhead for hardware TM events, single-instruction overhead for software TM events, and utilizing a low logic area of 2.3% per processor core, we run TM benchmarks to evaluate various different levels of profiling detail with an average runtime overhead of 6%. We demonstrate the usefulness of such detailed examination of SW/HW transactional behavior in two parts: (i) we speed up a TM benchmark by 24.1%, and (ii) we closely inspect transactions to point out pathologies.
Oriol Arcas-Abella, Philipp Kirchhofer, Nehir Sönmez, Martin Schindewolf, Osman S. Unsal, Wolfgang Karl, Adrián Cristal
FCCM6
2012 What scientific applications can benefit from hardware transactional memory?
abstract
Achieving efficient and correct synchronization of multiple threads is a difficult and error-prone task at small scale and, as we march towards extreme scale computing, will be even more challenging when the resulting application is supposed to utilize millions of cores efficiently. Transactional Memory (TM) is a promising technique to ease the burden on the programmer, but only recently has become available on commercial hardware in the new Blue Gene/Q system and hence the real benefit for realistic applications has not been studied yet. This paper presents the first performance results of TM embedded into OpenMP on a prototype system of BG/Q and characterizes code properties that will likely lead to benefits when augmented with TM primitives. We first study the influence of thread count, environment variables and memory layout on TM performance and identify code properties that will yield performance gains with TM. Second, we evaluate the combination of OpenMP with multiple synchronization primitives on top of MPI to determine suitable task to thread ratios per node. Finally, we condense our findings into a set of best practices. These are applied to a Monte Carlo Benchmark and a Smoothed Particle Hydrodynamics method. In both cases an optimized TM version, executed with 64 threads on one node, outperforms a simple TM implementation. MCB with optimized TM yields a speedup of 27.45 over baseline.
Martin Schindewolf, Barna L. Bihari, John C. Gyllenhaal, Martin Schulz 0001, Amy Wang, Wolfgang Karl
SC6
2012 A survey on hardware-aware and heterogeneous computing on multicore processors and accelerators
abstract
SUMMARY In the last few years, the landscape of parallel computing has been subject to profound and highly dynamic changes. The paradigm shift towards multicore and manycore technologies coupled with accelerators in a heterogeneous environment is offering a great potential of computing power for scientific and industrial applications. However, for one to take full advantage of these new technologies, holistic approaches coupling the expertise ranging from hardware architecture and software design to numerical algorithms are a pressing necessity. Parallel computing is no longer limited to supercomputers and is now much more diversified – with a multitude of technologies, architectures, and programming approaches leading to increased complexity for developers and engineers. In this work, we give – from the perspective of numerical simulation and applications – an overview of existing and emerging multicore and manycore technologies as well as accelerator concepts. We emphasize the challenges associated with high‐performance heterogeneous computing and discuss the interfaces needed to fill the gap between the hardware architecture and the implementation of efficient numerical algorithms. By means of this short survey – which stresses the necessity of hardware‐aware computing – we aim at giving assistance to users in scientific computing entering this fascinating field and help understanding associated issues and capabilities. Copyright © 2011 John Wiley & Sons, Ltd.
Rainer Buchty, Vincent Heuveline, Wolfgang Karl, Jan-Philipp Weiss
Concurr. Comput. Pract. Exp.3
2012 Seamlessly portable applications: Managing the diversity of modern heterogeneous systems
abstract
Nowadays, many possible configurations of heterogeneous systems exist, posing several new challenges to application development: different types of processing units usually require individual programming models with dedicated runtime systems and accompanying libraries. If these are absent on an end-user system, e.g. because the respective hardware is not present, an application linked against these will break. This handicaps portability of applications being developed on one system and executed on other, differently configured heterogeneous systems. Moreover, the individual profit of different processing units is normally not known in advance. In this work, we propose a technique to effectively decouple applications from their accelerator-specific parts, respectively code. These parts are only linked on demand and thereby an application can be made portable across systems with different accelerators. As there are usually multiple hardware-specific implementations for a certain task, e.g., a CPU and a GPU version, a method is required to determine which are usable at all and which one is most suitable for execution on the current system. With our approach, application and hardware programmers can express the requirements and the abilities of the application and the hardware-specific implementations in a simplified manner. During runtime, the requirements and abilities are compared with regard to the present hardware in order to determine the usable implementations of a task. If multiple implementations are usable, an online-learning history-based selector is employed to determine the most efficient one. We show that our approach chooses the fastest usable implementation dynamically on several systems while introducing only a negligible overhead itself. Applied to an MPI application, our mechanism enables exploitation of local accelerators on different heterogeneous hosts without preliminary knowledge or modification of the application.
Mario Kicherer, Fabian Nowak, Rainer Buchty, Wolfgang Karl
ACM Trans. Archit. Code Optim.4
2011 Introduction
Wolfgang Karl, Samuel Thibault, Stanimire Tomov, Taisuke Boku
Euro-Par (2)1
2011 Cost-aware function migration in heterogeneous systems
abstract
Today's approaches towards heterogeneous computing rely on either the programmer or dedicated programming models to efficiently integrate heterogeneous components. In this work, we propose an adaptive cost-aware function-migration mechanism built on top of a light-weight hardware abstraction layer. With this mechanism, the highly dynamic task of choosing the most beneficial processing unit will be hidden from the programmer while causing only minor variation in the work and program flow. The migration mechanism transparently adapts to the current workload and system environment without the necessity of JIT compilation or binary translation.
Mario Kicherer, Rainer Buchty, Wolfgang Karl
HiPEAC3
2011 Digital On-demand Computing Organism - Interaction between Monitoring and Middleware
abstract
Organic Computing is a vital and promising research area. Inspired by nature, organic computing research wants to learn and adopt from techniques and properties of nature. The goal is to acquire the so called self-X properties like self-organization and self-healing. The DodOrg project introduces such an organic computing system for real-time applications, a whole new computing system from the bottom to the top. In this paper, we present the interaction between organic middleware and monitoring. Our results showed very promising results and only a small overhead for monitoring and the artificial hormone system based middleware.
Alexander von Renteln, Uwe Brinkschulte, David Kramer, Wolfgang Karl, Christian Schuck, Jürgen Becker 0001
ISORC4
2010 Cyberaide onServe: Software as a Service on Production Grids
abstract
The Software as a Service (SaaS) methodology is a key paradigm of Cloud computing. In this paper, we focus on an interesting topic - to implement a Cloud computing functionality, the SaaS model, on existing production Grid infrastructures. In general, production Grids employ a Job-Submission-Execution (JSE) model with rigid access interfaces. In this paper we develop the Cyberaide onServe, a lightweight middleware with a virtual appliance. The Cyberaide onServe implements the SaaS methodology on production Grids by translating the SaaS model to the JSE model. The Cyberaide onServe virtual appliance is deployed on demand, hosts applications as Web services, accepts Web service invocations, and finally the Cyberaide onServe executes them on production Grids. We have deployed the Cyberaide onServe on the TeraGrid infrastructure and test results show Cyberaide onServe can provide the SaaS functionality with good performance.
Tobias Kurze, Lizhe Wang 0001, Gregor von Laszewski, Jie Tao 0001, Marcel Kunze, David Kramer, Wolfgang Karl
ICPP7
2010 From source code to runtime behaviour: Software metrics help to select the computer architecture
Frank Eichinger, David Kramer, Klemens Böhm, Wolfgang Karl
Knowl. Based Syst.4
2009 Introduction
Pedro C. Diniz, Ben H. H. Juurlink, Alain Darte, Wolfgang Karl
Euro-Par4
2008 Scientific Cloud Computing: Early Definition and Experience
abstract
Cloud computing emerges as a new computing paradigm which aims to provide reliable, customized and QoS guaranteed computing dynamic environments for end-users. This paper reviews recent advances of Cloud computing, identifies the concepts and characters of scientific Clouds, and finally presents an example of scientific Cloud for data centers
Lizhe Wang 0001, Jie Tao 0001, Marcel Kunze, Alvaro Canales Castellanos, David Kramer, Wolfgang Karl
HPCC6
2008 Evaluating the Cache Architecture of Multicore Processors
abstract
Microprocessor architecture for both commercial and academical purpose is coming into a new generation: multiprocessors on a chip. Together with this novel architecture, questions and research topics also arise. For example, how to design the on-chip caches to avoid memory operations becoming the performance bottleneck? In this work, we study the impact of various cache architectures on the execution behavior of multi-threading applications. We focus on four general design issues: cache structure, configuration parameters, coherence influence, and prefetching strategies. The study is based on a self- developed cache simulator that models the functionality of a multicore cache hierarchy with arbitrary levels and various organizations. The achieved results can direct both hardware and program developers to optimize their cache designs or the program codes.
Jie Tao 0001, Marcel Kunze, Wolfgang Karl
PDP3
2007 A Profiling Tool for Detecting Cache-Critical Data Structures
Jie Tao 0001, Tobias Gaugler, Wolfgang Karl
Euro-Par3
2007 Optimizing Cache Performance of the Discrete Wavelet Transform Using a Visualization Tool
abstract
The 2D DWT consists of two 1D DWT in both directions: horizontal filtering processes the rows followed by vertical filtering processes the columns. It is well known that a straightforward implementation of the vertical filtering shows quite different performance with various working set sizes. The only reasonable explanation for this has to be the access behavior of the cache memory. As known, vertical filtering has mapping conflicts in the cache with a working set size that is power of two. However, it is not clear how this conflict forms and whether cache problems exist with other data sizes. Such knowledge is the base for efficient code optimization. In order to acquire this knowledge and to achieve more accurate optimization potentials, we apply a cache visualization tool to examine the runtime cache activities of the vertical implementation. We find that besides mapping conflicts, vertical filtering also shows a large number of capacity misses. More specifically, the visualization tool allows us to detect the parameters related to the strategies. This guarantees the feasibility of the optimization. Our initial experimental results on several different architectures show an up to 215% gain in execution time compared to an already optimized baseline implementation.
Jie Tao 0001, Asadollah Shahbahrami, Ben H. H. Juurlink, Rainer Buchty, Wolfgang Karl, Stamatis Vassiliadis
ISM5
2006 A network agent for diagnosis and analysis of real-time Ethernet networks
abstract
Within the field of automation technology the use of Industrial Ethernet is rising. This in turn demands devices capable of precisely recording, analyzing, and manipulating communication data for diagnostic purposes. Existing solutions so far lack required flexibility or are unable to cope with sustained Gigabit-per-second data streams. This is especially true for general-purpose approaches employing ordinary network adapters and plain software-based analysis.In this paper we describe a flexible and lightweight network agent for real-time, high-performance networks. This agent is capable of handling sustained data rates up to 2x 1GBit/s while offering real-time event-triggers, 10ns-resolution timestamps, real-time filtering, and statistics functions. An auxiliary processing unit as well as a modular software environment allow customization for a variety of tasks. The agent is realized as a dual processor SoC design on a Xilinx Virtex-II Pro FPGA.
Hans-Peter Löb, Rainer Buchty, Wolfgang Karl
CASES3
2006 Topic 7: Parallel Computer Architecture and Instruction Level Parallelism
Eduard Ayguadé, Wolfgang Karl, Koen De Bosschere, Jean-Francois Collard
Euro-Par2
2006 Supporting Cache Locality Optimization with a Toolset
Jie Tao 0001, Wolfgang Karl
Euro-Par2
2005 YACO: A User Conducted Visualization Tool for Supporting Cache Optimization
Boris Quaing, Jie Tao 0001, Wolfgang Karl
HPCC3
2005 Monitoring cache behavior on parallel SMP architectures and related programming tools
Thomas Brandes, Helmut Schwamborn, Michael Gerndt, Jürgen Jeitner, Edmond Kereku, Martin Schulz 0001, Holger Brunst, Wolfgang E. Nagel, Reinhard Neumann, Ralph Müller-Pfefferkorn, Bernd Trenkler, Wolfgang Karl, Jie Tao 0001, Hans-Christian Hoppe
Future Gener. Comput. Syst.12
2005 Simulation as a tool for optimizing memory accesses on NUMA machines
Jie Tao 0001, Martin Schulz 0001, Wolfgang Karl
Perform. Evaluation3
2004 Topic 8: Parallel Computer Architecture and Instruction-Level Parallelism
Kemal Ebcioglu, Wolfgang Karl, André Seznec, Marco Aldinucci
Euro-Par2
2004 Impact of Cache Coherence Models on Performance of OpenMP Applications
Jie Tao 0001, Wolfgang Karl
Euro-Par2
2003 SMiLE: an integrated, multi-paradigm software infrastructure for SCI-basedclusters
Martin Schulz 0001, Jie Tao 0001, Carsten Trinitis, Wolfgang Karl
Future Gener. Comput. Syst.4
2003 ARS: an adaptive runtime system for locality optimization
Jie Tao 0001, Martin Schulz 0001, Wolfgang Karl
Future Gener. Comput. Syst.3
2002 SMiLE: An Integrated, Multi-Paradigm Software Infrastructure for SCI-Based Clusters
abstract
The availability of a comprehensive software infrastructure is essential for the success a parallel architecture. In order to allow for the greatest possible flexibility, an infrastructure has to be designed in an integrated, easy-to-use manner and with the support of multiple programming paradigms and models to address a wide base of codes. SMiLE provides such an infrastructure for SCI (Scalable Coherent Interface) based clusters. It includes support for both a large range of message passing libraries as well as for almost arbitrary shared memory programming models. In addition, SMiLE also contains initial work on appropriate tool sets for performance optimizations. The complete infrastructure is implemented in way that is as closely relate d to the underlying hardware and is therefore capable of exploiting the benefits of the underlying network fabric and offering them to the user without significant overheads.
Martin Schulz 0001, Jie Tao 0001, Carsten Trinitis, Wolfgang Karl
CCGRID4
2001 OpenSESAME: An Intuitive Dependability Modeling Environment Supporting Inter-Component Dependencies
abstract
The paper proposes a novel modeling method for the evaluation of dependability measures of highly available systems. The proposed method, which has been implemented in the tool OpenSESAME (Simple but Extensive Structured Availability Modeling Environment), combines the advantages of Boolean methods and state space based methods. The tool supports the modeler with a set of well-defined, structured, intuitive input diagrams and tables, which are automatically transformed into GSPNs (Generalized Stochastic Petri Nets) for evaluation. To show the usefulness of the proposed method, it is applied to a model of a typical CompactPCI-based high availability system as can be found in the telecommunications area.
Max Walter, Carsten Trinitis, Wolfgang Karl
PRDC3
2000 NEPHEW: Applying a Toolset for the Efficient Deployment of a Medical Image Application on SCI-Based Clusters
Wolfgang Karl, Martin Schulz 0001, Martin Völk, Sibylle Ilse Ziegler
Euro-Par1
2000 Electrical phenomena during Hot Swap events
abstract
The exchange of a computer system's components during operation can be accomplished by the so called Hot Swap technology. This technology makes it possible to continuously run a computer system without the necessity of a shutdown for maintenance purposes, e.g. upgrading of a network adapter. Thus the overall uptime of a system can be drastically increased. The Hot Swap capability has been integrated into the so called CompactPCI technology. This paper summarizes the investigations that have been carried out with respect to the electrical behavior during Hot Swap events. Several simulations were performed with the H-SPICE program by AMP, Harrisburg, PA. After an introduction to live insertion phenomena in general, the CompactPCI specific simulations are described and recommendations for designing a Hot Swap capable system from an electrical point of view are given.
Carsten Trinitis, Wolfgang Karl, Markus Leberecht
PRDC2
1999 Memory Characteristics of Iterative Methods
abstract
Conventional implementations of iterative numerical algorithms, especially multigrid methods, merely reach a disappointing small percentage of the theoretically available CPU performance when applied to representative large problems.One of the most important reasons for this phenomenon is that the current DRAM technology cannot provide the data fast enough to keep the CPU busy.Although the fundamentals of cache optimizations are quite simple, current compilers cannot optimize even elementary iterative schemes.In this paper, we analyze the memory and cache behavior of iterative methods with extensive profiling and describe program transformation techniques to improve the cache performance of two-and three-dimensional multigrid algorithms.This project is partially funded by DFG Ru 422/7-1,2.1 All benchmarks in the article were compiled with native FORTRAN77 compilers and aggressive optimizations enabled.On the Intel platform we used egcs (V2.91.60).The platforms include an Intel PentiumII Xeon PC (450 MHz, 450 MFLOPS), a SUN Ultra 60 (296 MHz, 592 MFLOPS), a HP SPP2200 Convex Exemplar Node (200 MHz, 800 MFLOPS), a Compaq PWS 500au (500 MHz, 1 GFLOPS), and a Compaq XP1000 (500 MHz, 1 GFLOPS).1
Christian Weiß 0001, Wolfgang Karl, Markus Kowarschik, Ulrich Rüde
SC2
1998 Exploiting Spatial and Temporal Locality of Accesses: A New Hardware-Based Monitoring Approach for DSM Systems
Robert Hockauf, Wolfgang Karl, Markus Leberecht, Michael Oberhuber, Michael Wagner 0003
Euro-Par2