EDBT 2026 Demo / reviewers in the wild / expert
Wolfgang Karl
dblp:52/6717
· DBLP profile ↗
37ranked-venue papers
2as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 2 first-authorSecurity and privacy · 2Software engineering, systems software and programming languages · 2Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Parallel and multicore computing · 49% High-performance computing · 31% Hardware accelerators and domain-specific architectures · 12% |
Topics — the 17 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › accelerator orchestration
accelerator selection |
0.1 | 1 | 2012 | Seamlessly portable applications: Managing the diversity of modern heterogeneous systems · ACM Trans. Archit. Code Optim. 2012 |
High-performance computing
application portability |
0.1 | 1 | 2012 | Seamlessly portable applications: Managing the diversity of modern heterogeneous systems · ACM Trans. Archit. Code Optim. 2012 |
Parallel and multicore computing
parallel programming runtimes |
0.1 | 1 | 2012 | Seamlessly portable applications: Managing the diversity of modern heterogeneous systems · ACM Trans. Archit. Code Optim. 2012 |
High-performance computing
performance optimization at scale |
0.1 | 1 | 2012 | What scientific applications can benefit from hardware transactional memory? · SC 2012 |
Parallel and multicore computing
synchronization |
0.1 | 1 | 2012 | What scientific applications can benefit from hardware transactional memory? · SC 2012 |
Parallel and multicore computing
transactional memory |
0.1 | 1 | 2012 | What scientific applications can benefit from hardware transactional memory? · SC 2012 |
Parallel and multicore computing › parallel programming models › message passing
MPI applications |
0.0 | 1 | 2012 | Seamlessly portable applications: Managing the diversity of modern heterogeneous systems · ACM Trans. Archit. Code Optim. 2012 |
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP |
0.0 | 1 | 2012 | What scientific applications can benefit from hardware transactional memory? · SC 2012 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2012 | What scientific applications can benefit from hardware transactional memory? · SC 2012 |
Memory systems › cache
cache optimization |
0.0 | 1 | 1999 | Memory Characteristics of Iterative Methods · SC 1999 |
Memory systems › cache
cache performance |
0.0 | 1 | 1999 | Memory Characteristics of Iterative Methods · SC 1999 |
High-performance computing
iterative methods |
0.0 | 1 | 1999 | Memory Characteristics of Iterative Methods · SC 1999 |
Performance modeling and evaluation › workload characterization
memory system behavior |
0.0 | 1 | 1999 | Memory Characteristics of Iterative Methods · SC 1999 |
High-performance computing › numerical linear algebra › linear solver › iterative linear solvers
multigrid method |
0.0 | 1 | 1999 | Memory Characteristics of Iterative Methods · SC 1999 |
High-performance computing
scientific computing systems |
0.0 | 1 | 1999 | Memory Characteristics of Iterative Methods · SC 1999 |
Memory systems › cache
cache behavior |
0.0 | 1 | 1999 | Memory Characteristics of Iterative Methods · SC 1999 |
Performance modeling and evaluation
profiling |
0.0 | 1 | 1999 | Memory Characteristics of Iterative Methods · SC 1999 |
Methods — techniques the papers use, named apart from their topics
performance characterization · 0.1online-learning history-based selection · 0.1best practices · 0.1program transformation · 0.0profiling · 0.0cache optimization · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | An Energy-Efficient Middleware for Computation Offloading in Real-Time Embedded SystemsabstractEmbedded systems have limited resources, such as computation capabilities and battery life. The Dynamic Voltage and Frequency Scaling (DVFS) technique is used to save energy by running the processor of the embedded system at low voltage and frequency levels. However, this prolongs the execution time, which may cause potential deadline misses for real-time tasks. In this paper, we propose a general-purpose middleware to reduce the energy consumption in embedded systems without violating the real-time constraints. The algorithms in the middleware adopt the computation offloading concept to reduce the workload on the processor of the embedded system by sending the computation-intensive tasks to a powerful server. The algorithms are further combined with the DVFS technique to find the running frequency (or speed) such that the energy consumption is minimized and the real-time constraints are satisfied. The evaluation shows that our approach reduces the average energy consumption down to nearly 60%, compared to executing all the tasks locally at the maximum processor speed. Anas Toma, Santiago Pagani, Jian-Jia Chen, Wolfgang Karl, Jörg Henkel |
RTCSA | 4 |
| 2015 | Automatic task mapping and heterogeneity-aware fault tolerance: The benefits for runtime optimization and application development
Mario Kicherer, Wolfgang Karl |
J. Syst. Archit. | 2 |
| 2015 | Combined hardware-software multi-parallel prefiltering on the Convey HC-1 for fast homology detection
Michael Bromberger, Fabian Nowak, Wolfgang Karl |
Parallel Comput. | 3 |
| 2013 | Topic 4: High-Performance Architectures and Compilers - (Introduction)
Denis Barthou, Wolfgang Karl, Ramón Doallo, Evelyn Duesterwald, Sami Yehia |
Euro-Par | 2 |
| 2013 | Evaluation of Two Formulations of the Conjugate Gradients Method with Transactional Memory
Martin Schindewolf, Björn Rocker, Wolfgang Karl, Vincent Heuveline |
Euro-Par | 3 |
| 2013 | Multi-parallel prefiltering on the convey HC-1 for supporting homology detectionabstractGene databases used in research are huge and still grow at a fast pace. Many comparisons need to be done when searching similar (homologous) sequences in these databases for a given query sequence. Therefore, highly parallel architectures and much bandwidth are required for handling processing and transferring massive amounts of data. The Convey HC-1 with four FPGAs and high memory bandwidth of up to 76.8 GB/s seems very suitable for supporting this task as other bioinformatics applications have already been greatly supported by the HC-1. We research accelerating an application for searching homologous sequences. Limited by FPGA size only, we present a design that calculates 3 prefiltering scores per FPGA concurrently, i.e. 12 calculations in total. This score calculation for database sequences against the query profile is done by a modified Smith-Waterman scheme that is internally parallelized 16*8=128 times in contrast to the SSE implementation where only 16-fold parallelism can be exploited and where memory bandwidth poses the limiting factor. Preloading the query profile, we are able to transform the memory-bound SSE implementation to a compute-bound FPGA design which is only limited by FPGA size. Despite much lower clock rates, the FPGAs outperform SSE for the calculation of the prefiltering scores by a factor of 4.46. We achieve application speedup of 1.79 against the original, unmodified state-of-the-art SSE-based implementation because the score calculation accounts for less than 63% of the application runtime. Fabian Nowak, Michael Bromberger, Martin Schindewolf, Wolfgang Karl |
EuroMPI | 4 |
| 2012 | A Scalable Monitoring Infrastructure for Self-Organizing Many-Core ArchitecturesabstractSelf-organizing principles can address the growing complexity and the huge challenge of management and efficient utilization of adaptive many-core architectures. Fundamental for realizing a self-organizing behavior within such architectures is a dedicated monitoring infrastructure that provides the essential information about the system status and system behavior for realizing the basic property of self-awareness. This paper therefore proposes a flexible, hierarchical and scalable monitoring infrastructure for self-organizing, adaptive many-core architectures. The employed basic monitoring unit in the bottom monitoring layer performs data aggregation and filtering and reduces the amount of data that must be processed in higher monitoring layers. The middle layer performs first data analysis and is further responsible for hiding the heterogeneity of the underlying hardware configuration to the topmost monitoring layer. The latter is finally responsible for detecting changes in the system behavior and realizing self-awareness. The proposed monitoring infrastructure was evaluated entirely using a simulation framework. Results show that the infrastructure is able of detecting changes in the system behavior of an entire many-core system causing only a minor system disturbance. Further, the prototypical implementation of the basic monitoring unit proved that it can be realized very efficiently in hardware. David Kramer, Wolfgang Karl |
DSD | 2 |
| 2012 | A Low-Overhead Profiling and Visualization Framework for Hybrid Transactional MemoryabstractMulti-core prototyping presents a good opportunity for establishing low overhead and detailed profiling and visualization in order to study new research topics. In this paper, we design and implement a low execution, low area overhead profiling mechanism and a visualization tool for observing Transactional Memory behaviors on FPGA. To achieve this, we non-disruptively create and bring out events on the fly and process them offline on a host. There, our tool regenerates the execution from the collected events and produces traces for comprehensively inspecting the behavior of interacting multithreaded programs. With zero execution overhead for hardware TM events, single-instruction overhead for software TM events, and utilizing a low logic area of 2.3% per processor core, we run TM benchmarks to evaluate various different levels of profiling detail with an average runtime overhead of 6%. We demonstrate the usefulness of such detailed examination of SW/HW transactional behavior in two parts: (i) we speed up a TM benchmark by 24.1%, and (ii) we closely inspect transactions to point out pathologies. Oriol Arcas-Abella, Philipp Kirchhofer, Nehir Sönmez, Martin Schindewolf, Osman S. Unsal, Wolfgang Karl, Adrián Cristal |
FCCM | 6 |
| 2012 | What scientific applications can benefit from hardware transactional memory?abstractAchieving efficient and correct synchronization of multiple threads is a difficult and error-prone task at small scale and, as we march towards extreme scale computing, will be even more challenging when the resulting application is supposed to utilize millions of cores efficiently. Transactional Memory (TM) is a promising technique to ease the burden on the programmer, but only recently has become available on commercial hardware in the new Blue Gene/Q system and hence the real benefit for realistic applications has not been studied yet. This paper presents the first performance results of TM embedded into OpenMP on a prototype system of BG/Q and characterizes code properties that will likely lead to benefits when augmented with TM primitives. We first study the influence of thread count, environment variables and memory layout on TM performance and identify code properties that will yield performance gains with TM. Second, we evaluate the combination of OpenMP with multiple synchronization primitives on top of MPI to determine suitable task to thread ratios per node. Finally, we condense our findings into a set of best practices. These are applied to a Monte Carlo Benchmark and a Smoothed Particle Hydrodynamics method. In both cases an optimized TM version, executed with 64 threads on one node, outperforms a simple TM implementation. MCB with optimized TM yields a speedup of 27.45 over baseline. Martin Schindewolf, Barna L. Bihari, John C. Gyllenhaal, Martin Schulz 0001, Amy Wang, Wolfgang Karl |
SC | 6 |
| 2012 | A survey on hardware-aware and heterogeneous computing on multicore processors and acceleratorsabstractSUMMARY In the last few years, the landscape of parallel computing has been subject to profound and highly dynamic changes. The paradigm shift towards multicore and manycore technologies coupled with accelerators in a heterogeneous environment is offering a great potential of computing power for scientific and industrial applications. However, for one to take full advantage of these new technologies, holistic approaches coupling the expertise ranging from hardware architecture and software design to numerical algorithms are a pressing necessity. Parallel computing is no longer limited to supercomputers and is now much more diversified – with a multitude of technologies, architectures, and programming approaches leading to increased complexity for developers and engineers. In this work, we give – from the perspective of numerical simulation and applications – an overview of existing and emerging multicore and manycore technologies as well as accelerator concepts. We emphasize the challenges associated with high‐performance heterogeneous computing and discuss the interfaces needed to fill the gap between the hardware architecture and the implementation of efficient numerical algorithms. By means of this short survey – which stresses the necessity of hardware‐aware computing – we aim at giving assistance to users in scientific computing entering this fascinating field and help understanding associated issues and capabilities. Copyright © 2011 John Wiley & Sons, Ltd. Rainer Buchty, Vincent Heuveline, Wolfgang Karl, Jan-Philipp Weiss |
Concurr. Comput. Pract. Exp. | 3 |
| 2012 | Seamlessly portable applications: Managing the diversity of modern heterogeneous systemsabstractNowadays, many possible configurations of heterogeneous systems exist, posing several new challenges to application development: different types of processing units usually require individual programming models with dedicated runtime systems and accompanying libraries. If these are absent on an end-user system, e.g. because the respective hardware is not present, an application linked against these will break. This handicaps portability of applications being developed on one system and executed on other, differently configured heterogeneous systems. Moreover, the individual profit of different processing units is normally not known in advance. In this work, we propose a technique to effectively decouple applications from their accelerator-specific parts, respectively code. These parts are only linked on demand and thereby an application can be made portable across systems with different accelerators. As there are usually multiple hardware-specific implementations for a certain task, e.g., a CPU and a GPU version, a method is required to determine which are usable at all and which one is most suitable for execution on the current system. With our approach, application and hardware programmers can express the requirements and the abilities of the application and the hardware-specific implementations in a simplified manner. During runtime, the requirements and abilities are compared with regard to the present hardware in order to determine the usable implementations of a task. If multiple implementations are usable, an online-learning history-based selector is employed to determine the most efficient one. We show that our approach chooses the fastest usable implementation dynamically on several systems while introducing only a negligible overhead itself. Applied to an MPI application, our mechanism enables exploitation of local accelerators on different heterogeneous hosts without preliminary knowledge or modification of the application. Mario Kicherer, Fabian Nowak, Rainer Buchty, Wolfgang Karl |
ACM Trans. Archit. Code Optim. | 4 |
| 2011 | Introduction
Wolfgang Karl, Samuel Thibault, Stanimire Tomov, Taisuke Boku |
Euro-Par (2) | 1 |
| 2011 | Cost-aware function migration in heterogeneous systemsabstractToday's approaches towards heterogeneous computing rely on either the programmer or dedicated programming models to efficiently integrate heterogeneous components. In this work, we propose an adaptive cost-aware function-migration mechanism built on top of a light-weight hardware abstraction layer. With this mechanism, the highly dynamic task of choosing the most beneficial processing unit will be hidden from the programmer while causing only minor variation in the work and program flow. The migration mechanism transparently adapts to the current workload and system environment without the necessity of JIT compilation or binary translation. Mario Kicherer, Rainer Buchty, Wolfgang Karl |
HiPEAC | 3 |
| 2011 | Digital On-demand Computing Organism - Interaction between Monitoring and MiddlewareabstractOrganic Computing is a vital and promising research area. Inspired by nature, organic computing research wants to learn and adopt from techniques and properties of nature. The goal is to acquire the so called self-X properties like self-organization and self-healing. The DodOrg project introduces such an organic computing system for real-time applications, a whole new computing system from the bottom to the top. In this paper, we present the interaction between organic middleware and monitoring. Our results showed very promising results and only a small overhead for monitoring and the artificial hormone system based middleware. Alexander von Renteln, Uwe Brinkschulte, David Kramer, Wolfgang Karl, Christian Schuck, Jürgen Becker 0001 |
ISORC | 4 |
| 2010 | Cyberaide onServe: Software as a Service on Production GridsabstractThe Software as a Service (SaaS) methodology is a key paradigm of Cloud computing. In this paper, we focus on an interesting topic - to implement a Cloud computing functionality, the SaaS model, on existing production Grid infrastructures. In general, production Grids employ a Job-Submission-Execution (JSE) model with rigid access interfaces. In this paper we develop the Cyberaide onServe, a lightweight middleware with a virtual appliance. The Cyberaide onServe implements the SaaS methodology on production Grids by translating the SaaS model to the JSE model. The Cyberaide onServe virtual appliance is deployed on demand, hosts applications as Web services, accepts Web service invocations, and finally the Cyberaide onServe executes them on production Grids. We have deployed the Cyberaide onServe on the TeraGrid infrastructure and test results show Cyberaide onServe can provide the SaaS functionality with good performance. Tobias Kurze, Lizhe Wang 0001, Gregor von Laszewski, Jie Tao 0001, Marcel Kunze, David Kramer, Wolfgang Karl |
ICPP | 7 |
| 2010 | From source code to runtime behaviour: Software metrics help to select the computer architecture
Frank Eichinger, David Kramer, Klemens Böhm, Wolfgang Karl |
Knowl. Based Syst. | 4 |
| 2009 | Introduction
Pedro C. Diniz, Ben H. H. Juurlink, Alain Darte, Wolfgang Karl |
Euro-Par | 4 |
| 2008 | Scientific Cloud Computing: Early Definition and ExperienceabstractCloud computing emerges as a new computing paradigm which aims to provide reliable, customized and QoS guaranteed computing dynamic environments for end-users. This paper reviews recent advances of Cloud computing, identifies the concepts and characters of scientific Clouds, and finally presents an example of scientific Cloud for data centers Lizhe Wang 0001, Jie Tao 0001, Marcel Kunze, Alvaro Canales Castellanos, David Kramer, Wolfgang Karl |
HPCC | 6 |
| 2008 | Evaluating the Cache Architecture of Multicore ProcessorsabstractMicroprocessor architecture for both commercial and academical purpose is coming into a new generation: multiprocessors on a chip. Together with this novel architecture, questions and research topics also arise. For example, how to design the on-chip caches to avoid memory operations becoming the performance bottleneck? In this work, we study the impact of various cache architectures on the execution behavior of multi-threading applications. We focus on four general design issues: cache structure, configuration parameters, coherence influence, and prefetching strategies. The study is based on a self- developed cache simulator that models the functionality of a multicore cache hierarchy with arbitrary levels and various organizations. The achieved results can direct both hardware and program developers to optimize their cache designs or the program codes. Jie Tao 0001, Marcel Kunze, Wolfgang Karl |
PDP | 3 |
| 2007 | A Profiling Tool for Detecting Cache-Critical Data Structures
Jie Tao 0001, Tobias Gaugler, Wolfgang Karl |
Euro-Par | 3 |
| 2007 | Optimizing Cache Performance of the Discrete Wavelet Transform Using a Visualization ToolabstractThe 2D DWT consists of two 1D DWT in both directions: horizontal filtering processes the rows followed by vertical filtering processes the columns. It is well known that a straightforward implementation of the vertical filtering shows quite different performance with various working set sizes. The only reasonable explanation for this has to be the access behavior of the cache memory. As known, vertical filtering has mapping conflicts in the cache with a working set size that is power of two. However, it is not clear how this conflict forms and whether cache problems exist with other data sizes. Such knowledge is the base for efficient code optimization. In order to acquire this knowledge and to achieve more accurate optimization potentials, we apply a cache visualization tool to examine the runtime cache activities of the vertical implementation. We find that besides mapping conflicts, vertical filtering also shows a large number of capacity misses. More specifically, the visualization tool allows us to detect the parameters related to the strategies. This guarantees the feasibility of the optimization. Our initial experimental results on several different architectures show an up to 215% gain in execution time compared to an already optimized baseline implementation. Jie Tao 0001, Asadollah Shahbahrami, Ben H. H. Juurlink, Rainer Buchty, Wolfgang Karl, Stamatis Vassiliadis |
ISM | 5 |
| 2006 | A network agent for diagnosis and analysis of real-time Ethernet networksabstractWithin the field of automation technology the use of Industrial Ethernet is rising. This in turn demands devices capable of precisely recording, analyzing, and manipulating communication data for diagnostic purposes. Existing solutions so far lack required flexibility or are unable to cope with sustained Gigabit-per-second data streams. This is especially true for general-purpose approaches employing ordinary network adapters and plain software-based analysis.In this paper we describe a flexible and lightweight network agent for real-time, high-performance networks. This agent is capable of handling sustained data rates up to 2x 1GBit/s while offering real-time event-triggers, 10ns-resolution timestamps, real-time filtering, and statistics functions. An auxiliary processing unit as well as a modular software environment allow customization for a variety of tasks. The agent is realized as a dual processor SoC design on a Xilinx Virtex-II Pro FPGA. Hans-Peter Löb, Rainer Buchty, Wolfgang Karl |
CASES | 3 |
| 2006 | Topic 7: Parallel Computer Architecture and Instruction Level Parallelism
Eduard Ayguadé, Wolfgang Karl, Koen De Bosschere, Jean-Francois Collard |
Euro-Par | 2 |
| 2006 | Supporting Cache Locality Optimization with a Toolset
Jie Tao 0001, Wolfgang Karl |
Euro-Par | 2 |
| 2005 | YACO: A User Conducted Visualization Tool for Supporting Cache Optimization
Boris Quaing, Jie Tao 0001, Wolfgang Karl |
HPCC | 3 |
| 2005 | Monitoring cache behavior on parallel SMP architectures and related programming tools
Thomas Brandes, Helmut Schwamborn, Michael Gerndt, Jürgen Jeitner, Edmond Kereku, Martin Schulz 0001, Holger Brunst, Wolfgang E. Nagel, Reinhard Neumann, Ralph Müller-Pfefferkorn, Bernd Trenkler, Wolfgang Karl, Jie Tao 0001, Hans-Christian Hoppe |
Future Gener. Comput. Syst. | 12 |
| 2005 | Simulation as a tool for optimizing memory accesses on NUMA machines
Jie Tao 0001, Martin Schulz 0001, Wolfgang Karl |
Perform. Evaluation | 3 |
| 2004 | Topic 8: Parallel Computer Architecture and Instruction-Level Parallelism
Kemal Ebcioglu, Wolfgang Karl, André Seznec, Marco Aldinucci |
Euro-Par | 2 |
| 2004 | Impact of Cache Coherence Models on Performance of OpenMP Applications
Jie Tao 0001, Wolfgang Karl |
Euro-Par | 2 |
| 2003 | SMiLE: an integrated, multi-paradigm software infrastructure for SCI-basedclusters
Martin Schulz 0001, Jie Tao 0001, Carsten Trinitis, Wolfgang Karl |
Future Gener. Comput. Syst. | 4 |
| 2003 | ARS: an adaptive runtime system for locality optimization
Jie Tao 0001, Martin Schulz 0001, Wolfgang Karl |
Future Gener. Comput. Syst. | 3 |
| 2002 | SMiLE: An Integrated, Multi-Paradigm Software Infrastructure for SCI-Based ClustersabstractThe availability of a comprehensive software infrastructure is essential for the success a parallel architecture. In order to allow for the greatest possible flexibility, an infrastructure has to be designed in an integrated, easy-to-use manner and with the support of multiple programming paradigms and models to address a wide base of codes. SMiLE provides such an infrastructure for SCI (Scalable Coherent Interface) based clusters. It includes support for both a large range of message passing libraries as well as for almost arbitrary shared memory programming models. In addition, SMiLE also contains initial work on appropriate tool sets for performance optimizations. The complete infrastructure is implemented in way that is as closely relate d to the underlying hardware and is therefore capable of exploiting the benefits of the underlying network fabric and offering them to the user without significant overheads. Martin Schulz 0001, Jie Tao 0001, Carsten Trinitis, Wolfgang Karl |
CCGRID | 4 |
| 2001 | OpenSESAME: An Intuitive Dependability Modeling Environment Supporting Inter-Component DependenciesabstractThe paper proposes a novel modeling method for the evaluation of dependability measures of highly available systems. The proposed method, which has been implemented in the tool OpenSESAME (Simple but Extensive Structured Availability Modeling Environment), combines the advantages of Boolean methods and state space based methods. The tool supports the modeler with a set of well-defined, structured, intuitive input diagrams and tables, which are automatically transformed into GSPNs (Generalized Stochastic Petri Nets) for evaluation. To show the usefulness of the proposed method, it is applied to a model of a typical CompactPCI-based high availability system as can be found in the telecommunications area. Max Walter, Carsten Trinitis, Wolfgang Karl |
PRDC | 3 |
| 2000 | NEPHEW: Applying a Toolset for the Efficient Deployment of a Medical Image Application on SCI-Based Clusters
Wolfgang Karl, Martin Schulz 0001, Martin Völk, Sibylle Ilse Ziegler |
Euro-Par | 1 |
| 2000 | Electrical phenomena during Hot Swap eventsabstractThe exchange of a computer system's components during operation can be accomplished by the so called Hot Swap technology. This technology makes it possible to continuously run a computer system without the necessity of a shutdown for maintenance purposes, e.g. upgrading of a network adapter. Thus the overall uptime of a system can be drastically increased. The Hot Swap capability has been integrated into the so called CompactPCI technology. This paper summarizes the investigations that have been carried out with respect to the electrical behavior during Hot Swap events. Several simulations were performed with the H-SPICE program by AMP, Harrisburg, PA. After an introduction to live insertion phenomena in general, the CompactPCI specific simulations are described and recommendations for designing a Hot Swap capable system from an electrical point of view are given. Carsten Trinitis, Wolfgang Karl, Markus Leberecht |
PRDC | 2 |
| 1999 | Memory Characteristics of Iterative MethodsabstractConventional implementations of iterative numerical algorithms, especially multigrid methods, merely reach a disappointing small percentage of the theoretically available CPU performance when applied to representative large problems.One of the most important reasons for this phenomenon is that the current DRAM technology cannot provide the data fast enough to keep the CPU busy.Although the fundamentals of cache optimizations are quite simple, current compilers cannot optimize even elementary iterative schemes.In this paper, we analyze the memory and cache behavior of iterative methods with extensive profiling and describe program transformation techniques to improve the cache performance of two-and three-dimensional multigrid algorithms.This project is partially funded by DFG Ru 422/7-1,2.1 All benchmarks in the article were compiled with native FORTRAN77 compilers and aggressive optimizations enabled.On the Intel platform we used egcs (V2.91.60).The platforms include an Intel PentiumII Xeon PC (450 MHz, 450 MFLOPS), a SUN Ultra 60 (296 MHz, 592 MFLOPS), a HP SPP2200 Convex Exemplar Node (200 MHz, 800 MFLOPS), a Compaq PWS 500au (500 MHz, 1 GFLOPS), and a Compaq XP1000 (500 MHz, 1 GFLOPS).1 Christian Weiß 0001, Wolfgang Karl, Markus Kowarschik, Ulrich Rüde |
SC | 2 |
| 1998 | Exploiting Spatial and Temporal Locality of Accesses: A New Hardware-Based Monitoring Approach for DSM Systems
Robert Hockauf, Wolfgang Karl, Markus Leberecht, Michael Oberhuber, Michael Wagner 0003 |
Euro-Par | 2 |