VLDB 2026 Research / reviewers in the wild / expert
Eduard Mehofer
dblp:01/6706
· DBLP profile ↗
15ranked-venue papers
1as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 77% High-performance computing · 23% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 2 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization › parallel language compilation
data-parallel compilation |
0.0 | 1 | 2002 | Distribution Assignment Placement: Effective Optimization of Redistribution Costs · IEEE Trans. Parallel Distributed Syst. 2002 |
Parallel and multicore computing
data-parallel programming |
0.0 | 1 | 2002 | Distribution Assignment Placement: Effective Optimization of Redistribution Costs · IEEE Trans. Parallel Distributed Syst. 2002 |
Methods — techniques the papers use, named apart from their topics
redundancy elimination · 0.1dead code elimination · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Runtime and energy constrained work scheduling for heterogeneous systems
Valon Raca, Seeun William Umboh, Eduard Mehofer, Bernhard Scholz |
J. Supercomput. | 3 |
| 2020 | clusterCL: comprehensive support for multi-kernel data-parallel applications in heterogeneous asymmetric clusters
Valon Raca, Eduard Mehofer |
J. Supercomput. | 2 |
| 2016 | Optimal Time and Energy Efficient Work Distributions in Heterogeneous SystemsabstractHeterogeneous cluster architectures with different types of compute devices are wide-spread in the field of high performance computing introducing new kind of challenges. Time-efficient distribution of work onto heterogeneous devices is much more intricate than in the homogeneous case. Moreover, the use of heterogeneous devices makes it necessary to address energy efficiency as well. This results in bi-criteria optimization problems whereby often time efficiency and energy efficiency are conflicting objectives pushing in different directions. In this paper we define four optimization problems which are of particular interest for developers, give a formal description of these problems, and try to find optimal solutions for them efficiently with minimal effort. Essential is the fact that a user may specify time or energy constraints for which optimal work distributions onto the devices are determined. In addition, we raise the question how to recommend reasonable solutions, if no constraints have been specified. An experimental section examines the solutions obtained from our algorithms. Valon Raca, Eduard Mehofer, Marcus Hudec |
PDP | 2 |
| 2015 | Device-Sensitive Framework for Handling Heterogeneous Asymmetric Clusters EfficientlyabstractHeterogeneous systems with different types of compute devices are common nowadays in the field of High Performance Computing (HPC). This heterogeneity is not limited to compute devices, but also includes cluster nodes with different hardware configurations leading to asymmetric cluster architectures. In such a hierarchical system OpenCL is not sufficient any more. Support is required to distribute the work efficiently onto the non-identical cluster nodes. Different behavior of the individual compute devices with respect to execution time and energy consumption has to be taken into account to meet the demands of the user. Our framework provides a transparent view on the different compute devices alleviating the programmer to deal with the hardware architecture and device execution behavior explicitly. Besides efficiency considerations, the device-sensitive feature includes in addition handling of device failures and appropriate recovery actions. Experiments show that our framework succeeds in distributing the work onto compute devices efficiently. Valon Raca, Eduard Mehofer |
SBAC-PAD | 2 |
| 2014 | Modeling and optimizing large-scale data flows
Alexander Wöhrer, Peter Brezany, Ivan Janciak, Eduard Mehofer |
Future Gener. Comput. Syst. | 4 |
| 2012 | Optimization Techniques and Performance Analyses of Two Life Science Algorithms for Novel GPU ArchitecturesabstractIn this paper we evaluate two life science algorithms, namely Needleman-Wunsch sequence alignment and Direct Coulomb Summation, for GPUs. Whereas for Needleman-Wunsch it is difficult to get good performance numbers, Direct Coulomb Summation is particularly suitable for graphics cards. We present several optimization techniques, analyze the theoretical potential of the optimizations with respect to the algorithms, and measure the effect on execution times. We target the recent NVIDIA Fermi architecture to evaluate the performance impacts of novel hardware features like the cache subsystem on optimizing transformations. We compare the execution times of CUDA and OpenCL code versions for Fermi and predecessor models with parallel OpenMP versions executed on the main CPU. David Dilch, Eduard Mehofer |
PDP | 2 |
| 2009 | Experimental Study of Multithreading to Improve Memory Hierarchy Performance of Multi-core Processors for Scientific ApplicationsabstractIn this paper we study performance characteristics and parallelization strategies for recently shipped, powerful multi-core processors - IBM Power6 and Sun T2 Plus - for high-end scientific computing. Central aspect is data locality. First, we investigate the impacts of good and bad data locality by modifying data accesses. Next, we study the impact of multithreading with respect to data locality based on the data-parallel programming approach. The level of parallelism is increased by assigning multiple threads onto one core in order to hide processor stalls caused by bad data locality. We measure the impacts of data locality and multithreading in terms of execution times and bandwidth for synthetic micro-benchmarks, a matrix multiplication kernel, and an application from Bioinformatics. The results indicate that substantial performance improvements can be obtained with minor effort by utilizing multithreading. Enes Bajrovic, Eduard Mehofer |
CISIS | 2 |
| 2003 | Partial Redundancy Elimination with Predication Techniques
Bernhard Scholz, Eduard Mehofer, R. Nigel Horspool |
Euro-Par | 2 |
| 2002 | A Representation for Bit Section Based Analysis and Optimization
Rajiv Gupta 0001, Eduard Mehofer, Youtao Zhang |
CC | 2 |
| 2002 | Distribution Assignment Placement: Effective Optimization of Redistribution CostsabstractData locality and workload balance are key factors for getting high performance out of data-parallel programs on multiprocessor architectures. Data-parallel languages such as High-Performance Fortran (HPF) thus offer means allowing a programmer both to specify data distributions and to change them dynamically in order to maintain these properties. On the other hand, redistributions can be quite expensive and can significantly degrade a program's performance. They must thus be reduced to a minimum. In this article, we present a novel, aggressive approach for avoiding unnecessary remappings, which works by eliminating partially dead and partially redundant distribution changes. Basically, this approach evolves from extending and combining two algorithms for these optimizations, each achieving optimal results on its own. In distinction to the sequential setting, the data-parallel setting leads naturally to a family of algorithms of varying power and efficiency, allowing requirement-customized solutions. The power and flexibility of the new approach are demonstrated by various examples, which range from typical HPF fragments to real-world programs. Performance measurements underline its importance and show its effectiveness on different hardware platforms and in different settings. Jens Knoop, Eduard Mehofer |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2001 | A Novel Probabilistic Data Flow Framework
Eduard Mehofer, Bernhard Scholz |
CC | 1 |
| 2001 | Development and performance analysis of real-world applications for distributed and parallel architecturesabstractAbstract Several large real‐world applications have been developed for distributed and parallel architectures. We examine two different program development approaches. First, the usage of a high‐level programming paradigm which reduces the time to create a parallel program dramatically but sometimes at the cost of a reduced performance; a source‐to‐source compiler, has been employed to automatically compile programs—written in a high‐level programming paradigm—into message passing codes. Second, a manual program development by using a low‐level programming paradigm—such as message passing—enables the programmer to fully exploit a given architecture at the cost of a time‐consuming and error‐prone effort. Performance tools play a central role in supporting the performance‐oriented development of applications for distributed and parallel architectures. SCALA—a portable instrumentation, measurement, and post‐execution performance analysis system for distributed and parallel programs—has been used to analyze and to guide the application development, by selectively instrumenting and measuring the code versions, by comparing performance information of several program executions, by computing a variety of important performance metrics, by detecting performance bottlenecks, and by relating performance information back to the input program. We show several experiments of SCALA when applied to real‐world applications. These experiments are conducted for a NEC Cenju‐4 distributed‐memory machine and a cluster of heterogeneous workstations and networks. Copyright © 2001 John Wiley & Sons, Ltd. Thomas Fahringer, Peter Blaha, A. Hössinger, J. Luitz, Eduard Mehofer, Hans Moritsch, Bernhard Scholz |
Concurr. Comput. Pract. Exp. | 5 |
| 1999 | Buffer-Safe and Cost-Driven Communication Optimization
Thomas Fahringer, Eduard Mehofer |
J. Parallel Distributed Comput. | 2 |
| 1998 | Problem and Machine Sensitive Communication Optimization
Thomas Fahringer, Eduard Mehofer |
International Conference on Supercomputing | 2 |
| 1997 | Optimal Distribution Assignment Placement
Jens Knoop, Eduard Mehofer |
Euro-Par | 2 |