VLDB 2026 Research / reviewers in the wild / expert
Hans Moritsch
dblp:07/462
· DBLP profile ↗
7ranked-venue papers
1as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Parallel and multicore computing · 83% Performance modeling and evaluation · 17% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
autotuning |
0.1 | 1 | 2012 | A multi-objective auto-tuning framework for parallel codes · SC 2012 |
Compilers and program optimization › loop optimization
loop tiling |
0.1 | 1 | 2012 | A multi-objective auto-tuning framework for parallel codes · SC 2012 |
Parallel and multicore computing
parallel programming runtimes |
0.1 | 1 | 2012 | A multi-objective auto-tuning framework for parallel codes · SC 2012 |
Parallel and multicore computing › parallel computing › parallel application performance
parallel program performance analysis |
0.0 | 1 | 2001 | On using SCALEA for performance analysis of distributed and parallel programs · SC 2001 |
Parallel and multicore computing
data distribution |
0.0 | 1 | 1993 | Dynamic data distributions in Vienna Fortran · SC 1993 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 1993 | Dynamic data distributions in Vienna Fortran · SC 1993 |
Compilers and program optimization › parallel language compilation
data-parallel compilation |
0.0 | 1 | 1993 | Dynamic data distributions in Vienna Fortran · SC 1993 |
Methods — techniques the papers use, named apart from their topics
runtime adaptation · 0.3multi-objective optimization · 0.3compiler auto-tuning · 0.3instrumentation · 0.0hardware profiling · 0.0dynamic code region call graph · 0.0data distribution · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | A multi-objective auto-tuning framework for parallel codesabstractIn this paper we introduce a multi-objective autotuning framework comprising compiler and runtime components. Focusing on individual code regions, our compiler uses a novel search technique to compute a set of optimal solutions, which are encoded into a multi-versioned executable. This enables the runtime system to choose specifically tuned code versions when dynamically adjusting to changing circumstances. We demonstrate our method by tuning loop tiling in cache-sensitive parallel programs, optimizing for both runtime and efficiency. Our static optimizer finds solutions matching or surpassing those determined by exhaustively sampling the search space on a regular grid, while using less than 4% of the computational effort on average. Additionally, we show that parallelism-aware multi-versioning approaches like our own gain a performance improvement of up to 70% over solutions tuned for only one specific number of threads. Herbert Jordan, Peter Thoman, Juan José Durillo, Simone Pellegrini, Philipp Gschwandtner, Thomas Fahringer, Hans Moritsch |
SC | 7 |
| 2002 | On the Evaluation of JavaSymphony for Cluster ApplicationsabstractIn the past few years, increasing interest has been shown in using Java as a language for performance-oriented distributed and parallel computing. Most Java-based systems that support portable parallel and distributed computing either require the programmer to deal with intricate low level details of Java which can be a tedious, time-consuming and error-prone task, or prevent the programmer from controlling locality of data. In contrast to most existing systems, JavaSymphony - a class library written entirely in Java - allows to control parallelism, load balancing and locality at a high level. Objects can be explicitly distributed and migrated based on virtual architectures which impose a virtual hierarchy on a distributed/parallel system of physical computing nodes. The concept of blocking/nonblocking remote method invocation is used to exchange data among distributed objects and to process work by remote objects. We evaluate the JavaSymphony programming API for a variety of distributed/parallel algorithms which comprises backtracking, N-body, encryption/decryption algorithms and asynchronous nested optimization algorithms. Performance results are presented for both homogeneous and heterogeneous cluster architectures. Moreover, we compare JavaSymphony with an alternative well-known semi-automatic system. Thomas Fahringer, Alexandru Jugravu, Beniamino Di Martino, Salvatore Venticinque, Hans Moritsch |
CLUSTER | 5 |
| 2002 | High-performance numerical pricing methodsabstractAbstract The pricing of financial derivatives is an important field in finance and constitutes a major component of financial management applications. The uncertainty of future events often makes analytic approaches infeasible and, hence, time‐consuming numerical simulations are required. In the Aurora Financial Management System, pricing is performed on the basis of lattice representations of stochastic multidimensional scenario processes using the Monte Carlo simulation and Backward Induction methods, the latter allowing for the exploitation of shared‐memory parallelism. We present the parallelization of a Backward Induction numerical pricing kernel on a cluster of SMPs using HPF+, an extended version of High‐Performance Fortran. Based on language extensions for specifying a hierarchical mapping of data onto an SMP cluster, the compiler generates a hybrid‐parallel program combining distributed‐memory and shared‐memory parallelism. We outline the parallelization strategy adopted by the VFC compiler and present an experimental evaluation of the pricing kernel on an NEC SX‐5 vector supercomputer and a Linux SMP cluster, comparing a pure MPI version to a hybrid‐parallel MPI/OpenMP version. Copyright © 2002 John Wiley & Sons, Ltd. Hans Moritsch, Siegfried Benkner |
Concurr. Comput. Pract. Exp. | 1 |
| 2001 | On using SCALEA for performance analysis of distributed and parallel programsabstractIn this paper we give an overview of SCALEA, which is a new performance analysis tool for OpenMP, MPI, HPF, and mixed parallel/distributed programs. SCALEA instruments, executes and measures programs and computes a variety of performance overheads based on a novel overhead classification. Source code and HW-profiling is combined in a single system which significantly extends the scope of possible overheads that can be measured and examined, ranging from HW-counters, such as the number of cache misses or floating point operations, to more complex performance metrics, such as control or loss of parallelism. Moreover, SCALEA uses a new representation of code regions, called the dynamic code region call graph, which enables detailed overhead analysis for arbitrary code regions. An instrumentation description file is used to relate performance information to code regions of the input program and to reduce instrumentation overhead. Several experiments with realistic codes that cover MPI, OpenMP, HPF, and mixed OpenMP/MPI codes demonstrate the usefulness of SCALEA. Hong Linh Truong 0001, Thomas Fahringer, Georg Madsen, Allen D. Malony, Hans Moritsch, Sameer Shende |
SC | 5 |
| 2001 | Development and performance analysis of real-world applications for distributed and parallel architecturesabstractAbstract Several large real‐world applications have been developed for distributed and parallel architectures. We examine two different program development approaches. First, the usage of a high‐level programming paradigm which reduces the time to create a parallel program dramatically but sometimes at the cost of a reduced performance; a source‐to‐source compiler, has been employed to automatically compile programs—written in a high‐level programming paradigm—into message passing codes. Second, a manual program development by using a low‐level programming paradigm—such as message passing—enables the programmer to fully exploit a given architecture at the cost of a time‐consuming and error‐prone effort. Performance tools play a central role in supporting the performance‐oriented development of applications for distributed and parallel architectures. SCALA—a portable instrumentation, measurement, and post‐execution performance analysis system for distributed and parallel programs—has been used to analyze and to guide the application development, by selectively instrumenting and measuring the code versions, by comparing performance information of several program executions, by computing a variety of important performance metrics, by detecting performance bottlenecks, and by relating performance information back to the input program. We show several experiments of SCALA when applied to real‐world applications. These experiments are conducted for a NEC Cenju‐4 distributed‐memory machine and a cluster of heterogeneous workstations and networks. Copyright © 2001 John Wiley & Sons, Ltd. Thomas Fahringer, Peter Blaha, A. Hössinger, J. Luitz, Eduard Mehofer, Hans Moritsch, Bernhard Scholz |
Concurr. Comput. Pract. Exp. | 6 |
| 2000 | Evaluation of P3T+: A Performance Estimator for Distributed and Parallel ApplicationsabstractIn this paper, we report on experiences with P/sup 3/T+, a performance estimator for distributed and parallel programs which is used to examine at compile time the performance outcome of changes in code, problem and machine sizes, and target architectures. P/sup 3/T+ computes a variety of performance parameters including work distribution, number of transfers, amount of data transferred, transfer times, computation times, and number of cache misses. It is unique in that it models programs, code transformations and parallel and distributed architectures and derives a performance prediction based on all three of these elements. P/sup 3/T+ is the successor tool of P/sup 3/T which computed a similar set of performance parameters, however for parallel programs only. P/sup 3/T+ has been re-designed and re-implemented from scratch and goes beyond P/sup 3/T by extending the class of programs that cart be handled and by employing several novel estimation methods (symbolic analysis, simulation, pre-measured kernel codes, etc.). The core part of this paper reports on the evaluation of P/sup 3/T+ to demonstrate both accuracy and usefulness of this tool for realistic kernel codes taken from real-world applications (pricing of financial derivatives and quantum mechanical calculations of solids). Thomas Fahringer, A. Pozgaj, Hans Moritsch, J. Luitz |
IPDPS | 3 |
| 1993 | Dynamic data distributions in Vienna FortranabstractNo abstract available. Barbara M. Chapman, Piyush Mehrotra, Hans Moritsch, Hans P. Zima |
SC | 3 |