Fabian Nowak

dblp:43/6950 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 39% High-performance computing · 30% Hardware accelerators and domain-specific architectures · 30%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › accelerator orchestration
accelerator selection
0.112012
Seamlessly portable applications: Managing the diversity of modern heterogeneous systems · ACM Trans. Archit. Code Optim. 2012
High-performance computing
application portability
0.112012
Seamlessly portable applications: Managing the diversity of modern heterogeneous systems · ACM Trans. Archit. Code Optim. 2012
Parallel and multicore computing
parallel programming runtimes
0.112012
Seamlessly portable applications: Managing the diversity of modern heterogeneous systems · ACM Trans. Archit. Code Optim. 2012
Parallel and multicore computing › parallel programming models › message passing
MPI applications
0.012012
Seamlessly portable applications: Managing the diversity of modern heterogeneous systems · ACM Trans. Archit. Code Optim. 2012

Methods — techniques the papers use, named apart from their topics

online-learning history-based selection · 0.1
YearPublicationVenuePosition
2015 Combined hardware-software multi-parallel prefiltering on the Convey HC-1 for fast homology detection
Michael Bromberger, Fabian Nowak, Wolfgang Karl
Parallel Comput.2
2013 Multi-parallel prefiltering on the convey HC-1 for supporting homology detection
abstract
Gene databases used in research are huge and still grow at a fast pace. Many comparisons need to be done when searching similar (homologous) sequences in these databases for a given query sequence. Therefore, highly parallel architectures and much bandwidth are required for handling processing and transferring massive amounts of data. The Convey HC-1 with four FPGAs and high memory bandwidth of up to 76.8 GB/s seems very suitable for supporting this task as other bioinformatics applications have already been greatly supported by the HC-1. We research accelerating an application for searching homologous sequences. Limited by FPGA size only, we present a design that calculates 3 prefiltering scores per FPGA concurrently, i.e. 12 calculations in total. This score calculation for database sequences against the query profile is done by a modified Smith-Waterman scheme that is internally parallelized 16*8=128 times in contrast to the SSE implementation where only 16-fold parallelism can be exploited and where memory bandwidth poses the limiting factor. Preloading the query profile, we are able to transform the memory-bound SSE implementation to a compute-bound FPGA design which is only limited by FPGA size. Despite much lower clock rates, the FPGAs outperform SSE for the calculation of the prefiltering scores by a factor of 4.46. We achieve application speedup of 1.79 against the original, unmodified state-of-the-art SSE-based implementation because the score calculation accounts for less than 63% of the application runtime.
Fabian Nowak, Michael Bromberger, Martin Schindewolf, Wolfgang Karl
EuroMPI1
2012 Seamlessly portable applications: Managing the diversity of modern heterogeneous systems
abstract
Nowadays, many possible configurations of heterogeneous systems exist, posing several new challenges to application development: different types of processing units usually require individual programming models with dedicated runtime systems and accompanying libraries. If these are absent on an end-user system, e.g. because the respective hardware is not present, an application linked against these will break. This handicaps portability of applications being developed on one system and executed on other, differently configured heterogeneous systems. Moreover, the individual profit of different processing units is normally not known in advance. In this work, we propose a technique to effectively decouple applications from their accelerator-specific parts, respectively code. These parts are only linked on demand and thereby an application can be made portable across systems with different accelerators. As there are usually multiple hardware-specific implementations for a certain task, e.g., a CPU and a GPU version, a method is required to determine which are usable at all and which one is most suitable for execution on the current system. With our approach, application and hardware programmers can express the requirements and the abilities of the application and the hardware-specific implementations in a simplified manner. During runtime, the requirements and abilities are compared with regard to the present hardware in order to determine the usable implementations of a task. If multiple implementations are usable, an online-learning history-based selector is employed to determine the most efficient one. We show that our approach chooses the fastest usable implementation dynamically on several systems while introducing only a negligible overhead itself. Applied to an MPI application, our mechanism enables exploitation of local accelerators on different heterogeneous hosts without preliminary knowledge or modification of the application.
Mario Kicherer, Fabian Nowak, Rainer Buchty, Wolfgang Karl
ACM Trans. Archit. Code Optim.2