Frederik Thorøe

dblp:235/2463 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
parallelizing compiler
0.412019
Incremental flattening for nested data parallelism · PPoPP 2019
Parallel and multicore computing › data parallelism
nested data parallelism
0.412019
Incremental flattening for nested data parallelism · PPoPP 2019
Parallel and multicore computing
parallel programming models
0.412019
Incremental flattening for nested data parallelism · PPoPP 2019
Parallel and multicore computing › parallelizing compiler
parallel code generation
0.112019
Incremental flattening for nested data parallelism · PPoPP 2019

Methods — techniques the papers use, named apart from their topics

static clustering · 0.8auto-tuning · 0.8
YearPublicationVenuePosition
2019 Incremental flattening for nested data parallelism
abstract
Compilation techniques for nested-parallel applications that can adapt to hardware and dataset characteristics are vital for unlocking the power of modern hardware. This paper proposes such a technique, which builds on flattening and is applied in the context of a functional data-parallel language. Our solution uses the degree of utilized parallelism as the driver for generating a multitude of code versions, which together cover all possible mappings of the application's regular nested parallelism to the levels of parallelism supported by the hardware. These code versions are then combined into one program by guarding them with predicates, whose threshold values are automatically tuned to hardware and dataset characteristics. Our unsupervised method---of statically clustering datasets to code versions---is different from autotuning work that typically searches for the combination of code transformations producing a single version, best suited for a specific dataset or on average for all datasets.
Troels Henriksen, Frederik Thorøe, Martin Elsman, Cosmin E. Oancea
PPoPP2