Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Rodrigo Dominguez

dblp:76/7750 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2010
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization › vectorization
loop vectorization
0.112010
Data transformations enabling loop vectorization on multithreaded data parallel architectures · PPoPP 2010
Parallel and multicore computing
data parallelism
0.112010
Data transformations enabling loop vectorization on multithreaded data parallel architectures · PPoPP 2010
Parallel and multicore computing › data parallelism
SIMD vectorization
0.112010
Data transformations enabling loop vectorization on multithreaded data parallel architectures · PPoPP 2010

Methods — techniques the papers use, named apart from their topics

mathematical modeling of memory access patterns · 0.2
YearPublicationVenuePosition
2010 Data transformations enabling loop vectorization on multithreaded data parallel architectures
abstract
Loop vectorization, a key feature exploited to obtain high performance on Single Instruction Multiple Data (SIMD) vector architectures, is significantly hindered by irregular memory access patterns in the data stream. This paper describes data transformations that allow us to vectorize loops targeting massively multithreaded data parallel architectures. We present a mathematical model that captures loop-based memory access patterns and computes the most appropriate data transformations in order to enable vectorization. Our experimental results show that the proposed data transformations can significantly increase the number of loops that can be vectorized and enhance the data-level parallelism of applications. Our results also show that the overhead associated with our data transformations can be easily amortized as the size of the input data set increases. For the set of high performance benchmark kernels studied, we achieve consistent and significant performance improvements (up to 11.4X) by applying vectorization using our data transformation approach.
Byunghyun Jang, Perhaad Mistry, Dana Schaa, Rodrigo Dominguez, David R. Kaeli
PPoPP4