Sofia Dimoudi

dblp:250/2692 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 44% Hardware accelerators and domain-specific architectures · 22% Memory systems · 22%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
FFT-based convolution
0.412020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020
GPUs and heterogeneous computing
GPU kernel optimization
0.412020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy
0.412020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020
Memory systems
shared memory
0.412020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020
High-performance computing › tensor computation
convolution
0.112020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020
Storage systems
signal processing
0.112020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020

Methods — techniques the papers use, named apart from their topics

overlap-and-save · 0.4FFT · 0.4
YearPublicationVenuePosition
2020 GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory
abstract
We present an implementation of the overlap-and-save method, a method for the convolution of very long signals with short response functions, which is tailored to GPUs. We have implemented several FFT algorithms (using the CUDA programming language) which exploit GPU shared memory, allowing for GPU accelerated convolution. We compare our implementation with an implementation of the overlap-and-save algorithm utilizing the NVIDIA FFT library (cuFFT). We demonstrate that by using a shared memory based FFT we can achieved significant speed-ups for certain problem sizes and lower the memory requirements of the overlap-and-save method on GPUs.
Karel Adámek, Sofia Dimoudi, Michael B. Giles, Wes Armour
ACM Trans. Archit. Code Optim.2