Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Karel Adámek

dblp:154/6765 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
1since 2021 · last 2024
0000-0003-2797-0595ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Energy-efficient computing · 46% GPUs and heterogeneous computing · 23% Hardware accelerators and domain-specific architectures · 12%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing
energy measurement
0.812024
Accurate and Convenient Energy Measurements for GPUs: A Detailed Study of NVIDIA GPU's Built-In Power Sensor · SC 2024
Hardware accelerators and domain-specific architectures
FFT-based convolution
0.412020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020
GPUs and heterogeneous computing
GPU kernel optimization
0.412020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy
0.412020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020
Memory systems
shared memory
0.412020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020
Energy-efficient computing
GPU power consumption
0.212024
Accurate and Convenient Energy Measurements for GPUs: A Detailed Study of NVIDIA GPU's Built-In Power Sensor · SC 2024
High-performance computing › tensor computation
convolution
0.112020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020
Storage systems
signal processing
0.112020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020

Methods — techniques the papers use, named apart from their topics

power sensor calibration · 0.8overlap-and-save · 0.4FFT · 0.4
YearPublicationVenuePosition
2024 Accurate and Convenient Energy Measurements for GPUs: A Detailed Study of NVIDIA GPU's Built-In Power Sensor
abstract
GPU has emerged as the go-to accelerator for HPC workloads, however its power consumption has become a major limiting factor for further scaling HPC systems. An accurate understanding of GPU power consumption is essential for further improving its energy efficiency, and consequently reducing the associated carbon footprint. Despite the limited documentation and lack of understanding, NVIDIA GPUs’ built-in power sensor is widely used in energy-efficient computing research. Our study seeks to elucidate the internal mechanisms of the power readings provided by nvidia-smi and assess the accuracy of the measurements. We evaluated over 70 different GPUs across 12 architectural generations, and identified several unforeseen problems that can lead to drastic under/overestimation of energy consumed, for example on the A100 and H100 GPUs only 25% of the runtime is sampled. We proposed several mitigations that could reduce the energy measurement error by an average of 35% in the test cases we present.
Karel Adámek, Wes Armour
SC2
2020 GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory
abstract
We present an implementation of the overlap-and-save method, a method for the convolution of very long signals with short response functions, which is tailored to GPUs. We have implemented several FFT algorithms (using the CUDA programming language) which exploit GPU shared memory, allowing for GPU accelerated convolution. We compare our implementation with an implementation of the overlap-and-save algorithm utilizing the NVIDIA FFT library (cuFFT). We demonstrate that by using a shared memory based FFT we can achieved significant speed-ups for certain problem sizes and lower the memory requirements of the overlap-and-save method on GPUs.
Karel Adámek, Sofia Dimoudi, Michael B. Giles, Wes Armour
ACM Trans. Archit. Code Optim.1
2018 Building the World's Largest Radio Telescope: The Square Kilometre Array Science Data Processor
abstract
The Square Kilometre Array (SKA) will be the largest radio telescope constructed to date and the largest Big Data project in the known Universe. The first phase of the project will generate 160 terabytes every second. This amounts to 5 zettabytes (5 million petabytes) of data that will be generated by the facility each year - a data rate equivalent to 5 times the estimated global internet traffic in 2015. These data need to be reduced and then continuously ingested by the SKA Science Data Processor (SDP). Within the SDP Consortium, we are contributing to various roles in the development of the telescope including building a lightweight end-to-end prototype of the major components of the SDP system - a project we call the SDP Integration Prototype (SIP). The aim is to build a mini, fully-operational SDP, for which we have been developing realistic SKA-like science pipelines that can handle these unprecedented data volumes.
Jamie S. Farnes, Ben Mort, Fred Dulwich, Karel Adámek, Anna Brown, Jan Novotný, Stef Salvini, Wes Armour
eScience4