Wes Armour

dblp:154/6449 · also Wesley Armour · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0003-1756-3064ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Optimization for machine learning · 58% Deep learning architectures and training · 20% Efficient and distributed learning · 17%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Energy-efficient computing · 45% GPUs and heterogeneous computing · 20% Hardware accelerators and domain-specific architectures · 10%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent
1.012026
Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training · AAAI 2026
Machine learning › Deep learning architectures and training › training optimization
large-batch training
1.012026
Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training · AAAI 2026
Machine learning › Optimization for machine learning › gradient-based optimization › gradient descent
natural gradient descent
1.012026
Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training · AAAI 2026
Machine learning › Optimization for machine learning
second-order optimization
1.012026
Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training · AAAI 2026
Machine learning › Efficient and distributed learning › inference efficiency
energy-efficient inference
0.912025
The Hidden Joules: Evaluating the Energy Consumption of Vision Backbones for Progress Towards More Efficient Model Inference · ICML 2025
Energy-efficient computing
energy measurement
0.812024
Accurate and Convenient Energy Measurements for GPUs: A Detailed Study of NVIDIA GPU's Built-In Power Sensor · SC 2024
Compilers and program optimization
vectorization
0.412020
Evaluating Auto-Vectorizing Compilers through Objective Withdrawal of Useful Information · ACM Trans. Archit. Code Optim. 2020
Performance modeling and evaluation
compiler performance evaluation
0.412020
Evaluating Auto-Vectorizing Compilers through Objective Withdrawal of Useful Information · ACM Trans. Archit. Code Optim. 2020
Hardware accelerators and domain-specific architectures
FFT-based convolution
0.412020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020
GPUs and heterogeneous computing
GPU kernel optimization
0.412020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy
0.412020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020
Memory systems
shared memory
0.412020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020
Computer vision › Image recognition and object detection
image classification
0.312025
The Hidden Joules: Evaluating the Energy Consumption of Vision Backbones for Progress Towards More Efficient Model Inference · ICML 2025
Energy-efficient computing
GPU power consumption
0.212024
Accurate and Convenient Energy Measurements for GPUs: A Detailed Study of NVIDIA GPU's Built-In Power Sensor · SC 2024
High-performance computing › tensor computation
convolution
0.112020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020
Storage systems
signal processing
0.112020
GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory · ACM Trans. Archit. Code Optim. 2020

Methods — techniques the papers use, named apart from their topics

energy efficiency scoring · 1.7empirical benchmarking · 1.7variance-aware update · 1.0kronecker-factored approximate curvature · 1.0fisher-orthogonal projection · 1.0proxy kernels · 0.9information withdrawal · 0.9power sensor calibration · 0.8overlap-and-save · 0.4FFT · 0.4
YearPublicationVenuePosition
2026 Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training
abstract
Modern GPUs are equipped with large amounts of high-bandwidth memory, enabling them to support mini-batch sizes of up to tens of thousands of training samples. However, most existing optimizers struggle to perform effectively at such a large batch size. As batch size increases, gradient noise decreases due to averaging over many samples, limiting the ability of first-order methods to escape sharp or suboptimal minima and reach the global minimum. Meanwhile, second-order methods like the natural gradient with Kronecker-Factored Approximate Curvature (KFAC) often require excessively high damping to remain stable at large batch sizes. This high damping effectively ``washes out" the curvature information that gives these methods their advantage, reducing their performance to that of simple gradient descent. In this paper, we introduce Fisher-Orthogonal Projection (FOP), a novel technique that restores the effectiveness of the second-order method at very large batch sizes, enabling scalable training with improved generalization and faster convergence. FOP constructs a variance-aware update direction by leveraging gradients from two sub-batches, enhancing the average gradient with a component of the gradient difference that is orthogonal to the average under the Fisher-metric. Through extensive benchmarks, we show that FOP accelerates convergence by ×1.2–1.3 over K-FAC and ×1.5–1.7 over SGD/AdamW at the same moderate batch sizes, while at extreme scales it achieves up to a ×7.5 speedup. Unlike other methods, FOP maintains small-batch accuracy when scaling to extremely large batch sizes. Moreover, it reduces Top-1 error by 2.3–3.3% on long-tailed CIFAR benchmarks, demonstrating robust generalization under severe class imbalance. Our lightweight, geometry-aware use of intra-batch variance makes natural-gradient optimization practical on modern data-centre GPUs. FOP is open-source and pip-installable, which can be integrated into existing training code with a single line and no extra configuration.
Yishun Lu, Wes Armour
AAAI2
2025 The Hidden Joules: Evaluating the Energy Consumption of Vision Backbones for Progress Towards More Efficient Model Inference
abstract
Deep learning has achieved significant success but poses increasing concerns about energy consumption and sustainability. Despite these concerns, there is a lack of understanding of their energy efficiency during inference. In this study, we conduct a comprehensive analysis of the inference energy consumption of 1,200 ImageNet classification models—the largest evaluation of its kind to date. Our findings reveal a steep decline in accuracy gains relative to the increase in energy usage, highlighting sustainability concerns in the pursuit of marginal improvements. We identify key factors contributing to energy consumption and demonstrate methods to improve energy efficiency. To promote more sustainable AI practices, we introduce an energy efficiency scoring system and develop an interactive web application that allows users to compare models based on accuracy and energy consumption. By providing extensive empirical data and practical tools, we aim to facilitate informed decision-making and encourage collaborative efforts in the development of energy-efficient AI technologies.
Wes Armour
ICML2
2024 Accurate and Convenient Energy Measurements for GPUs: A Detailed Study of NVIDIA GPU's Built-In Power Sensor
abstract
GPU has emerged as the go-to accelerator for HPC workloads, however its power consumption has become a major limiting factor for further scaling HPC systems. An accurate understanding of GPU power consumption is essential for further improving its energy efficiency, and consequently reducing the associated carbon footprint. Despite the limited documentation and lack of understanding, NVIDIA GPUs’ built-in power sensor is widely used in energy-efficient computing research. Our study seeks to elucidate the internal mechanisms of the power readings provided by nvidia-smi and assess the accuracy of the measurements. We evaluated over 70 different GPUs across 12 architectural generations, and identified several unforeseen problems that can lead to drastic under/overestimation of energy consumed, for example on the A100 and H100 GPUs only 25% of the runtime is sampled. We proposed several mitigations that could reduce the energy measurement error by an average of 35% in the test cases we present.
Karel Adámek, Wes Armour
SC3
2022 Intensity-Sensitive Similarity Indexes for Image Quality Assessment
abstract
The importance of Image quality assessment (IQA) is ever increasing due to the fast paced advances in imaging technology and computer vision. Among the numerous IQA methods, Structural SIMilarity (SSIM) index and its variants are better matched to the perceived quality of the human visual system. However, SSIM methods are insufficiently sensitive, when images contain low information, where the important information only occupies a low proportion of the image while most of the image is noise-like, which is common in scientific data. Therefore, we propose two new IQA methods, InTensity Weighted SSIM index and Low-Information Similarity Index, for such low information images. In addition, auxiliary indexes are proposed to assist with the assessment. The application of these new IQA methods to natural images and field-specific images, such as radio astronomical images, medical images, and remote sensing images, are also demonstrated. The results show that our IQA methods perform better than state-of-the-art SSIM methods for differences in high-intensity parts of the input images and have similar performance to that of the original and gradient-based SSIM for differences in low-intensity parts. Different similarity indexes are suitable for different applications, which we demonstrate in our results.
Wes Armour
ICPR2
2020 GPU Fast Convolution via the Overlap-and-Save Method in Shared Memory
abstract
We present an implementation of the overlap-and-save method, a method for the convolution of very long signals with short response functions, which is tailored to GPUs. We have implemented several FFT algorithms (using the CUDA programming language) which exploit GPU shared memory, allowing for GPU accelerated convolution. We compare our implementation with an implementation of the overlap-and-save algorithm utilizing the NVIDIA FFT library (cuFFT). We demonstrate that by using a shared memory based FFT we can achieved significant speed-ups for certain problem sizes and lower the memory requirements of the overlap-and-save method on GPUs.
Karel Adámek, Sofia Dimoudi, Michael B. Giles, Wes Armour
ACM Trans. Archit. Code Optim.4
2020 Evaluating Auto-Vectorizing Compilers through Objective Withdrawal of Useful Information
abstract
The need for compilers to generate highly vectorized code is at an all-time high with the increasing vectorization capabilities of modern processors. To this end, the information that compilers have at their disposal, either through code analysis or via user annotations, is instrumental for auto-vectorization, and hence for the overall performance. However, the information that is available to compilers at compile time and its accuracy varies greatly, as does the resulting performance of vectorizing compilers. Benchmarks like the Test Suite for Vectorizing Compilers (TSVC) have been developed to evaluate the vectorization capability of such compilers. The overarching approach of TSVC and similar benchmarks is to evaluate the compilers under the best possible scenario (i.e., assuming that compilers have access to all useful contextual information at compile time). Although this idealistic view is useful to observe the capability of compilers for auto-vectorization, it is not a true reflection of the conditions found in real-world applications. In this article, we propose a novel method for evaluating the auto-vectorization capability of compilers. Instead of assuming that compilers have access to a wealth of information at compile time, we formulate a method to objectively supply or withdraw information that would otherwise aid the compiler in the auto-vectorization process. This method is orthogonal to the approach adopted by TSVC, and as such, it provides the means of assessing the capabilities of modern vectorizing compilers in a more detailed way. Using this new method, we exhaustively evaluated five industry-grade compilers (GNU, Intel, Clang, PGI, and IBM) on four representative vector platforms (AVX-2, AVX-512 (Skylake), AVX-512 (KNL), and AltiVec) using the modified version of TSVC and application-level proxy kernels. The results show the impact that withdrawing information has on the vectorization capabilities of each compiler and also prove the validity of the presented technique.
Sergi Siso, Wes Armour, Jeyan Thiyagalingam
ACM Trans. Archit. Code Optim.2
2018 Building the World's Largest Radio Telescope: The Square Kilometre Array Science Data Processor
abstract
The Square Kilometre Array (SKA) will be the largest radio telescope constructed to date and the largest Big Data project in the known Universe. The first phase of the project will generate 160 terabytes every second. This amounts to 5 zettabytes (5 million petabytes) of data that will be generated by the facility each year - a data rate equivalent to 5 times the estimated global internet traffic in 2015. These data need to be reduced and then continuously ingested by the SKA Science Data Processor (SDP). Within the SDP Consortium, we are contributing to various roles in the development of the telescope including building a lightweight end-to-end prototype of the major components of the SDP system - a project we call the SDP Integration Prototype (SIP). The aim is to build a mini, fully-operational SDP, for which we have been developing realistic SKA-like science pipelines that can handle these unprecedented data volumes.
Jamie S. Farnes, Ben Mort, Fred Dulwich, Karel Adámek, Anna Brown, Jan Novotný, Stef Salvini, Wes Armour
eScience8
2005 QCDgrid: A Grid Resource for Quantum Chromodynamics
James Perry, Lorna Smith, A. N. Jackson, R. D. Kenway, B. Joo, Christopher M. Maynard, Arthur S. Trew, D. Byrne, George Beckett, C. T. H. Davies, S. Downing, A. C. Irving, Craig McNeile, Z. Sroczynski, C. R. Allton, Wes Armour, J. M. Flynn
J. Grid Comput.16