Leighton Wilson

dblp:260/0797 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0003-1676-8156ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 45% Hardware accelerators and domain-specific architectures · 45% Performance modeling and evaluation · 5%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
collective communication
0.812024
Near-Optimal Wafer-Scale Reduce · HPDC 2024
Hardware accelerators and domain-specific architectures
wafer-scale architecture
0.812024
Near-Optimal Wafer-Scale Reduce · HPDC 2024
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.712023
Scaling the "Memory Wall" for Multi-Dimensional Seismic Processing with Algebraic Compression on Cerebras CS-2 Systems · SC 2023
High-performance computing
scientific computing systems
0.712023
Scaling the "Memory Wall" for Multi-Dimensional Seismic Processing with Algebraic Compression on Cerebras CS-2 Systems · SC 2023
High-performance computing › scientific computing
seismic processing
0.712023
Scaling the "Memory Wall" for Multi-Dimensional Seismic Processing with Algebraic Compression on Cerebras CS-2 Systems · SC 2023
Hardware accelerators and domain-specific architectures › many-core accelerator
wafer-scale engine
0.712023
Scaling the "Memory Wall" for Multi-Dimensional Seismic Processing with Algebraic Compression on Cerebras CS-2 Systems · SC 2023
Performance modeling and evaluation › performance prediction
execution time prediction
0.212024
Near-Optimal Wafer-Scale Reduce · HPDC 2024
Memory systems
memory bandwidth
0.212023
Scaling the "Memory Wall" for Multi-Dimensional Seismic Processing with Algebraic Compression on Cerebras CS-2 Systems · SC 2023

Methods — techniques the papers use, named apart from their topics

lower bound analysis · 0.8automatic code generation · 0.8tile low-rank matrix-vector multiplication · 0.7algebraic compression · 0.7
YearPublicationVenuePosition
2025 Slide FFT on a homogeneous mesh in wafer-scale computing
abstract
Abstract Searches for signals at low signal-to-noise ratios frequently involve correlations evaluated by the Fast Fourier Transform (FFT). To accelerate the discovery power of present and next-generation multi-messenger observatories, we here explore the implementation of FFT on wafer-scale engines. To minimize the memory overhead of the inherently non-local FFT algorithm on a homogeneous mesh of Processing Elements (PEs) with no global memory on the chip, we introduce a new synchronous slide operation (Slide) exploiting fast interconnect between adjacent PEs. The feasibility of compute-limited performance is demonstrated in linear scaling of Slide execution times with varying array sizes in preliminary benchmarks on the CS-2 WSE. As a first step, this benchmark appears promising for the proposed implementation of high-throughput FFT-based signal processing in multi-messenger astronomy.
Maurice H. P. M. van Putten, Leighton Wilson, Adam W. Lavely, Mark Hair
Discov. Comput.2
2024 Near-Optimal Wafer-Scale Reduce
abstract
Efficient Reduce and AllReduce communication collectives are a critical cornerstone of high-performance computing (HPC) applications. We present the first systematic investigation of Reduce and AllReduce on the Cerebras Wafer-Scale Engine (WSE). This architecture has been shown to achieve unprecedented performance both for machine learning workloads and other computational problems like FFT. We introduce a performance model to estimate the execution time of algorithms on the WSE and validate our predictions experimentally for a wide range of input sizes. In addition to existing implementations, we design and implement several new algorithms specifically tailored to the architecture. Moreover, we establish a lower bound for the runtime of a Reduce operation on the WSE. Based on our model, we automatically generate code that achieves near-optimal performance across the whole range of input sizes. Experiments demonstrate that our new Reduce and AllReduce algorithms outperform the current vendor solution by up to 3.27×. Additionally, our model predicts performance with less than 4% error. The proposed communication collectives increase the range of HPC applications that can benefit from the high throughput of the WSE. Our model-driven methodology demonstrates a disciplined approach that can lead the way to further algorithmic advancements on wafer-scale architectures.
Piotr Luczynski, Lukas Gianinazzi, Patrick Iff, Leighton Wilson, Daniele De Sensi, Torsten Hoefler
HPDC4
2023 Scaling the "Memory Wall" for Multi-Dimensional Seismic Processing with Algebraic Compression on Cerebras CS-2 Systems
abstract
We exploit the high memory bandwidth of AI-customized Cerebras CS-2 systems for seismic processing. By leveraging low-rank matrix approximation, we fit memory-hungry seismic applications onto memory-austere SRAM wafer-scale hardware, thus addressing a challenge arising in many wave-equation-based algorithms that rely on Multi-Dimensional Convolution (MDC) operators. Exploiting sparsity inherent in seismic data in the frequency domain, we implement embarrassingly parallel tile low-rank matrix-vector multiplications (TLR-MVM), which account for most of the elapsed time in MDC operations, to successfully solve the Multi-Dimensional Deconvolution (MDD) inverse problem. By reducing memory footprint along with arithmetic complexity, we fit a standard seismic benchmark dataset into the small local memories of Cerebras processing elements. Deploying TLR-MVM execution onto 48 CS-2 systems in support of MDD gives a sustained memory bandwidth of 92.58PB/s on 35, 784, 000 processing elements, a significant milestone that highlights the capabilities of AI-customized architectures to enable a new generation of seismic algorithms that will empower multiple technologies of our low-carbon future.
Hatem Ltaief, Yuxi Hong 0001, Leighton Wilson, Mathias Jacquelin, Matteo Ravasi, David E. Keyes
SC3