Joshua Randall 0001

dblp:207/2235 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0002-5154-8688ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Ember: A Compiler for Embedding Operations on Decoupled Access-Execute Architectures
abstract
Decoupled Access-Execute (DAE) architectures separate memory accesses from computation in two specialized units. This design is becoming increasingly popular among hyperscalers to accelerate irregular embedding lookups in recommendation models. In this paper, we first broaden the scope by demonstrating the benefits of DAE architectures across a wider range of irregular embedding operations in several machine learning models. Then, we propose the Ember compiler to automatically compile all of these embedding operations to DAE architectures. Conversely from other DAE compilers, Ember features multiple intermediate representations specifically designed for different optimization levels. In this way, Ember can implement all optimizations to match the performance of hand-written code, unlocking the full potential of DAE architectures at scale.
Marco Siracusa, Olivia Hsu, Víctor Soria 0001, Joshua Randall 0001, Arnaud Grasset, Eric Biscondi, Douglas J. Joseph, Randy Allen, Fredrik Kjolstad, Miquel Moretó, Adrià Armejach
CGO4
2026 Multi-Level TDA for Heterogeneous Arm Processors Using a Standardized Telemetry Framework
abstract
Performance analysis on modern superscalar processors is challenging due to complex pipelines in which hardware telemetry rarely provides actionable, easily interpretable insights. Software developers and system engineers therefore need quick, reliable, and portable analysis methods to tune software and hardware for efficient workload execution. This motivates the need for a standardized methodology, with metrics and events that have uniform interpretation across cores and platforms. This paper presents the first unified Top-Down Microarchitecture Analysis (TDA) solution for Arm CPUs, and the first to enable full TDA support on heterogeneous mobile-class platforms. The solution defines a complete four-level hierarchy of metrics derived from native hardware Performance Monitoring Unit (PMU) events built on explicit stall accounting. Evaluations on three Arm cores in a single SoC using Geekbench 6 workloads and the targeted ustress validation suite demonstrate the portability of the approach and the accuracy of its bottleneck attribution, exercising all TDA hierarchy levels to produce clear and consistent results. We further introduce the Arm telemetry framework and Arm Top-Down Tool, which use a standardized machine-readable JSON schema for multi-OS data collection and metric derivation. This approach ensures that metrics remain consistent across products, operating systems, and core configurations. Finally, we present an optimization case study on both mobile and serverclass Arm platforms, demonstrating applicability from edge to cloud deployments. Our solution provides a unified approach for cross-core, cross-generation, and cross-platform workload characterization and enables efficient, reproducible performance analysis on Arm-based devices.
Jumana Mundichipparakkal, Darin Greene, Joshua Randall 0001, Vincent Coubard, Michael Williams, René de Jong
ISPASS3
2023 A Tensor Marshaling Unit for Sparse Tensor Algebra on General-Purpose Processors
abstract
This paper proposes the Tensor Marshaling Unit (TMU), a near-core programmable dataflow engine for multicore architectures that accelerates tensor traversals and merging, the most critical operations of sparse tensor workloads running on today’s computing infrastructures. The TMU leverages a novel multi-lane design that enables parallel tensor loading and merging, which naturally produces vector operands that are marshaled into the core for efficient SIMD computation. The TMU supports all the necessary primitives to be tensor-format and tensor-algebra complete. We evaluate the TMU on a simulated multicore system using a broad set of tensor algebra workloads, achieving 3.6 ×, 2.8 ×, and 4.9 × speedups over memory-intensive, compute-intensive, and merge-intensive vectorized software implementations, respectively.
Marco Siracusa, Víctor Soria 0001, Francesco Sgherzi, Joshua Randall 0001, Douglas J. Joseph, Miquel Moretó, Adrià Armejach
MICRO4