Márton Erdos

dblp:238/5530 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-5146-4361ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 The Future of Instruction-Level Parallelism (ILP)
abstract
High-performance processors have long used instruction-level parallelism (ILP) to achieve performance, but in the past decade processor vendors have dramatically increased their reliance upon this technique. We therefore take another look at the theoretical limits of ILP, in order to evaluate challenges and opportunities for processor architectures. Using the dynamic dependency graph of general-purpose workloads, we find that the upper bound on ILP is surprisingly close to the IPC capabilities of current state-of-the-art cores. Our results suggest that there may be as little as a decade of further scaling on current trends before hardware capabilities exceed the ILP bound.
Alexandra W. Chadwick, Márton Erdos, Utpal Bora 0003, Akshay Bhosale, Bob Lytton, Giacomo Gabrielli, Timothy M. Jones 0001
ISPASS2
2025 LoopFrog: In-Core Hint-Based Loop Parallelization
abstract
To scale ILP, designers build deeper and wider out-of-order superscalar CPUs.However, this approach incurs quadratic scaling complexity, area, and energy costs with each generation.While small loops may benefit from increased instruction-window sizes and large loops may see speedups via thread-level parallelism across cores, there remains unexploited medium-granularity parallelism.We propose LoopFrog to tap into this potential by bringing thread-level speculation schemes into the modern era.LoopFrog runs multiple loop iterations from a single thread in parallel within the microarchitecture.The core can spawn future loop iterations as new microarchitectural threadlets based on compiler-inserted hints, which can leapfrog execution beyond the parent thread's instruction window, exposing a new, medium-grained parallelism, orthogonal to traditional ILP and TLP.LoopFrog monitors data dependencies between executing threadlets, forwards data for true dependencies and squashes speculative threadlets on ordering violations.Using an LLVM-based compiler to insert hints, we achieve a geometric mean loop speedup of 43%, translating to whole-program speedups of 9.2% on SPEC CPU 2006 and 9.5% on SPEC CPU 2017 benchmarks, with only modest area and power overheads.
Márton Erdos, Utpal Bora 0001, Akshay Bhosale, Bob Lytton, Ali Mustafa Zaidi, Alexandra W. Chadwick, Giacomo Gabrielli, Timothy M. Jones 0001
MICRO1
2025 Ghost Threading: Helper-Thread Prefetching for Real Systems
Akshay Bhosale, Utpal Bora 0001, Alexandra W. Chadwick, Márton Erdos, Giacomo Gabrielli, Timothy M. Jones 0001
MICRO5
2024 OptiWISE: Combining Sampling and Instrumentation for Granular CPI Analysis
abstract
Despite decades of improvement in compiler technology, it remains necessary to profile applications to improve performance. Existing profiling tools typically either sample hardware performance counters or instrument the program with extra instructions to analyze its execution. Both techniques are valuable with different strengths and weaknesses, but do not always correctly identify optimization opportunities. We present OPTIWISE, a profiling tool that runs the program twice, once with low-overhead sampling to accurately measure performance, and once with instrumentation to accurately capture control flow and execution counts. OPTIWISE then combines this information to give a highly detailed per-instruction CPI metric by computing the ratio of samples to execution counts, as well as aggregated information such as costs per loop, source-code line, or function. We evaluate OPTIWISE to show it has an overhead of 8.1× geomean, and 57× worst case on SPEC CPU2017 benchmarks. Using OPTIWISE, we present case studies of optimizing selected SPEC benchmarks on a modern x86 server processor. The per-instruction CPI metrics quickly reveal problems such as costly mispredicted branches and cache misses, which we use to manually optimize for effective performance improvements.
Alexandra W. Chadwick, Márton Erdos, Utpal Bora 0003, Ilias Vougioukas, Giacomo Gabrielli, Timothy M. Jones 0001
CGO3
2022 MineSweeper: a "clean sweep" for drop-in use-after-free prevention
abstract
Low-level languages, which require manual memory management from the programmer, remain in wide use for performance-critical applications. Memory-safety bugs are common, and now a major source of exploits. In particular, a use-after-free bug occurs when an object is erroneously deallocated, whilst pointers to it remain active in memory, and those (dangling) pointers are later used to access the object. An attacker can reallocate the memory area backing an erroneously freed object, then overwrite its contents, injecting carefully chosen data into the host program, thus altering its execution and achieving privilege escalation.
Márton Erdos, Sam Ainsworth 0001, Timothy M. Jones 0001
ASPLOS1
2019 The janus triad: exploiting parallelism through dynamic binary modification
abstract
We present a unified approach for exploiting thread-level, data-level, and memory-level parallelism through a same-ISA dynamic binary modifier guided by static binary analysis. A static binary analyser first examines an executable and determines the operations required to extract parallelism at runtime, encoding them as a series of rewrite rules that a dynamic binary modifier uses to perform binary transformation. We demonstrate this framework by exploiting three different kinds of parallelism to perform automatic vectorisation, software prefetching, and automatic parallelisation together on legacy application binaries. Software prefetch insertion alone achieves an average speedup of 1.2x, comparing favourably with an automatic compiler pass. Automatic vectorisation brings speedups of 2.7x on the TSVC benchmarks, significantly beating a compiler approach for some workloads. Finally, combining prefetching, vectorisation, and parallelisation realises a speedup of 3.8x on a representative application loop.
Ruoyu Zhou, George Wort, Márton Erdos, Timothy M. Jones 0001
VEE3