Amanieu D'Antras

dblp:178/1281 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
3 papers
Runtime systems and virtual machines · 100%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Processor architecture and microarchitecture · 87% Embedded and real-time systems · 13%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Runtime systems and virtual machines › binary translation
dynamic binary translation
0.832017
Low overhead dynamic binary translation on ARM · PLDI 2017
MAMBO: A Low-Overhead Dynamic Binary Modification Tool for ARM · ACM Trans. Archit. Code Optim. 2016
Optimizing Indirect Branches in Dynamic Binary Translators · ACM Trans. Archit. Code Optim. 2016
Processor architecture and microarchitecture
branch prediction
0.212016
Optimizing Indirect Branches in Dynamic Binary Translators · ACM Trans. Archit. Code Optim. 2016
Processor architecture and microarchitecture › instruction set architecture › RISC
ARMv8
0.112017
Low overhead dynamic binary translation on ARM · PLDI 2017
Processor architecture and microarchitecture
instruction set architecture
0.112017
Low overhead dynamic binary translation on ARM · PLDI 2017
Processor architecture and microarchitecture › instruction set architecture › RISC
ARM architecture
0.112016
MAMBO: A Low-Overhead Dynamic Binary Modification Tool for ARM · ACM Trans. Archit. Code Optim. 2016
Embedded and real-time systems
embedded processor
0.112016
MAMBO: A Low-Overhead Dynamic Binary Modification Tool for ARM · ACM Trans. Archit. Code Optim. 2016

Methods — techniques the papers use, named apart from their topics

software return address stack · 0.5profiling · 0.5hash table · 0.5
YearPublicationVenuePosition
2018 Optimising Dynamic Binary Modification Across ARM Microarchitectures
abstract
Dynamic Binary Modification (DBM) is a technique for modifying applications transparently while they are executed, working at the level of native code. However, DBM introduces a performance overhead, which in some cases can dominate execution time, making many uses impractical. The ARM hardware ecosystem poses unique challenges for high performance DBM systems because of the large number and wide range of capabilities of the commercially available implementations: from single issue, in order cores up to 6-issue out-of-order cores and including less traditional implementations. These variations raise the question of whether it is possible to develop DBM optimisations which either improve or, at the very least, do not affect performance on all available systems and microarchitectures. To answer this question, the performance of three new optimisations for the MAMBO DBM system has been evaluated on five systems using different microarchitectures. For comparison, the overhead of DynamoRIO, a high performance DBM system which was recently ported to the ARM architecture, is also evaluated.
Cosmin Gorgovan, Amanieu D'Antras, Mikel Luján
ICPE2
2017 Low overhead dynamic binary translation on ARM
abstract
The ARMv8 architecture introduced AArch64, a 64-bit execution mode with a new instruction set, while retaining binary compatibility with previous versions of the ARM architecture through AArch32, a 32-bit execution mode. Most hardware implementations of ARMv8 processors support both AArch32 and AArch64, which comes at a cost in hardware complexity.
Amanieu D'Antras, Cosmin Gorgovan, Jim D. Garside, Mikel Luján
PLDI1
2017 HyperMAMBO-X64: Using Virtualization to Support High-Performance Transparent Binary Translation
abstract
Current computer architectures --- ARM, MIPS, PowerPC, SPARC, x86 --- have evolved from a 32-bit architecture to a 64-bit one. Computer architects often consider whether it could be possible to eliminate hardware support for a subset of the instruction set as to reduce hardware complexity, which could improve performance, reduce power usage and accelerate processor development. This paper considers the scenario where we want to eliminate 32-bit hardware support from the ARMv8 architecture.
Amanieu D'Antras, Cosmin Gorgovan, Jim D. Garside, John Goodacre, Mikel Luján
VEE1
2016 Optimizing Indirect Branches in Dynamic Binary Translators
abstract
Dynamic binary translation is a technology for transparently translating and modifying a program at the machine code level as it is running. A significant factor in the performance of a dynamic binary translator is its handling of indirect branches. Unlike direct branches, which have a known target at translation time, an indirect branch requires translating a source program counter address to a translated program counter address every time the branch is executed. This translation can impose a serious runtime penalty if it is not handled efficiently. MAMBO-X64, a dynamic binary translator that translates 32-bit ARM (AArch32) code to 64-bit ARM (AArch64) code, uses three novel techniques to improve the performance of indirect branch translation. Together, these techniques allow MAMBO-X64 to achieve a very low performance overhead of only 10% on average compared to native execution of 32-bit programs. Hardware-assisted function returns use a software return address stack to predict the targets of function returns, making use of several novel optimizations while also exploiting hardware return address prediction. This technique has a significant impact on most benchmarks, reducing binary translation overhead compared to native execution by 40% on average and by 90% on some benchmarks. Branch table inference , an algorithm for detecting and translating branch tables, can reduce the overhead of translated code by up to 40% on some SPEC CPU2006 benchmarks. The remaining indirect branches are handled using a fast atomic hash table , which is optimized to work with multiple threads. This last technique translates indirect branches using a single shared hash table while avoiding expensive synchronization in performance-critical lookup code. This allows the performance to be on par with thread-private hash tables while having superior memory scalability.
Amanieu D'Antras, Cosmin Gorgovan, Jim D. Garside, Mikel Luján
ACM Trans. Archit. Code Optim.1
2016 MAMBO: A Low-Overhead Dynamic Binary Modification Tool for ARM
abstract
As the ARM architecture expands beyond its traditional embedded domain, there is a growing interest in dynamic binary modification (DBM) tools for general-purpose multicore processors that are part of the ARM family. Existing DBM tools for ARM suffer from introducing large overheads in the execution of applications. The specific questions that this article addresses are (i) how to develop such DBM tools for the ARM architecture and (ii) whether new optimisations are plausible and needed. We describe the general design of MAMBO, a new DBM tool for ARM, which we release together with this publication, and introduce novel optimisations to handle indirect branches. In addition, we explore scenarios in which it may be possible to relax the transparency offered by DBM tools to allow extra optimisations to be applied. These scenarios arise from analysing the most typical usages: for example, application binaries without handcrafted assembly. The performance evaluation shows that MAMBO introduces small overheads for SPEC CPU2006 and PARSEC 3.0 when comparing with the execution times of the unmodified programs: a geometric mean overhead of 28% on a Cortex-A9 and of 34% on a Cortex-A15 for CPU2006, and between 27% and 32%, depending on the number of threads, for PARSEC on a Cortex-A15.
Cosmin Gorgovan, Amanieu D'Antras, Mikel Luján
ACM Trans. Archit. Code Optim.2