Igor Böhm

dblp:02/8025 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
0since 2021 · last 2020
0009-0009-5723-7811ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3Software engineering, systems software and programming languages · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Runtime systems and virtual machines · 65% Concurrent programming · 35%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Runtime systems and virtual machines › binary translation
dynamic binary translation
0.622020
Fast and Correct Load-Link/Store-Conditional Instruction Handling in DBT Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Generalized just-in-time trace compilation using a parallel task farm in a dynamic binary translator · PLDI 2011
Concurrent programming › synchronization
synchronization primitives
0.412020
Fast and Correct Load-Link/Store-Conditional Instruction Handling in DBT Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Parallel and multicore computing › transactional memory
hardware transactional memory
0.412020
Fast and Correct Load-Link/Store-Conditional Instruction Handling in DBT Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation
0.112011
Generalized just-in-time trace compilation using a parallel task farm in a dynamic binary translator · PLDI 2011
Runtime systems and virtual machines › dynamic compilation › just-in-time compilation
trace-based compilation
0.112011
Generalized just-in-time trace compilation using a parallel task farm in a dynamic binary translator · PLDI 2011
Parallel and multicore computing
parallelizing compiler
0.112011
Generalized just-in-time trace compilation using a parallel task farm in a dynamic binary translator · PLDI 2011

Methods — techniques the papers use, named apart from their topics

page translation cache · 0.9hardware transactional memory · 0.9compare-and-swap emulation · 0.9task farm · 0.2dynamic work scheduling · 0.2
YearPublicationVenuePosition
2020 Fast and Correct Load-Link/Store-Conditional Instruction Handling in DBT Systems
abstract
Dynamic binary translation (DBT) requires the implementation of load-link/store-conditional (LL/SC) primitives for guest systems that rely on this form of synchronization. When targeting, e.g., ×86 host systems, LL/SC guest instructions are typically emulated using atomic compare-and-swap (CAS) instructions on the host. Whilst this direct mapping is efficient, this approach is problematic due to subtle differences between LL/SC and CAS semantics. In this article, we demonstrate that this is a real problem, and we provide code examples that fail to execute correctly on QEMU and a commercial DBT system, which both use the CAS approach to LL/SC emulation. We then develop two novel and provably correct LL/SC emulation schemes: 1) a purely software-based scheme, which uses the DBT system's page translation cache for correctly selecting between fast, but unsynchronized, and slow, but fully synchronized memory accesses and 2) a hardware-accelerated scheme that leverages hardware transactional memory (HTM) provided by the host. We have implemented these two schemes in the Synopsys DesignWare ARC nSIM DBT system, and we evaluate our implementations against full applications, and targeted microbenchmarks. We demonstrate that our novel schemes are not only correct but also deliver competitive performance on-par or better than the widely used, but broken CAS scheme.
Martin Kristien, Tom Spink, Brian Campbell 0001, Susmit Sarkar, Ian Stark, Björn Franke, Igor Böhm, Nigel P. Topham
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2019 Mitigating JIT compilation latency in virtual execution environments
abstract
Many Virtual Execution Environments (VEEs) rely on Just-in-time (JIT) compilation technology for code generation at runtime, e.g. in Dynamic Binary Translation (DBT) systems or language Virtual Machines (VMs). While JIT compilation improves native execution performance as opposed to e.g. interpretive execution, the JIT compilation process itself introduces latency. In fact, for highly optimizing JIT compilers or compilers not specifically designed for JIT compilation, e.g. LLVM, this latency can cause a substantial overhead. While existing work has introduced asynchronously decoupled JIT compilation task farms to hide this JIT compilation latency, we show that this on its own is not sufficient to mitigate the impact of JIT compilation latency on overall performance. In this paper, we introduce a novel JIT compilation scheduling policy, which performs continuous low-cost profiling of code regions already dispatched for JIT compilation, right up to the point where compilation commences. We have integrated our novel JIT compilation scheduling approach into a commercial LLVM-based DBT system and demonstrate speedups of 1.32x on average, and up to 2.31x, over its state-of-the-art concurrent task-farm based JIT compilation scheme across the SPEC CPU2006 and BioPerf benchmark suites.
Martin Kristien, Tom Spink, Harry Wagstaff, Björn Franke, Igor Böhm, Nigel P. Topham
VEE5
2012 Efficiently parallelizing instruction set simulation of embedded multi-core processors using region-based just-in-time dynamic binary translation
abstract
Embedded systems, as typified by modern mobile phones, are already seeing a drive toward using multi-core processors. The number of cores will likely increase rapidly in the future. Engineers and researchers need to be able to simulate systems, as they are expected to be in a few generations time, running simulations of many-core devices on today's multi-core machines. These requirements place heavy demands on the scalability of simulation engines, the fastest of which have typically evolved from just-in-time (Jit) dynamic binary translators (Dbt).
Stephen C. Kyle, Igor Böhm, Björn Franke, Hugh Leather, Nigel P. Topham
LCTES2
2011 Generalized just-in-time trace compilation using a parallel task farm in a dynamic binary translator
abstract
Dynamic Binary Translation (DBT) is the key technology behind cross-platform virtualization and allows software compiled for one Instruction Set Architecture (ISA) to be executed on a processor supporting a different ISA. Under the hood, DBT is typically implemented using Just-In-Time (JIT) compilation of frequently executed program regions, also called traces. The main challenge is translating frequently executed program regions as fast as possible into highly efficient native code. As time for JIT compilation adds to the overall execution time, the JIT compiler is often decoupled and operates in a separate thread independent from the main simulation loop to reduce the overhead of JIT compilation. In this paper we present two innovative contributions. The first contribution is a generalized trace compilation approach that considers all frequently executed paths in a program for JIT compilation, as opposed to previous approaches where trace compilation is restricted to paths through loops. The second contribution reduces JIT compilation cost by compiling several hot traces in a concurrent task farm. Altogether we combine generalized light-weight tracing, large translation units, parallel JIT compilation and dynamic work scheduling to ensure timely and efficient processing of hot traces. We have evaluated our industry-strength, LLVM-based parallel DBT implementing the ARCompact ISA against three benchmark suites (EEMBC, BioPerf and SPEC CPU2006) and demonstrate speedups of up to 2.08 on a standard quad-core Intel Xeon machine. Across short- and long-running benchmarks our scheme is robust and never results in a slowdown. In fact, using four processors total execution time can be reduced by on average 11.5% over state-of-the-art decoupled, parallel (or asynchronous) JIT compilation.
Igor Böhm, Tobias J. K. Edler von Koch, Stephen C. Kyle, Björn Franke, Nigel P. Topham
PLDI1
2010 Integrated instruction selection and register allocation for compact code generation exploiting freeform mixing of 16- and 32-bit instructions
abstract
For memory constrained embedded systems code size is at least as important as performance. One way of increasing code density is to exploit compact instruction formats, e.g. ARM Thumb, where the processor either operates in standard or compact instruction mode. The ARCompact ISA considered in this paper is different in that it allows freeform mixing of 16- and 32-bit instructions without a mode switch. Compact 16-bit instructions can be used anywhere in the code given that additional register constraints are satisfied. In this paper we present an integrated instruction selection and register allocation methodology and develop two approaches for mixed-mode code generation: a simple opportunistic scheme and a more advanced feedback-guided instruction selection scheme. We have implemented a code generator targeting the ARCompact ISA and evaluated its effectiveness against the ARC750D embedded processor and the EEMBC benchmark suite. On average, we achieve a code size reduction of 16.7% across all benchmarks whilst at the same time improving performance by on average 17.7%.
Tobias J. K. Edler von Koch, Igor Böhm, Björn Franke
CGO2