EDBT 2026 Demo / reviewers in the wild / expert
Florian Huemer
dblp:173/7149
· DBLP profile ↗
9ranked-venue papers
6as first author
5since 2021 · last 2026
0000-0002-2776-7768ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Partial-Evaluation Templates: Accelerating Partial Evaluation with Pre-compiled TemplatesabstractPartial evaluation allows specializing programs by substituting some of their inputs with static values and evaluating the resulting static expressions. When applied to a language interpreter, it specializes the interpreter to a fixed source program, resulting in a specialized interpreter that existing compilers can optimize and generate efficient machine code for. Although this approach achieves good peak performance in just-in-time compilers, partial evaluation is costly with respect to compilation time. In the case of bytecode interpreters, partial evaluation must process and transform each bytecode instruction. Since instructions with the same opcodes share common functionality, partial evaluation applies the same transformations repeatedly. Consequently, partial evaluation significantly increases the compilation time of just-in-time compilation.We propose partial-evaluation templates, i.e., reusable collections of ahead-of-time-compiled graphs that are generated by a hand-written cogen approach. The templates are parametric in their static inputs, and cover common functionality in a bytecode interpreter implementation. At compile time, templates specialize themselves by substituting their static inputs, reducing themselves to a single pre-compiled graph that the compiler inlines and optimizes. This reduces the size of the intermediate representation and speeds up partial evaluation, as well as optimizations.We integrated our approach into GraalVM, a state-of-the-art high-performance polyglot virtual machine, and GraalWasm, a bytecode-interpreter-based WebAssembly runtime. We extended the existing GraalVM compiler with a binding-time analysis that generates the templates and introduced a specializer for the templates into the existing partial evaluator. We enabled template generation for nearly all opcodes in GraalWasm, which reduced partial-evaluation time by up to 36 % and warmup time by up to 17 % without impacting peak performance. Florian Huemer, Aleksandar Prokopec, David Leopoldseder, Raphael Mosaner, Hanspeter Mössenböck |
CGO | 1 |
| 2024 | QDI Binary Comparator Networks and their Application in Combinational LogicabstractBinary comparator networks have already been used in quasi delay-insensitive (QDI) circuit design for the construction of completion detectors. This paper demonstrates how they can also be utilized to implement a wide range of Boolean functions for combinational QDI logic blocks. Designing combinational logic for QDI circuits poses several challenges because data must be processed in an encoded form (usually dual-rail) and it must be ensured that the resulting circuits do not contain hazards or orphans. These design constraints impose a considerable hardware overhead on the resulting circuits. Hence, over the years numerous design and optimization strategies have been proposed. We show that our comparator-network-based construction approach yields promising results for certain types of functions compared to other common QDI design styles. Florian Huemer |
DDECS | 1 |
| 2024 | Taking a Closer Look: An Outlier-Driven Approach to Compilation-Time Optimization
Florian Huemer, David Leopoldseder, Aleksandar Prokopec, Raphael Mosaner, Hanspeter Mössenböck |
ECOOP | 1 |
| 2022 | On SAT-Based Model Checking of Speed-Independent CircuitsabstractFormal verification plays an important role in the quality assurance of digital circuits. Apart from the now standard equivalence checking between design steps, functional correctness can be proven with model checking. In one approach, a Boolean satisfiability (SAT) problem describing the circuit’s implementation and expected properties is generated for each of a bounded number of time steps and fed to a SAT solver. In synchronous circuits, the time steps correspond to cycles of the global clock. The execution of asynchronous, specifically speed-independent (SI) circuits, however, relies on local handshakes instead of a global time reference. This absence of a global clock requires a different approach for choosing time steps for the SAT problem.This paper presents how bounded, SAT-based model checking can be used on SI asynchronous circuits. We aim to give a general and accessible introduction to this topic, highlight the inherent computational complexity and show that setting up a basic model checker for SI circuits is possible with quite simple means, without any reliance on (expensive) commercial tools. For our reference implementation used in the provided examples we use the open source Z3 solver. Florian Huemer, Robert Najvirt, Andreas Steininger |
DDECS | 1 |
| 2021 | An Automated Setup for Large-Scale Simulation-Based Fault-Injection Experiments on Asynchronous Digital CircuitsabstractExperimental fault injection is an essential tool in the assessment and verification of fault-tolerance properties. Often, in these experiments it is impossible to reasonably cover the huge parameter space spanned by target state and fault parameters, and compromises or restrictions must be made. This is even more pronounced for asynchronous circuits where a convenient discretization of time through a synchronous clock is not possible. In this paper we present a fault-injection toolset that allows for a very efficient injection and data processing, thus bringing studies with many billions of meaningful injections into asynchronous targets within reach. The key ingredients of our solution are an auto-setup feature capable of optimizing parameter values, seamless distribution of the simulation load to many host computers, and efficient arrangement of the important settings and readings in a database. We will use the example of a comparative study of different asynchronous pipeline styles to motivate the need for such an approach and illustrate its benefits. Patrick Behal, Florian Huemer, Robert Najvirt, Andreas Steininger |
DSD | 2 |
| 2018 | Using a Duplex Time-to-Digital Converter for Metastability Characterization of an FPGAabstractIn view of the increasing number of clock domains found in modern ASICs, the precise characterization of metastability at their boundaries becomes crucial. In some cases, the conventional approach does not provide a sufficient level of detail information. As an alternative approach, the use of a time-to-digital converter based on a tapped delay line has been proposed. In this paper we extend the latter by an additional tapped delay line thus allowing to further refine the concept. We present the underlying concept, its implementation, as well as experimental measurements on an FPGA platform that reveals significant variations in the metastable behavior of different FPGA boards of the same type. Florian Huemer, Thomas Polzer, Andreas Steininger |
DDECS | 1 |
| 2017 | Measuring metastability using a time-to-digital converterabstractIn view of the numerous clock domain crossings found in modern systems-on-chip and multicore architectures precise metastability characterization is a fundamental task. We propose a conceptually novel approach for the experimental assessment of upset rate over resolution time that is usually employed to extract the relevant characteristics. Our method is based on connecting a time-to-digital converter to the output of the flip flop under test, rather than using a phase shifted clock, as conventionally done. We present the details of an FPGA implementation of our approach and show its feasibility through an experimental evaluation, whose results favorably match those obtained by the conventional method. The benefits of the novel scheme are the ability to perform a calibration for the delay steps, a speed-up of the measurement process, and the availability of a more comprehensive and ordered measurement data set. Thomas Polzer, Florian Huemer, Andreas Steininger |
DDECS | 2 |
| 2016 | A new coding scheme for fault tolerant 4-phase delay-insensitive codesabstractDelay-insensitive (DI) codes are usually prone to transient faults occurring during an ongoing transmission. For most DI code words even a single transient can turn an incomplete transmission into a complete code word, which is different from the originally sent one. In this paper we therefore propose a novel two-step data encoding scheme that combines DI and error detecting codes. Our solution exploits the inherent fault resilience of DI codes to achieve a low coding overhead. For analyzing this fault resilience we use methods from graph theory. In contrast to existing approaches we carefully avoid the introduction of timing assumptions to mask faults. The presented coding scheme is generic and can, in principle, be used with any 4-phase DI code. We show how to apply it to m-of-n codes and analyze the resulting coding efficiency. Florian Huemer, Jakob Lechner, Andreas Steininger |
ICCD | 1 |
| 2015 | Methods for analysing and improving the fault resilience of delay-insensitive codesabstractDelay-insensitive (DI) codes are usually prone to transient faults occurring during an ongoing transmission. For most DI codewords even a single transient can turn an incomplete transmission into a complete codeword, which is different from the originally sent codeword. Unless further redundant information is provided, the receiver has no means to detect such a transmission fault. In this paper we therefore propose two methods to systematically increase redundancy, either by i) building resilient subcodes, or by ii) using a two-step data encoding where error detecting codes are appropriately combined with delay-insensitive codes. In contrast to existing approaches we carefully avoid the introduction of timing assumptions to mask faults. Both methods are generic and can be used for any 4-phase DI code. In this paper we apply them to m-of-n codes, Berger and Zero-Sum codes and thoroughly analyse the efficiency of the resulting coding schemes. Jakob Lechner, Andreas Steininger, Florian Huemer |
ICCD | 3 |