VLDB 2026 Research / reviewers in the wild / expert
Simon Rokicki
dblp:199/8680
· DBLP profile ↗
20ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-0195-096XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 6 first-author · 7 since 2021Software engineering, systems software and programming languages · 8 · 4 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automatic Extraction of Timing Models for WCET Estimation From a High-Level Synthesis FlowabstractReal-time, domain-specific processors require faithful timing models for WCET analysis. However, existing models are typically hand-crafted from sparse documentation, making them error-prone and difficult to maintain. This work aims to automatically extract WCET timing models from single-issue in-order processor pipelines generated by High-Level Synthesis (HLS). By deriving timing models directly from the SpecHLS intermediate representation, the models are faithful by construction. Experimental results show that our timing-model extraction process generalizes across diverse RISC-V core variants and yields WCET estimates within 0.48% on average of those from a handcrafted model, on the Mälardalen WCET benchmarks. Thomas Feuilletin, Dylan Leothaud, Simon Rokicki, Steven Derrien, Isabelle Puaut |
DATE | 3 |
| 2026 | Area Efficient Speculative Loop Pipelining for High-Level SynthesisabstractHigh-Level Synthesis (HLS) allows the automatic generation of efficient circuit designs for computation-intensive kernels, but it lacks flexibility when dealing with irregular control flow. Dynamic and speculative HLS techniques are used to address this issue. These techniques outperform state-of-the-art HLS in kernel execution times but introduce a significant area overhead. In contrast, state-of-the-art HLS easily highlights and exploits resource-sharing opportunities. In this work, we show how to adapt an existing speculative HLS approach to take advantage of well-known static resource sharing mechanisms. Our results show a decrease of the area cost by 34% on average. Dylan Leothaud, Simon Rokicki, Steven Derrien, Isabelle Puaut |
DATE | 2 |
| 2026 | WCET Analysis of HLS-Generated Processors Using Abstract InterpretationabstractDeriving sound and precise timing models remains one of the main obstacles to static Worst-Case Execution Time (WCET) analysis. Modern processors exhibit diverse and evolving microarchitectures, making manual construction of timing models labor-intensive, error-prone, and difficult to adapt across processor variants. High-Level Synthesis (HLS) enables rapid customization of processor cores and architectural exploration, offering an opportunity to automate not only hardware generation but also the derivation of associated timing models. This paper presents an automated WCET analysis for HLS-generated processors based on abstract interpretation. We exploit the internal Gated-SSA representation of the HLS flow to automatically extract an abstract timing model capturing speculation and stall mechanisms. WCET estimation at the basic block level is then formulated as an exploration of abstract microarchitectural states within a basic block. The approach safely accounts for timing anomalies, while remaining scalable thanks to an efficient state-merging strategy. Integrated into the Heptane WCET tool and evaluated on Mälardalen benchmarks and a RISC-V, the method achieves the same tightness as a handcrafted timing model, while improving over a previously proposed automated approach. Thomas Feuilletin, Dylan Leothaud, Simon Rokicki, Steven Derrien, Isabelle Puaut |
ECRTS | 3 |
| 2025 | Exploring Speculation Barriers for RISC-V Selective Speculation
Herinomena Andrianatrehina, Ronan Lashermes, Joseph Paturel, Simon Rokicki, Thomas Rubiano |
ARES (2) | 4 |
| 2025 | Ahead of Time Generation for GPSA Protection in RISC-V Embedded CoresabstractState-of-the-art hardware countermeasures against fault attacks are based, among others, on control-flow and code integrity checking. Generalized Path Signature Analysis and Continuous Signature Monitoring can assert these integrity properties. However, many implementations of such mechanisms require a dedicated compiler flow and do not support indirect jumps, while others have prohibitive overheads. This work proposes a technique based on a ahead-of-time analysis to generate those signatures, associated with a hardware/software runtime handling indirect jumps while executing unmodified off-the-shelf RISC-V binaries. The proposed approach has been implemented on a pipelined processor, and experimental results show an average slowdown of$\times 1.82$and an area overhead of at least$\times 1.3$compared to unprotected implementations. Louis Savary, Simon Rokicki, Steven Derrien |
ASAP | 2 |
| 2025 | Optimizing Recovery Logic in Speculative High-Level SynthesisabstractHigh-Level Synthesis (HLS) excels at handling compute-intensive loops with straightforward control but struggles to identify parallelism in kernels with complex and irregular control-flow. To address this, novel scheduling techniques based on speculation have been introduced. While these methods outperform traditional static scheduling, they also introduce significant area overhead, particularly in the rollback control logic. Optimizing the cost of this rollback control logic remains an open challenge. In this work, we show how it is possible to simplify and/or eliminate rollback logic using a combination of static analysis and linear programming. Our results show improvements in both execution throughput and area cost. Dylan Leothaud, Jean-Michel Gorius, Simon Rokicki, Steven Derrien |
DAC | 3 |
| 2025 | Hardware/Software Runtime for GPSA Protection in RISC-V Embedded CoresabstractState-of-the-art hardware countermeasures against fault attacks are based, among others, on control-flow and code integrity checking. Generalized Path Signature Analysis and Continuous Signature Monitoring can assert these integrity properties. However, supporting such mechanisms requires a dedicated compiler flow and does not support indirect jumps. This work proposes a technique based on a hardware/software runtime to generate those signatures while executing unmodified off-the-shelf RISC-V binaries. To the best of our knowledge, this is the first solution for providing this level of protection against fault injection on unmodified binaries. The proposed approach has been implemented on a pipelined processor, and experimental results show an average slowdown of ×3.35 and an area overhead of at least ×1.86 compared to unprotected implementations. Louis Savary, Simon Rokicki, Steven Derrien |
DATE | 2 |
| 2024 | A Unified Memory Dependency Framework for Speculative High-Level SynthesisabstractHeterogeneous hardware platforms that leverage application-specific hardware accelerators are becoming increasingly popular as the demand for high-performance compute intensive applications rises. The design of such high-performance hardware accelerators is a complex task. High-Level Synthesis (HLS) promises to ease this process by synthesizing hardware from a high-level algorithmic description. Recent works have demonstrated that speculative execution can be inferred from the latter by leveraging compilation transformation and analysis techniques in HLS flows. However, existing work on speculative HLS lacks support for the intricate memory interactions in data-processing applications. In this paper, we introduce a unified memory speculation framework, which allows aggressive scheduling and high-throughput accelerator synthesis in the presence of complex memory dependencies. We show that our technique can generate high-throughput designs for various applications and describe a complete implementation inside an existing speculative HLS toolchain. Jean-Michel Gorius, Simon Rokicki, Steven Derrien |
CC | 2 |
| 2024 | Efficient Design Space Exploration for Dynamic & Speculative High-Level SynthesisabstractHigh-Level Synthesis performs well for compute-intensive loops with regular control but struggles to uncover parallelism in kernels with complex control-flow. Novel scheduling techniques based on dynamic scheduling and speculation have been proposed to address this issue. Although they outperform classical static scheduling techniques, they also come at a significant area overhead. Precisely determining where and by how much to apply these techniques remains an open problem, which we address in this work through an efficient exploration algorithm (combining pruning and search heuristics). We show that our approach can explore large solution spaces while producing efficient solutions. Dylan Leothaud, Jean-Michel Gorius, Simon Rokicki, Steven Derrien |
FPL | 3 |
| 2022 | RT-DFI: Optimizing Data-Flow Integrity for Real-Time SystemsabstractInternational audience Nicolas Bellec 0001, Guillaume Hiet, Simon Rokicki, Frédéric Tronel, Isabelle Puaut |
ECRTS | 3 |
| 2022 | Design Exploration of RISC-V Soft-Cores through Speculative High-Level SynthesisabstractThe RISC- V ecosystem is quickly growing and has gained a lot of traction in the FPGA community, as it permits free customization of both ISA and micro- architectural features. However, the design of the cor- responding micro-architecture is costly and error-prone. We address this issue by providing a flow capable of automatically synthesizing pipelined micro-architectures directly from an Instruction Set Simulator in C/C++. Our flow is based on HLS technology and bridges part of the gap between Instruction Set Processor design flows and High- Level Synthesis tools by taking advantage of speculative loop pipelining. Our results show that our flow is general enough to support a variety of ISA and micro-architectural extensions, and is capable of producing circuits that are competitive with manually designed cores. Jean-Michel Gorius, Simon Rokicki, Steven Derrien |
FPT | 2 |
| 2020 | GhostBusters: Mitigating Spectre Attacks on a DBT-Based ProcessorabstractUnveiled early 2018, the Spectre vulnerability affects most of the modern high-performance processors. Spectre variants exploit the speculative execution mechanisms and a cache side-channel attack to leak secret data. As of today, the main countermeasures consist of turning off the speculation, which drastically reduces the processor performance. In this work, we focus on a different kind of micro-architecture: the DBT based processors, such as Transmeta Crusoe [1], NVidia Denver [2], or Hybrid-DBT [3]. Instead of using complex outof-order (OoO) mechanisms, those cores combines a software Dynamic Binary Translation mechanism (DBT) and a parallel in-order architecture, typically a VLIW core. The DBT is in charge of translating and optimizing the binaries before their execution. Studies show that DBT based processors can reach the performance level of OoO cores for regular enough applications. In this paper, we demonstrate that, even if those processors do not use OoO execution, they are still vulnerable to Spectre variants, because of the DBT optimizations. However, we also demonstrate that those systems can easily be patched, as the DBT is done in software and has fine-grained control over the optimization process. Simon Rokicki |
DATE | 1 |
| 2020 | Attack Detection Through Monitoring of Timing Deviations in Embedded Real-Time SystemsabstractReal-time embedded systems (RTES) are required to interact more and more with their environment, thereby increasing their attack surface. Recent security breaches on car brakes and other critical components have already proven the feasibility of attacks on RTES. Such attacks may change the control-flow of the programs, which may lead to violations of the system’s timing constraints. In this paper, we present a technique to detect attacks in RTES based on timing information. Our technique, designed for single-core processors, is based on a monitor implemented in hardware to preserve the predictability of instrumented programs. The monitor uses timing information (Worst-Case Execution Time - WCET) of code regions to detect attacks. The proposed technique guarantees that attacks that delay the run-time of any region beyond its WCET are detected. Since the number of regions in programs impacts the memory resources consumed by the hardware monitor, our method includes a region selection algorithm that limits the amount of memory consumed by the monitor. An implementation of the hardware monitor and its simulation demonstrates the practicality of our approach. In particular, an experimental study evaluates the attack detection latency. Nicolas Bellec 0001, Simon Rokicki, Isabelle Puaut |
ECRTS | 2 |
| 2020 | Toward Speculative Loop Pipelining for High-Level SynthesisabstractLoop pipelining (LP) is a key optimization in modern high-level synthesis (HLS) tools for synthesizing efficient hardware datapaths. Existing techniques for automatic LP are limited by static analysis that cannot precisely analyze loops with data-dependent control flow and/or memory accesses. We propose a technique for speculative LP that handles both control-flow and memory speculations in a unified manner. Our approach is entirely expressed at the source level, allowing a seamless integration to development flows using HLS. Our evaluation shows significant improvement in throughput over standard LP. Steven Derrien, Thibaut Marty, Simon Rokicki, Tomofumi Yuki |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Aggressive Memory Speculation in HW/SW Co-Designed MachinesabstractSingle-ISA heterogeneous systems (such as ARM big.LITTLE) are an attractive solution for embedded platforms as they expose performance/energy trade-offs directly to the operating system. Recent works have demonstrated the ability to increase their efficiency by using VLIW cores, supported through Dynamic Binary Translation (DBT) to maintain the illusion of a single-ISA system. However, VLIW cores cannot rival with Out-of-Order (OoO) cores when it comes to performance, mainly because they do not use speculative execution. In this work, we study how it is possible to use memory dependency speculation during the DBT process. Our approach enables fine-grained speculation optimizations thanks to a combination of hardware and software. Our results show that our approach leads to a geo-mean speed-up of 10% at the price of a 7% area overhead. Simon Rokicki, Erven Rohou, Steven Derrien |
DATE | 1 |
| 2019 | What You Simulate Is What You Synthesize: Designing a Processor Core from C++ SpecificationsabstractThe following topics are dealt with: learning (artificial intelligence); neural nets; field programmable gate arrays; integrated circuit design; logic design; optimisation; network routing; low-power electronics; cryptography; electronic design automation. Simon Rokicki, Davide Pala, Joseph Paturel, Olivier Sentieys |
ICCAD | 1 |
| 2019 | Hybrid-DBT: Hardware/Software Dynamic Binary Translation Targeting VLIWabstractIn order to provide dynamic adaptation of the performance/energy tradeoff, systems today rely on heterogeneous multicore architectures (different micro-architectures on a chip). These systems are limited to single-ISA approaches to enable transparent migration between the different cores. To offer more tradeoff, we can integrate statically scheduled micro-architecture and use dynamic binary translation (DBT) for task migration. However, in a system where performance and energy consumption are a prime concern, the translation overhead has to be kept as low as possible. In this paper, we present Hybrid-DBT, an open-source, hardware accelerated DBT system targeting VLIW cores. Three different hardware accelerators have been designed to speed-up critical steps of the translation process. Experimental study shows that the accelerated steps are two orders of magnitude faster than their software equivalent. The impact on the total execution time of applications and the quality of generated binaries are also measured. Simon Rokicki, Erven Rohou, Steven Derrien |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | Supporting runtime reconfigurable VLIWs cores through dynamic binary translationabstractSingle ISA-Heterogeneous multi-cores such as the ARM big.LITTLE have proven to be an attractive solution to explore different energy/performance trade-offs. Such architectures combine Out of Order cores with smaller in-order ones to offer different power/energy profiles. They however do not really exploit the characteristics of workloads (compute-intensive vs. control dominated). In this work, we propose to enrich these architectures with runtime configurable VLIW cores, which are very efficient at compute-intensive kernels. To preserve the single ISA programming model, we resort to Dynamic Binary Translation, and use this technique to enable dynamic code specialization for Runtime Reconfigurable VLIWs cores. Our proposed DBT framework targets the RISC-V ISA, for which both OoO and in-order implementations exist. Our experimental results show that our approach can lead to best-case performance and energy efficiency when compared against static VLIW configurations. Simon Rokicki, Erven Rohou, Steven Derrien |
DATE | 1 |
| 2018 | Hybrid Obfuscation to Protect Against Disclosure Attacks on Embedded MicroprocessorsabstractThe risk of code reverse-engineering is particularly acute for embedded processors which often have limited available resources to protect program information. Previous efforts involving code obfuscation provide some additional security against reverse- engineering of programs, but the security benefits are typically limited and not quantifiable. Hence, new approaches to code protection and creation of associated metrics are highly desirable. This paper has two main contributions. We propose the first hybrid diversification approach for protecting embedded software and we provide statistical metrics to evaluate the protection. Diversification is achieved by combining hardware obfuscation at the microarchitecture level and the use of software-level obfuscation techniques tailored to embedded systems. Both measures are based on a compiler which generates obfuscated programs, and an embedded processor implemented in an FPGA with a randomized Instruction Set Architecture (ISA) encoding to execute the hybrid obfuscated program. We employ a fine-grained, hardware-enforced access control mechanism for information exchange with the processor and hardware-assisted booby traps to actively counteract manipulation attacks. It is shown that our approach is effective against a wide variety of possible information disclosure attacks in case of a physically present adversary. Moreover, we propose a novel statistical evaluation methodology that provides a security metric for hybrid-obfuscated programs. Marc Fyrbiak, Simon Rokicki, Nicolai Bissantz, Russell Tessier, Christof Paar |
IEEE Trans. Computers | 2 |
| 2017 | Hardware-accelerated dynamic binary translationabstractDynamic Binary Translation (DBT) is often used in hardware/software co-design to take advantage of an architecture model while using binaries from another one. The co-development of the DBT engine and of the execution architecture leads to architecture with special support to these mechanisms. In this work, we propose a hardware accelerated Dynamic Binary Translation where the first steps of the DBT process are fully accelerated in hardware. Results shows that using our hardware accelerators leads to a speed-up of 8× and a cost in energy 18× lower, compared with an equivalent software approach. Simon Rokicki, Erven Rohou, Steven Derrien |
DATE | 1 |