VLDB 2026 Research / reviewers in the wild / expert
Steven Derrien
dblp:28/3875
· DBLP profile ↗
56ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0002-6281-083XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 47 · 7 first-author · 9 since 2021Software engineering, systems software and programming languages · 16 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automatic Extraction of Timing Models for WCET Estimation From a High-Level Synthesis FlowabstractReal-time, domain-specific processors require faithful timing models for WCET analysis. However, existing models are typically hand-crafted from sparse documentation, making them error-prone and difficult to maintain. This work aims to automatically extract WCET timing models from single-issue in-order processor pipelines generated by High-Level Synthesis (HLS). By deriving timing models directly from the SpecHLS intermediate representation, the models are faithful by construction. Experimental results show that our timing-model extraction process generalizes across diverse RISC-V core variants and yields WCET estimates within 0.48% on average of those from a handcrafted model, on the Mälardalen WCET benchmarks. Thomas Feuilletin, Dylan Leothaud, Simon Rokicki, Steven Derrien, Isabelle Puaut |
DATE | 4 |
| 2026 | Area Efficient Speculative Loop Pipelining for High-Level SynthesisabstractHigh-Level Synthesis (HLS) allows the automatic generation of efficient circuit designs for computation-intensive kernels, but it lacks flexibility when dealing with irregular control flow. Dynamic and speculative HLS techniques are used to address this issue. These techniques outperform state-of-the-art HLS in kernel execution times but introduce a significant area overhead. In contrast, state-of-the-art HLS easily highlights and exploits resource-sharing opportunities. In this work, we show how to adapt an existing speculative HLS approach to take advantage of well-known static resource sharing mechanisms. Our results show a decrease of the area cost by 34% on average. Dylan Leothaud, Simon Rokicki, Steven Derrien, Isabelle Puaut |
DATE | 3 |
| 2026 | WCET Analysis of HLS-Generated Processors Using Abstract InterpretationabstractDeriving sound and precise timing models remains one of the main obstacles to static Worst-Case Execution Time (WCET) analysis. Modern processors exhibit diverse and evolving microarchitectures, making manual construction of timing models labor-intensive, error-prone, and difficult to adapt across processor variants. High-Level Synthesis (HLS) enables rapid customization of processor cores and architectural exploration, offering an opportunity to automate not only hardware generation but also the derivation of associated timing models. This paper presents an automated WCET analysis for HLS-generated processors based on abstract interpretation. We exploit the internal Gated-SSA representation of the HLS flow to automatically extract an abstract timing model capturing speculation and stall mechanisms. WCET estimation at the basic block level is then formulated as an exploration of abstract microarchitectural states within a basic block. The approach safely accounts for timing anomalies, while remaining scalable thanks to an efficient state-merging strategy. Integrated into the Heptane WCET tool and evaluated on Mälardalen benchmarks and a RISC-V, the method achieves the same tightness as a handcrafted timing model, while improving over a previously proposed automated approach. Thomas Feuilletin, Dylan Leothaud, Simon Rokicki, Steven Derrien, Isabelle Puaut |
ECRTS | 4 |
| 2025 | Ahead of Time Generation for GPSA Protection in RISC-V Embedded CoresabstractState-of-the-art hardware countermeasures against fault attacks are based, among others, on control-flow and code integrity checking. Generalized Path Signature Analysis and Continuous Signature Monitoring can assert these integrity properties. However, many implementations of such mechanisms require a dedicated compiler flow and do not support indirect jumps, while others have prohibitive overheads. This work proposes a technique based on a ahead-of-time analysis to generate those signatures, associated with a hardware/software runtime handling indirect jumps while executing unmodified off-the-shelf RISC-V binaries. The proposed approach has been implemented on a pipelined processor, and experimental results show an average slowdown of$\times 1.82$and an area overhead of at least$\times 1.3$compared to unprotected implementations. Louis Savary, Simon Rokicki, Steven Derrien |
ASAP | 3 |
| 2025 | Optimizing Recovery Logic in Speculative High-Level SynthesisabstractHigh-Level Synthesis (HLS) excels at handling compute-intensive loops with straightforward control but struggles to identify parallelism in kernels with complex and irregular control-flow. To address this, novel scheduling techniques based on speculation have been introduced. While these methods outperform traditional static scheduling, they also introduce significant area overhead, particularly in the rollback control logic. Optimizing the cost of this rollback control logic remains an open challenge. In this work, we show how it is possible to simplify and/or eliminate rollback logic using a combination of static analysis and linear programming. Our results show improvements in both execution throughput and area cost. Dylan Leothaud, Jean-Michel Gorius, Simon Rokicki, Steven Derrien |
DAC | 4 |
| 2025 | Hardware/Software Runtime for GPSA Protection in RISC-V Embedded CoresabstractState-of-the-art hardware countermeasures against fault attacks are based, among others, on control-flow and code integrity checking. Generalized Path Signature Analysis and Continuous Signature Monitoring can assert these integrity properties. However, supporting such mechanisms requires a dedicated compiler flow and does not support indirect jumps. This work proposes a technique based on a hardware/software runtime to generate those signatures while executing unmodified off-the-shelf RISC-V binaries. To the best of our knowledge, this is the first solution for providing this level of protection against fault injection on unmodified binaries. The proposed approach has been implemented on a pipelined processor, and experimental results show an average slowdown of ×3.35 and an area overhead of at least ×1.86 compared to unprotected implementations. Louis Savary, Simon Rokicki, Steven Derrien |
DATE | 3 |
| 2024 | A Unified Memory Dependency Framework for Speculative High-Level SynthesisabstractHeterogeneous hardware platforms that leverage application-specific hardware accelerators are becoming increasingly popular as the demand for high-performance compute intensive applications rises. The design of such high-performance hardware accelerators is a complex task. High-Level Synthesis (HLS) promises to ease this process by synthesizing hardware from a high-level algorithmic description. Recent works have demonstrated that speculative execution can be inferred from the latter by leveraging compilation transformation and analysis techniques in HLS flows. However, existing work on speculative HLS lacks support for the intricate memory interactions in data-processing applications. In this paper, we introduce a unified memory speculation framework, which allows aggressive scheduling and high-throughput accelerator synthesis in the presence of complex memory dependencies. We show that our technique can generate high-throughput designs for various applications and describe a complete implementation inside an existing speculative HLS toolchain. Jean-Michel Gorius, Simon Rokicki, Steven Derrien |
CC | 3 |
| 2024 | Efficient Design Space Exploration for Dynamic & Speculative High-Level SynthesisabstractHigh-Level Synthesis performs well for compute-intensive loops with regular control but struggles to uncover parallelism in kernels with complex control-flow. Novel scheduling techniques based on dynamic scheduling and speculation have been proposed to address this issue. Although they outperform classical static scheduling techniques, they also come at a significant area overhead. Precisely determining where and by how much to apply these techniques remains an open problem, which we address in this work through an efficient exploration algorithm (combining pruning and search heuristics). We show that our approach can explore large solution spaces while producing efficient solutions. Dylan Leothaud, Jean-Michel Gorius, Simon Rokicki, Steven Derrien |
FPL | 4 |
| 2023 | Automatic Algorithm-Based Fault Tolerance (AABFT) of Stencil ComputationsabstractIn this work, we study fault tolerance of transient errors, such as those occurring due to cosmic radiation or hardware component aging and degradation, using Algorithm-Based Fault Tolerance (ABFT). ABFT methods typically work by adding some additional computation in the form of invariant checksums which, by definition, should not change as the program executes. By computing and monitoring checksums, it is possible to detect errors by observing differences in the checksum values. However, this is challenging for two key reasons: (1) it requires careful manual analysis of the input program, and (2) care must be taken to subsequently carry out the checksum computations efficiently enough for it to be worth it. Prior work has shown how to apply ABFT schemes with low overhead for a variety of input programs. Here, we focus on a subclass of programs called stencil applications, an important class of computations found widely in various scientific computing domains. We propose a new compilation scheme to automatically analyze and generate the checksum computations. To the best of our knowledge, this is the first work to do such a thing in a compiler. We show that low overhead code can be easily generated and provide a preliminary evaluation of the tradeoff between performance and effectiveness. Louis Narmour, Steven Derrien, Sanjay V. Rajopadhye |
PACT | 2 |
| 2023 | Increasing FPGA Accelerators Memory Bandwidth With a Burst-Friendly Memory LayoutabstractOffloading compute-intensive kernels to hardware accelerators relies on the large degree of parallelism offered by these platforms. However, the effective bandwidth of the memory interface often causes a bottleneck, hindering the accelerator’s effective performance. Techniques enabling data reuse, such as tiling, lower the pressure on memory traffic but do not fully exploit the bandwidth. A further increase in bandwidth utilization is possible by using burst rather than element-wise accesses, provided the data is contiguous in memory. In this article, we propose a memory allocation technique, and provide a proof-of-concept source-to-source compiler pass, that enables such burst transfers by modifying the data layout in external memory. Our experiments show the new memory allocation yields close to 100% bandwidth utilization while the memory engines occupy less than 5% of the field-programmable gate array (FPGA) logic area. Corentin Ferry, Tomofumi Yuki, Steven Derrien, Sanjay V. Rajopadhye |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Design Exploration of RISC-V Soft-Cores through Speculative High-Level SynthesisabstractThe RISC- V ecosystem is quickly growing and has gained a lot of traction in the FPGA community, as it permits free customization of both ISA and micro- architectural features. However, the design of the cor- responding micro-architecture is costly and error-prone. We address this issue by providing a flow capable of automatically synthesizing pipelined micro-architectures directly from an Instruction Set Simulator in C/C++. Our flow is based on HLS technology and bridges part of the gap between Instruction Set Processor design flows and High- Level Synthesis tools by taking advantage of speculative loop pipelining. Our results show that our flow is general enough to support a variety of ISA and micro-architectural extensions, and is capable of producing circuits that are competitive with manually designed cores. Jean-Michel Gorius, Simon Rokicki, Steven Derrien |
FPT | 3 |
| 2020 | Application-Specific Arithmetic in High-Level Synthesis ToolsabstractThis work studies hardware-specific optimization opportunities currently unexploited by high-level synthesis compilers. Some of these optimizations are specializations of floating-point operations that respect the usual semantics of the input program without changing the numerical result. Some other optimizations, locally triggered by the programmer thanks to a pragma, assume a different semantics, where floating-point code is interpreted as the specification of computation with real numbers. The compiler is then in charge to ensure an application-level accuracy constraint expressed in the pragma and has the freedom to use non-standard arithmetic hardware when more efficient. These two classes of optimizations are prototyped in the GeCoS source-to-source compiler and evaluated on the Polybench and EEMBC benchmark suites. Latency is reduced by up to 93%, and resource usage is reduced by up to 58%. Yohann Uguen, Florent de Dinechin, Victor Lezaud, Steven Derrien |
ACM Trans. Archit. Code Optim. | 4 |
| 2020 | Toward Speculative Loop Pipelining for High-Level SynthesisabstractLoop pipelining (LP) is a key optimization in modern high-level synthesis (HLS) tools for synthesizing efficient hardware datapaths. Existing techniques for automatic LP are limited by static analysis that cannot precisely analyze loops with data-dependent control flow and/or memory accesses. We propose a technique for speculative LP that handles both control-flow and memory speculations in a unified manner. Our approach is entirely expressed at the source level, allowing a seamless integration to development flows using HLS. Our evaluation shows significant improvement in throughput over standard LP. Steven Derrien, Thibaut Marty, Simon Rokicki, Tomofumi Yuki |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | Safe Overclocking for CNN Accelerators Through Algorithm-Level Error DetectionabstractIn this article, we propose a technique for improving the efficiency of convolutional neural network hardware accelerators based on timing speculation (overclocking) and fault tolerance. We augment the accelerator with a lightweight error detection mechanism to protect against timing errors in convolution layers, enabling aggressive timing speculation. The error detection mechanism we have developed works at the algorithm-level, utilizing algebraic properties of the computation, allowing the full implementation to be realized using high-level synthesis tools. Our prototype on ZC706 demonstrated up to 60% higher throughput with negligible area overhead for various wordlength implementations. Thibaut Marty, Tomofumi Yuki, Steven Derrien |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Aggressive Memory Speculation in HW/SW Co-Designed MachinesabstractSingle-ISA heterogeneous systems (such as ARM big.LITTLE) are an attractive solution for embedded platforms as they expose performance/energy trade-offs directly to the operating system. Recent works have demonstrated the ability to increase their efficiency by using VLIW cores, supported through Dynamic Binary Translation (DBT) to maintain the illusion of a single-ISA system. However, VLIW cores cannot rival with Out-of-Order (OoO) cores when it comes to performance, mainly because they do not use speculative execution. In this work, we study how it is possible to use memory dependency speculation during the DBT process. Our approach enables fine-grained speculation optimizations thanks to a combination of hardware and software. Our results show that our approach leads to a geo-mean speed-up of 10% at the price of a 7% area overhead. Simon Rokicki, Erven Rohou, Steven Derrien |
DATE | 3 |
| 2019 | Hiding Communication Delays in Contention-Free Execution for SPM-Based Multi-Core ArchitecturesabstractMulti-core systems using ScratchPad Memories (SPMs) are attractive architectures for executing time-critical embedded applications, because they provide both predictability and performance. In this paper, we propose a scheduling technique that jointly selects SPM contents off-line, in such a way that the cost of SPM loading/unloading is hidden. Communications are fragmented to augment hiding possibilities. Experimental results show the effectiveness of the proposed technique on streaming applications and synthetic task-graphs. The overlapping of communications with computations allows the length of generated schedules to be reduced by 4% on average on streaming applications, with a maximum of 16%, and by 8% on average for synthetic task graphs. We further show on a case study that generated schedules can be implemented with low overhead on a predictable multi-core architecture (Kalray MPPA). Benjamin Rouxel, Stefanos Skalistis, Steven Derrien, Isabelle Puaut |
ECRTS | 3 |
| 2019 | Reconciling Compiler Optimizations and WCET Estimation Using Iterative CompilationabstractStatic Worst-Case Execution Time (WCET) estimation techniques operate upon the binary code of a program in order to provide the necessary input for schedulability analysis techniques. Compilers used to generate this binary code include tens of optimizations, that can radically change the flow information of the program. Such information is hard to be maintained across optimization passes and may render automatic extraction of important flow information, such as loop bounds, impossible. Thus, compiler optimizations, especially the sophisticated optimizations of mainstream compilers, are typically avoided. In this work, we explore for the first time iterative-compilation techniques that reconcile compiler optimizations and static WCET estimation. We propose a novel learning technique that selects sequences of optimizations that minimize the WCET estimate of a given program. We experimentally evaluate the proposed technique using an industrial WCET estimation tool (AbsInt aiT) over a set of 46 benchmarks from four different benchmarks suites, including reference WCET benchmark applications, image processing kernels and telecommunication applications. Experimental results show that WCET estimates are reduced on average by 20.3% using the proposed technique, as compared to the best compiler optimization level applicable. Mickaël Dardaillon, Stefanos Skalistis, Isabelle Puaut, Steven Derrien |
RTSS | 4 |
| 2019 | Hybrid-DBT: Hardware/Software Dynamic Binary Translation Targeting VLIWabstractIn order to provide dynamic adaptation of the performance/energy tradeoff, systems today rely on heterogeneous multicore architectures (different micro-architectures on a chip). These systems are limited to single-ISA approaches to enable transparent migration between the different cores. To offer more tradeoff, we can integrate statically scheduled micro-architecture and use dynamic binary translation (DBT) for task migration. However, in a system where performance and energy consumption are a prime concern, the translation overhead has to be kept as low as possible. In this paper, we present Hybrid-DBT, an open-source, hardware accelerated DBT system targeting VLIW cores. Three different hardware accelerators have been designed to speed-up critical steps of the translation process. Experimental study shows that the accelerated steps are two orders of magnitude faster than their software equivalent. The impact on the total execution time of applications and the quality of generated binaries are also measured. Simon Rokicki, Erven Rohou, Steven Derrien |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Using polyhedral techniques to tighten WCET estimates of optimized code: A case study with array contractionabstractThe ARGO H2020 European project aims at developing a Worst-Case Execution Time (WCET)-aware parallelizing compilation toolchain. This toolchain operates on Scilab and XCoS inputs, and targets ScratchPad memory (SPM)-based multi-cores. Data-layout and loop transformations play a key role in this flow as they improve SPM efficiency and reduce the number of accesses to shared main memory. In this paper1, we study how these transformations impact WCET estimates of sequential codes. We demonstrate that they can bring significant improvements of WCET estimates (up to 2.7 χ) provided that the WCET analysis process is guided with automatically generated flow annotations obtained using polyhedral counting techniques. Thomas Lefeuvre, Imen Fassi, Christoph Cullmann, Gernot Gebhard, Emin-Koray Kasnakli, Isabelle Puaut, Steven Derrien |
DATE | 7 |
| 2018 | Supporting runtime reconfigurable VLIWs cores through dynamic binary translationabstractSingle ISA-Heterogeneous multi-cores such as the ARM big.LITTLE have proven to be an attractive solution to explore different energy/performance trade-offs. Such architectures combine Out of Order cores with smaller in-order ones to offer different power/energy profiles. They however do not really exploit the characteristics of workloads (compute-intensive vs. control dominated). In this work, we propose to enrich these architectures with runtime configurable VLIW cores, which are very efficient at compute-intensive kernels. To preserve the single ISA programming model, we resort to Dynamic Binary Translation, and use this technique to enable dynamic code specialization for Runtime Reconfigurable VLIWs cores. Our proposed DBT framework targets the RISC-V ISA, for which both OoO and in-order implementations exist. Our experimental results show that our approach can lead to best-case performance and energy efficiency when compared against static VLIW configurations. Simon Rokicki, Erven Rohou, Steven Derrien |
DATE | 3 |
| 2018 | Enabling Overclocking Through Algorithm-Level Error DetectionabstractIn this paper, we propose a technique for improving the efficiency of hardware accelerators based on timing speculation (overclocking) and fault tolerance. We augment the accelerator with a lightweight error detection mechanism to protect against timing errors, enabling aggressive timing speculation. We demonstrate the validity of our approach for the convolution layers in convolutional neural networks. We present an implementation of a fault-tolerant convolution layer accelerator combined with the lightweight error detection. The error detection mechanism we have developed works at the algorithm-level, utilizing algebraic properties of the computation, allowing the full implementation to be realized using High-Level Synthesis tools. Our prototype on ZC706 demonstrated 68% - 77% higher throughput with negligible overhead. Thibaut Marty, Tomofumi Yuki, Steven Derrien |
FPT | 3 |
| 2017 | WCET-aware parallelization of model-based applications for multi-cores: The ARGO approachabstractParallel architectures are nowadays not only confined to the domain of high performance computing, they are also increasingly used in embedded time-critical systems. The ARGO H2020 project1provides a programming paradigm and associated tool flow to exploit the full potential of architectures in terms of development productivity, time-to-market, exploitation of the platform computing power and guaranteed real-time performance. In this paper we give an overview of the objectives of ARGO and explore the challenges introduced by our approach. Steven Derrien, Isabelle Puaut, Panayiotis Alefragis, Marcus Bednara, Harald Bucher, Clément David, Yann Debray, Umut Durak, Imen Fassi, Christian Ferdinand, Damien Hardy, Angeliki Kritikakou, Gerard K. Rauwerda, Simon Reder, Martin Sicks, Timo Stripf, Kim Sunesen, Timon D. ter Braak, Nikos S. Voros, Jürgen Becker 0001 |
DATE | 1 |
| 2017 | Superword level parallelism aware word length optimizationabstractMany embedded processors do not support floating-point arithmetic in order to comply with strict cost and power consumption constraints. But, they generally provide support for SIMD as a mean to improve performance for little cost overhead. Achieving good performance when targeting such processors requires the use of fixed-point arithmetic and efficient exploitation of SIMD data-path. To reduce time-to-market, automatic SIMDization - such as superword level parallelism (SLP) extraction - and float-to-fixed-point conversion methodologies have been proposed. In this paper we show that applying these transformations independently is not efficient. We propose a SLP-aware word length optimization algorithm to jointly perform float-to-fixed-point conversion and SLP extraction. We implement the proposed approach in a source-to-source compiler framework and evaluate it on several embedded processors. Experimental results illustrate the validity of our approach. Ali El Moussawi, Steven Derrien |
DATE | 2 |
| 2017 | Hardware-accelerated dynamic binary translationabstractDynamic Binary Translation (DBT) is often used in hardware/software co-design to take advantage of an architecture model while using binaries from another one. The co-development of the DBT engine and of the execution architecture leads to architecture with special support to these mechanisms. In this work, we propose a hardware accelerated Dynamic Binary Translation where the first steps of the DBT process are fully accelerated in hardware. Results shows that using our hardware accelerators leads to a speed-up of 8× and a cost in energy 18× lower, compared with an equivalent software approach. Simon Rokicki, Erven Rohou, Steven Derrien |
DATE | 3 |
| 2017 | A High-Level Synthesis Approach Optimizing Accumulations in Floating-Point Programs Using Custom Formats and OperatorsabstractMany case studies have demonstrated the potential of Field-Programmable Gate Arrays (FPGAs) as accelerators for a wide range of applications. FPGAs offer massive parallelism and programmability at the bit level. This enables programmers to exploit a range of techniques that avoid many bottlenecks of classical von Neumann computing. However, development costs for FPGAs are orders of magnitude higher than classical programming. A solution would be the use of High-Level Synthesis (HLS) tools, which use C as a hardware description language. However, the C language was designed to be executed on general purpose processors, not to generate hardware. Its datatypes and operators are limited to a small number (more or less matching the hardware operators present in mainstream processors), and HLS tools inherit these limitations. To better exploit the freedom offered by hardware and FPGAs, HLS vendors have enriched the C language with integer and fixed-point types of arbitrary size. Still, the operations on these types remain limited to the basic arithmetic and logic ones. In floating point, the current situation is even worse. The operator set is limited, and the sizes are restricted to 32 and 64 bits. Besides, most recent compilers, including the HLS ones, attempt to follow established standards, in particular C11 and IEEE-754. This ensures bit-exact compatibility with software, but greatly reduces the freedom of optimization by the compiler. For instance, a floating point addition is not associative even though its real equivalent is. In the present work we attempt to give the compiler more freedom. For this, we sacrifice the strict respect of the IEEE-754 and C11 standards, but we replace it with the strict respect of a high-level accuracy specification expressed by the programmer through a pragma. The case study in this work is a program transformation that applies to floating-point additions on a loop's critical path. It decomposes them into elementary steps, resizes the corresponding subcomponents to guarantee some user-specified accuracy, and merges and reorders these components to improve performance. The result of this complex sequence of optimizations could not be obtained from an operator generator, since it involves global loop information. For this purpose, we used a compilation flow involving one or several source-to-source transformations operating on the code given to HLS tools (Figure 1).The proposed transformation already works very well on 3 of the 10 FPMarks where it improves both latency and accuracy by an order of magnitude for comparable area. For 2 more benchmarks, the latency is not improved (but not degraded either) due to current limitations of HLS tools. This defines short-term future work. The main result of this work is that HLS tools also have the potential to generate efficient designs for handling floating-point computations in a completely non-standard way. In the longer term, we believe that HLS flows can not only import application-specific operators from the FPGA literature, they can also improve them using high-level, program-level information. Yohann Uguen, Florent de Dinechin, Steven Derrien |
FCCM | 3 |
| 2017 | One size does not fit all: Implementation trade-offs for iterative stencil computations on FPGAsabstractIterative stencils are kernels in various application domains such as numerical simulations and medical imaging, that merit FPGA acceleration. The best architecture depends on many factors such as the target platform, off-chip memory bandwidth, problem size, and performance requirements. We generate a family of FPGA stencil accelerators targeting emerging System on Chip platforms, (e.g., Xilinx Zynq or Intel SoC). Our designs come with design knobs to explore trade-offs. We also propose performance models to hone in on the most interesting design points, and show how they accurately lead to optimal designs. The optimal choice depends on problem sizes and performance goals. Gaël Deest, Tomofumi Yuki, Sanjay V. Rajopadhye, Steven Derrien |
FPL | 4 |
| 2017 | Bridging high-level synthesis and application-specific arithmetic: The case study of floating-point summationsabstractFPGAs are well known for their ability to perform non-standard computations not supported by classical microprocessors. Many libraries of highly customizable application-specific IPs have exploited this capability. However, using such IPs usually requires handcrafted HDL, hence significant design efforts. High Level Synthesis (HLS) lowers the design effort thanks to the use of C/C++ dialects for programming FPGAs. However, high-level C language becomes a hindrance when one wants to express non-standard computations: this languages was designed for programming microprocessors and carries with it many restrictions due to this paradigm. This is especially true when computing with floating-point, whose data-types and evaluation semantics are defined by the IEEE-754 and C11 standards. If the high-level specification was a computation on the reals, then HLS imposes a very restricted implementation space. This work attempts to bridge FPGA application-specific efficiency and HLS ease of use. It specifically targets the ubiquitous floating-point summation-reduction pattern. A source-to-source compiler transforms selected floating-point additions into sequences of simpler operators using non-standard arithmetic formats. This improves performance and accuracy for several benchmarks, while keeping the ease of use of a high-level C description. Yohann Uguen, Florent de Dinechin, Steven Derrien |
FPL | 3 |
| 2017 | Tightening Contention Delays While Scheduling Parallel Applications on Multi-core ArchitecturesabstractMulti-core systems are increasingly interesting candidates for executing parallel real-time applications, in avionic, space or automotive industries, as they provide both computing capabilities and power efficiency. However, ensuring that timing constraints are met on such platforms is challenging, because some hardware resources are shared between cores. Assuming worst-case contentions when analyzing the schedulability of applications may result in systems mistakenly declared unschedulable, although the worst-case level of contentions can never occur in practice. In this paper, we present two contention-aware scheduling strategies that produce a time-triggered schedule of the application’s tasks. Based on knowledge of the application’s structure, our scheduling strategies precisely estimate the effective contentions, in order to minimize the overall makespan of the schedule. An Integer Linear Programming (ILP) solution of the scheduling problem is presented, as well as a heuristic solution that generates schedules very close to ones of the ILP (5% longer on average), with a much lower time complexity. Our heuristic improves by 19% the overall makespan of the resulting schedules compared to a worst-case contention baseline. Benjamin Rouxel, Steven Derrien, Isabelle Puaut |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2016 | System level synthesis for virtual memory enabled hardware threads
Nicolas Estibals, Gaël Deest, Ali El Moussawi, Steven Derrien |
DATE | 4 |
| 2014 | Low Power Reconfigurable Controllers for Wireless Sensor Network NodesabstractA key concern in the design of controllers in wireless sensor network (WSN) nodes is the flexibility to execute different control tasks involving sensing, communications and computational resources of the node. In this paper, low power flexible controllers for WSN nodes based on reconfigurable microtasks composed of an FSM and datapath are presented. Coarse grain power gating opportunities are exploited in FSM and datapath for low power operation in reconfigurable microtasks. Power estimation results on typical benchmark microtasks show a 2× to 5× improvement in energy efficiency w.r.t a microcontroller at a cost of 5× relative to a microtask implemented as an ASIC with higher NRE costs. Vivek D. Tovinakere, Olivier Sentieys, Steven Derrien, Christophe Huriaux |
FCCM | 3 |
| 2014 | Toward scalable source level accuracy analysis for floating-point to fixed-point conversionabstractIn embedded systems, many numerical algorithms are implemented with fixed-point arithmetic to meet area cost and power constraints. Fixed-point encoding decisions can significantly affect cost and performance. To evaluate their impact on accuracy, designers resort to simulations. Their high running-time prevents thorough exploration of the design-space. To address this issue, analytical modeling techniques have been proposed, but their applicability is limited by scalability issues. In this paper, we extend these techniques to a larger class of programs. We use polyhedral methods to extract a more compact, graph-based representation of the program. We validate our approach with a several image and signal processing algorithms. Gaël Deest, Tomofumi Yuki, Olivier Sentieys, Steven Derrien |
ICCAD | 4 |
| 2013 | Runtime dependency analysis for loop pipelining in high-level synthesisabstractResearch on High-Level Synthesis has mainly focused on applications with statically determinable characteristics and current tools often perform poorly in presence of data-dependent memory accesses. The reason is that they rely on conservative static scheduling strategies, which lead to inefficient implementations. In this work, we propose to address this issue by leveraging well-known techniques used in superscalar processors to perform runtime memory disambiguation. Our approach, implemented as a source-to-source transformation at the C level, demonstrates significant performance improvements for a moderate increase in area while retaining portability among HLS tools. Mythri Alle, Antoine Morvan, Steven Derrien |
DAC | 3 |
| 2013 | Component-Level Datapath Merging in System-Level Design of Wireless Sensor Node Controllers for FPGA-Based ImplementationsabstractWireless Sensor Networks (WSNs) are relatively new and challenging research area for embedded design automation. Engineering a WSN node hardware is a difficult job as the design must satisfy several constraints. Among these constraints, overall energy consumption and node size, are the two most significant constraints. WSN node platforms have until recently been designed using off-the-shelf low-power microprocessors (MCUs), even though energy profile of these MCUs is not suitable for ultra low-power sensor nodes. On the other hand, WSN-specific hardware accelerators have also been proposed that have excellent energy profile but lack in flexibility, need higher design efforts and have huge non-recurring engineering (NRE) costs. In this work, we propose an automated system level design flow for an intermediate approach, based on the concept of data path merging (DPM) where several hardware accelerators (called micro-tasks) share a common customized data path, to have an improvement in flexibility and silicon area with possible increase in dynamic power consumption for the control/processing part of the sensor node targeted for field programmable gate array (FPGA)-based implementation. Our experiments show that component-level DPM yields to savings from 20%, to 75% for various FPGA resources like I/O ports, area for combinational and sequential logic, and static power consumption. Muhammad Adeel Pasha, Steven Derrien, Olivier Sentieys |
DSD | 2 |
| 2013 | Using Model Types to Support Contract-Aware Model Substitutability
Wuliang Sun, Benoît Combemale, Steven Derrien, Robert B. France |
ECMFA | 3 |
| 2013 | Derivation of efficient FSM from loop nestsabstractPipelined execution is one of the most important optimizations in hardware design to improve hardware utilization rate, and hence the throughput. Loop pipelining is a transformation available in High Level Synthesis tools to execute multiple iterations of a loop in a pipeline. Nested loop pipelining is a related technique that improves hardware utilization rate when the iteration count of the innermost loop is small. However, it is also known to increase the complexity of the control, and hence degrading frequency. In this paper, we present an automatic transformation targeting HLS that improves the effectiveness of nested loop pipelining, by efficient implementations of the control-path. Specifically, we present (i) an analytical model that captures the trade-off between gain in cycles and loss in frequency, (ii), automatic derivation of efficient Finite State Machine from loop nests, and (iii) an efficient implementation of the derived FSM that improves the performance of synthesized hardware. Tomofumi Yuki, Antoine Morvan, Steven Derrien |
FPT | 3 |
| 2013 | GeCoS: A framework for prototyping custom hardware design flowsabstractGeCoS is an open source framework that provide a highly productive environment for hardware design. GeCoS primarily targets custom hardware design using High Level Synthesis, distinguishing itself from classical compiler infrastructures. Compiling for custom hardware makes use of domain specific semantics that are not considered by general purpose compilers. Finding the right balance between various performance criteria, such as area, speed, and accuracy, is the goal, contrary to the typical goal in high performance context to maximize speed. The GeCoS infrastructure facilitates the prototyping of hardware design flows, going beyond compiler analyses and transformations. Hardware designers must interact with the compiler for design space exploration, and it is important to be able to give instant feedback to the users. Antoine Floch, Tomofumi Yuki, Ali El Moussawi, Antoine Morvan, Kevin J. M. Martin, Maxime Naullet, Mythri Alle, Ludovic L'Hours, Nicolas Simon, Steven Derrien, François Charot, Christophe Wolinski, Olivier Sentieys |
SCAM | 10 |
| 2013 | Polyhedral Bubble Insertion: A Method to Improve Nested Loop Pipelining for High-Level SynthesisabstractHigh-level synthesis (HLS) allows hardware to be directly produced from behavioral description in C/C++, thus accelerating the design process. Loop pipelining is a key transformation of HLS, as it improves the throughput of the design at the price of a small hardware overhead. However, for small loops, its use often results in a poor hardware utilization due to the pipeline latency overhead. Overlapping the iterations of the whole loop nest instead of only overlapping the innermost loop is a way to overcome this difficulty, but currently available techniques are restricted to perfectly nested loops with constant bounds, involving uniform dependences only. Using the polyhedral model, we extend the applicability of the nested loop pipelining transformation by proposing a new legality check and a new loop correction technique, called polyhedral bubble insertion. This method was implemented in a source-to-source compiler targeting HLS, and results on benchmark kernels show that polyhedral bubble insertion is effective in practice on a much larger class of loop nests. Antoine Morvan, Steven Derrien, Patrice Quinton |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | A semiempirical model for wakeup time estimation in power-gated logic clustersabstractWakeup time is an important overhead that must be determined for effective power gating, particularly in logic clusters that undergo frequent mode transitions for run-time leakage power reduction. In this paper, a semiempirical model for virtual supply voltage in terms of basic parameters of the power-gated circuit is presented. Hence a closed-form expression for estimation of wakeup time of a power-gated logic cluster is derived. Experimental results of application of the model to ISCAS85 benchmark circuits show that wakeup time may be estimated within an average error of 16.3% across 22x variation in sleep transistor sizes and 13x variation in circuit sizes with significant speedup in computation time compared to SPICE level circuit simulations. Vivek D. Tovinakere, Olivier Sentieys, Steven Derrien |
DAC | 3 |
| 2012 | From Scilab to High Performance Embedded Multicore Systems: The ALMA ApproachabstractThe mapping process of high performance embedded applications to today's multiprocessor system on chip devices suffers from a complex tool chain and programming process. The problem here is the expression of parallelism with a pure imperative programming language which is commonly C. This traditional approach limits the mapping, partitioning and the generation of optimized parallel code, and consequently the achievable performance and power consumption of applications from different domains. The Architecture oriented paraLlelization for high performance embedded Multicore systems using scilAb (ALMA) European project aims to bridge these hurdles through the introduction and exploitation of a Scilab-based toolchain which enables the efficient mapping of applications on multiprocessor platforms from high level of abstraction. This holistic solution of the toolchain allows the complexity of both the application and the architecture to be hidden, which leads to a better acceptance, reduced development cost, and shorter time-to-market. Driven by the technology restrictions in chip design, the end of exponential growth of clock speeds, and an unavoidable increasing request of computing performance, ALMA is a fundamental step forward in the necessary introduction of novel computing paradigms and methodologies. Jürgen Becker 0001, Timo Stripf, Oliver Wolf, Michael Hübner 0001, Steven Derrien, Daniel Ménard, Olivier Sentieys, Gerard K. Rauwerda, Kim Sunesen, Nikolaos Kavvadias, Kostas Masselos, George Goulas, Panayiotis Alefragis, Nikos S. Voros, Dimitrios Kritharidis, Nikolaos Mitas, Diana Göhringer |
DSD | 5 |
| 2012 | On Model Subtyping
Clément Guy, Benoît Combemale, Steven Derrien, Jim Steel, Jean-Marc Jézéquel |
ECMFA | 3 |
| 2012 | Bridging the chasm between MDE and the world of compilation
Jean-Marc Jézéquel, Benoît Combemale, Steven Derrien, Clément Guy, Sanjay V. Rajopadhye |
Softw. Syst. Model. | 3 |
| 2012 | System-Level Synthesis for Wireless Sensor Node Controllers: A Complete Design FlowabstractWireless sensor networks (WSN) is a new and very challenging research field for embedded system design automation. Engineering a WSN node hardware platform is known to be a tough challenge, as the design must enforce many severe constraints, among which energy dissipation is by far the most important one. WSN node devices have until now been designed using off-the-shelf low-power microcontroller units (MCUs), even if their power dissipation is still an issue and hinders the widespread use of this new technology. In this work, we propose a complete system-level flow for an alternative approach based on the concept of hardware microtasks, which relies on hardware specialization and power gating to drastically improve the energy efficiency of the computational/control part of the node. Our case study shows that power savings between one to two orders of magnitude are possible w.r.t. MCU-based implementations. Muhammad Adeel Pasha, Steven Derrien, Olivier Sentieys |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2011 | Efficient nested loop pipelining in high level synthesis using polyhedral bubble insertionabstractLoop pipelining is a key transformation in high-level synthesis tools as it helps maximizing both computational throughput and hardware utilization. Nevertheless, it somewhat looses its efficiency when dealing with small trip-count inner loops, as the pipeline latency overhead quickly limits its efficiency. Even if it is possible to overcome this limitation by pipelining the execution of a whole loop nest, the applicability of nested loop pipelining has so far been limited to a very narrow subset of loops, namely perfectly nested loops with constant bounds. In this work we propose to extend the applicability of nested-loop pipelining to imperfectly nested loops with affine dependencies by leveraging on the so-called polyhedral model. We show how such loop nest can be analyzed, and under certain conditions, how one can modify the source code in order to allow nested loop pipeline to be applied using a method called polyhedral bubble insertion. We also discuss the implementation of our method in a source-to-source compiler specifically targeted at High-Level Synthesis tools. Antoine Morvan, Steven Derrien, Patrice Quinton |
FPT | 2 |
| 2011 | Model-Driven Engineering and Optimizing Compilers: A Bridge Too Far?
Antoine Floch, Tomofumi Yuki, Clément Guy, Steven Derrien, Benoît Combemale, Sanjay V. Rajopadhye, Robert B. France |
MoDELS | 4 |
| 2010 | A complete design-flow for the generation of ultra low-power WSN node architectures based on micro-taskingabstractWireless Sensor Networks (WSN) are a new and very challenging research field for embedded system design automation, as their design must enforce stringent constraints in terms of power and cost. WSN node devices have until now been designed using off-the-shelf low-power microcontroller units (MCUs), even if their power dissipation is still an issue and hinders the wide-spreading of this new technology. In this paper, we propose a new architectural model for WSN nodes (and its complete design-flow from C downto synthesizable VHDL) based on the notion of micro-tasks. Our approach combines hardware specialization and power-gating so as to provide an ultra low-power solution for WSN node design. Our first estimates show that power savings by one to two orders of magnitude are possible w.r.t. MCU-based implementations. Muhammad Adeel Pasha, Steven Derrien, Olivier Sentieys |
DAC | 2 |
| 2010 | System Level Synthesis for Ultra Low-Power Wireless Sensor NodesabstractEngineering hardware platform for a Wireless Sensor Network (WSN) node is known to be a tough challenge, as the design must enforce many severe constraints, among which energy dissipation is by far the most challenging one. Today, most of the WSN node platforms are based on low cost and low-power programmable micro controllers, even if it is acknowledged that their energy efficiency remains limited and hinders the wide-spreading of WSN to new applications. In this paper, we propose a complete system level flow for an alternative approach based on the concept of hardware micro-tasks, which relies on hardware specialization and power gating to dramatically improve the energy efficiency of the computational part of the node. Early estimates show power saving by more than one order of magnitude over MCU-based implementations. Muhammad Adeel Pasha, Steven Derrien, Olivier Sentieys |
DSD | 2 |
| 2010 | Accelerating HMMER on FPGA using parallel prefixes and reductionsabstractHMMER is a widely used tool in bioinformatic, based on Profile Hidden Markov Models. The computation kernels of HMMER i.e. MSV and P7Viterbi are very compute intensive and data dependencies restrict to sequential execution. In this paper, we propose an original parallelization scheme for HMMER by rewriting their mathematical formulation, to expose the hidden potential parallelization opportunities. Our parallelization scheme targets FPGA technology, and our architecture can achieve 10 times speedup compared with that of latest HMMER3 SSE version, while not compromising on sensitivity of original algorithm. Naeem Abbas, Steven Derrien, Sanjay V. Rajopadhye, Patrice Quinton |
FPT | 2 |
| 2009 | Ultra Low-power FSM for Control Oriented ApplicationsabstractIn this paper, we propose an approach combining the use of distributed hardware tasks implemented as finite state machines (FSM) and power gating techniques to obtain ultra low-power implementations. We target for control dominated applications represented as control task graphs, and propose a complete flow including a C to hardware task compiler. Our approach is validated experimentally and shows impressive improvement over software implementation on leading edge low-power microcontrollers such as the MSP430. Muhammad Adeel Pasha, Steven Derrien, Olivier Sentieys |
ISCAS | 2 |
| 2007 | Parallelizing HMMER for Hardware Acceleration on FPGAsabstractProfile based Hidden Markov Model is a widely used tool in bioinformatics. While being very valuable to biologists, it is extremely compute intensive and suffers from prohibitive execution time. We propose an original parallelization scheme of the hmmsearch tool for FPGA technology. We show how to derive a flexible and generic hardware architecture which accelerates the hmmsearch main kernel by two orders of magnitude without modifying its original algorithm. Steven Derrien, Patrice Quinton |
ASAP | 1 |
| 2006 | Acceleration of a content-based image-retrieval application on the RDISK clusterabstractBecause of the growing use of multimedia content over Internet, content-based image retrieval (CBIR) has recently received a lot of interest. While accurate search techniques based on local image descriptors exist, they suffer from very long execution time. We propose to accelerate CBIR on the RDISK machine, a cluster of FPGA-enhanced hard-drives, that follows the philosophy of smart-disks. Our platform combines coarse and fine grain parallelism thanks to the concurrent use of the cluster nodes and of a programmable logic device. The implementation of the CBIR application on this mixed hardware/software platform follows a strict methodology, that was validated on realistic data-set (image database of more than 30,000 images). This methodology allows us to adapt the original algorithm to suit a hardware implementation, and to select the values of some key design parameters to maximize global performance. Our preliminary results indicate that speed-ups between 120 and 200 could be obtained for a cluster of 32 nodes compared with a software implementation running on a standard desktop PC. Auguste Noumsi, Steven Derrien, Patrice Quinton |
IPDPS | 2 |
| 2005 | Hardware/Software Interface for Multi-Dimensional Processor ArraysabstractOn most recent systems on chip, the performance bottleneck is the on-chip communication medium, bus or network. Multimedia applications require a large communication bandwidth between the processor and graphic hardware accelerators, hence an efficient communication scheme using burst mode is mandatory. In the context of data-flow hardware accelerators, we approach this problem as a classical resource-constrained problem. We explain how to use recent optimization techniques so as to define a conflict-free schedule of input/output for multi-dimensional processor arrays (e.g. 2D grids). This schedule is static and allows us to perform further optimizations such as grouping successive data in packets to operate in burst mode. We also present an effective VHDL implementation on FPGA and compare our approach to a run-time congestion resolution showing important gains in hardware area. Alain Darte, Steven Derrien, Tanguy Risset |
ASAP | 2 |
| 2005 | Cluster of re-configurable nodes for scanning large genomic banks
Stéphane Guyetant, Mathieu Giraud, Ludovic L'Hours, Steven Derrien, Stéphane Rubini, Dominique Lavenier, Frédéric Raimbault |
Parallel Comput. | 4 |
| 2001 | Combining Instruction and Loop Level Parallelism for FPGAs
Steven Derrien, Sanjay V. Rajopadhye, Susmita Sur-Kolay |
FCCM | 1 |
| 2001 | Loop Tiling for Reconfigurable Accelerators
Steven Derrien, Sanjay V. Rajopadhye |
FPL | 1 |
| 2000 | FCCMS and the Memory WallabstractAlthough there has been considerable work in the conventional general purpose processors community on how to tackle an important looming problem, we are not aware of any similar effort for custom computing machines. The aim of this paper is to analyze the state of the art, pose the relevant questions, and indicate a preliminary solution vis a vis the following question: how will custom computing machines face the memory wall. Steven Derrien, Sanjay V. Rajopadhye |
FCCM | 1 |
| 2000 | Approximating a Single Viewpoint in Panoramic Imaging DevicesabstractPanoramic cameras, which image a very large field of view, are useful devices for mobile robots that must move rapidly and securely in their environments. Recent panoramic cameras present a very wide field of view from a single viewpoint. A single viewpoint is useful in mobile robotics for a number of reasons, including perspective reprojection and stereo analysis. However, the requirement of single viewpoint for panoramic cameras restricts the optical and geometrical design of these devices. In this paper we present a method for approximating a single viewpoint in panoramic devices that allows much greater freedom in design. We illustrate the method with a compact catadioptric device using a spherical mirror and standard optics, and apply it to perspective reprojection. The resultant panoramic camera has been integrated as a surveillance device on a small mobile robot. Steven Derrien, Kurt Konolige |
ICRA | 1 |