EDBT 2026 Demo / reviewers in the wild / expert
Olivier Muller
dblp:84/6627
· DBLP profile ↗
23ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-4182-0502ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Network Folding for Resource-Efficient Implementation of Stream-Dataflow Deep Neural Network Inference on FPGAsabstractDeep Neural Networks (DNNs) have achieved state-of-the-art accuracy across various domains, often surpassing human performance. However, this accuracy necessitates significant computational and storage overhead, complicating their deployment on edge devices. To address these challenges, research has increasingly focused on optimizing hardware design constraints, including power efficiency, silicon area, and system scalability. Van-Quan Pham, Adrien Prost-Boucle, Olivier Muller, Frédéric Pétrot |
CF | 3 |
| 2026 | Future cardiovascular events prediction from invasive coronary angiography: A graph representation learning perspectiveabstractAbstract Improving risk stratification for coronary artery disease, the leading cause of death worldwide, continues to present a daily challenge in clinical practice, highlighting the urgent need for innovative approaches to early prediction of future cardiovascular events. In this work, we propose AngioGraphCAD, a deep learning based framework that employs graph neural networks to leverage geometry features and a masked attention to fuse geometry features from multiple coronary stenoses for future events prediction at both lesion and patient level from invasive coronary angiography. AngioGraphCAD is evaluated across two clinical cohorts at the lesion level and one datatset at the patient level, achieving superior performance compared to clinical measures. This is the first study that highlights the importance of geometry information in advancing future events prediction from invasive coronary angiography. Given the significance of the clinical question and the innovative nature of the proposed methodology, this work could pave the way for the development of an AI framework fueled by patient-specific data in cardiology, potentially revolutionizing personalized decision-making in managing coronary artery diseases for individual patients. Xiaowu Sun, Theofilos Belmpas, Ortal Yona Senouf, Emmanuel Abbe, Pascal Frossard, Bernard De Bruyne, Denise Auberson, Olivier Muller, Stéphane Fournier, Thabo Mahendiran, Dorina Thanou |
Medical Image Anal. | 8 |
| 2023 | Can Knowledge Transfer Techniques Compensate for the Limited Myocardial Infarction Data by Leveraging Hæmodynamics? An in silico Study
Riccardo Tenderini, Federico Betti 0002, Ortal Yona Senouf, Olivier Muller, Simone Deparis, Annalisa Buffa, Emmanuel Abbe |
AIME | 4 |
| 2023 | A Chisel Framework for Flexible Design Space Exploration through a Functional ApproachabstractAs the need for efficient digital circuits is ever growing in the industry, the design of such systems remains daunting, requiring both expertise and time. In an attempt to close the gap between software development and hardware design, powerful features such as functional and object-oriented programming have been used to define new languages, known as Hardware Construction Languages. In this article, we investigate the usage of such languages—more precisely, of Chisel—in the context of Design Space Exploration, and propose a novel design methodology to build custom and adaptable design flows. We apply an innovative functional approach to define flexible strategies for design space exploration, based on the composition of basic exploration steps, and provide a library of basic strategies along with a proof-of-concept framework—which we believe to be the first Chisel-based DSE framework. This framework fully integrates within the ecosystem of Chisel to allow users to define their DSE processes in the same framework (and language) they use to describe their designs. We demonstrate our methodology through several use cases, illustrating how our functional approach makes it possible to consider various metrics of interest when building exploration processes—in particular, we provide a quality of service -driven exploration example. The methodology presented in this work makes use of designers’ expertise to reduce the time required for hardware design, in particular for Design Space Exploration, and its application should ease digital design and enhance hardware developers’ productivity. Bruno Ferres, Olivier Muller, Frédéric Rousseau 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2021 | Integrating Quick Resource Estimators in Hardware Construction Framework for Design Space ExplorationabstractHardware design processes often come with time-consuming iteration loops, as feedbacks generally result of long synthesis runs. It is even more true when multiple different implementations need to be compared to perform Design Space Exploration (DSE). In order to accelerate such flows and increase agility of developers — closing the gap with software development methodologies — we propose to use quick feedback generating transforms based on RTL circuit analysis for quicker convergence of exploration. We also introduce an Hardware Construction Language (HCL) based methodology to build explorable circuit generators, and demonstrate such usage over a General Matrix Multiply (GEMM) Chisel implementation. We demonstrates that using RTL estimation early in the exploration process results in ×7 less synthesis runs and ×4.1 faster convergence than an exhaustive synthesis process, and still achieves state of the art performances when targetting a Xilinx VC709 FPGA. Bruno Ferres, Olivier Muller, Frédéric Rousseau 0001 |
RSP | 2 |
| 2021 | Hardware Context Switch-based Cryptographic Accelerator for Handling Multiple StreamsabstractThe confidentiality and integrity of a stream has become one of the biggest issues in telecommunication. The best available algorithm handling the confidentiality of a data stream is the symmetric key block cipher combined with a chaining mode of operation such as cipher block chaining (CBC) or counter mode (CTR). This scheme is difficult to accelerate using hardware when multiple streams coexist. This is caused by the computation time requirement and mainly by management of the streams. In most accelerators, computation is treated at the block-level rather than as a stream, making the management of multiple streams complex. This article presents a solution combining CBC and CTR modes of operation with a hardware context switching. The hardware context switching allows the accelerator to treat the data as a stream. Each stream can have different parameters: key, initialization value, state of counter. Stream switching was managed by the hardware context switching mechanism. A high-level synthesis tool was used to generate the context switching circuit. The scheme was tested on three cryptographic algorithms: AES, DES, and BC3. The hardware context switching allowed the software to manage multiple streams easily, efficiently, and rapidly. The software was freed of the task of managing the stream state. Compared to the original algorithm, about 18%–38% additional logic elements were required to implement the CBC or CTR mode and the additional circuits to support context switching. Using this method, the performance overhead when treating multiple streams was low, and the performance was comparable to that of existing hardware accelerators not supporting multiple streams. Arif Sasongko, I. M. Narendra Kumara, Arief Wicaksana, Frédéric Rousseau 0001, Olivier Muller |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2020 | (System)Verilog to Chisel Translation for Faster Hardware DesignabstractBringing agility to hardware developments has been a long-running goal for hardware communities struggling with limitations of current hardware description languages such as (System)Verilog or VHDL. The numerous recent Hardware Construction Languages such as Chisel are providing enhanced ways to design complex hardware architectures with notable academic and industrial successes. While the latter environments are now mature and perfectly suited for brand new projects, migrating partially or entirely existing Verilog code-base proves to be a challenging and very time-consuming process. Successful migrations need to be able to leverage finely tuned existing hardware descriptions as a basis to build complex systems through simple iterations. This article introduces sv2chisel, an open-source automated (System)Verilog to Chisel translator as entry point for this iterative migration processes. Our tool achieved the proper translation, with on-par resource usage of a real-world production FPGA design at OVHcloud as well as two independent open-source Verilog projects: a MIPS core and the size-optimized 32-bit RISC-V core PicoRV32. Jean Bruant, Pierre-Henri Horrein, Olivier Muller, Tristan Groleat, Frédéric Pétrot |
RSP | 3 |
| 2019 | Efficient Decompression of Binary Encoded Balanced Ternary SequencesabstractA balanced ternary digit, known as a trit, takes its values in {-1, 0, 1}. It can be encoded in binary as {11, 00, 01} for the direct use in digital circuits. In this brief, we study the decompression of a sequence of bits into a sequence of binary encoded balanced ternary digits. We first show that it is useless, in practice, to compress sequences of more than five ternary values. We then provide two mappings, one to map 5 bits to 3 trits and one to map 8 bits to 5 trits. Both mappings were obtained by human analysis and lead to Boolean implementations that compare quite favorably with others obtained by tweaking assignment or encoding optimization tools. However, mappings that lead to better implementations may be feasible. Olivier Muller, Adrien Prost-Boucle, Alban Bourge, Frédéric Pétrot |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | Message-Oriented Devices on FPGAsabstractEmbedded systems increasingly include an FPGA for performance or power efficiency. Fortunately, FPGA makers provide efficient tools to develop and assemble multiple Intellectual Properties (IPs) as devices on FPGAs. Unfortunately, the integration challenge does not stop there, hardware devices are only usable if they have available software device drivers executing on the processor subsystem. To better approach this end-to-end integration challenge, across both hardware and software, we argue that both sides need to evolve. We propose to take a step towards message-based interfaces for hardware devices integrated on an FPGA. Our goal is to deliver to the FPGA market the plug-and-play value of the USB stack with essentially no performance overhead, negligible power increase, and a reasonable surface cost. We have implemented our proposal on the Xilinx Zynq SoC, combining ARM cores and an FPGA, demonstrating the feasibility of the approach. Thomas Baumela, Olivier Gruber, Olivier Muller, Frédéric Pétrot |
RSP | 3 |
| 2017 | Prototyping dynamic task migration on heterogeneous reconfigurable systemsabstractReconfigurable devices, such as FPGAs, have been known to offer an excellent performance and a high efficiency in computation. Due to their improving capacity and more efficient architecture recently, there are growing interests in using FPGAs as coprocessors in reconfigurable systems. However, FPGAs still lack the support in dynamic scheduling, e.g. to manage multiple tasks or users in a system. Performing runtime task relocation or load distribution is not possible unless the reconfigurable system supports dynamic task migration. Such ability requires the automation of configuration and context management in reconfigurable architecture, which is not available in the existing solutions. Arief Wicaksana, Alban Bourge, Olivier Muller, Arif Sasongko, Frédéric Rousseau 0001 |
RSP | 3 |
| 2016 | HLS-Based Methodology for Fast Iterative Development Applied to Elliptic Curve ArithmeticabstractHigh-Level Synthesis (HLS) is used by hardware developers to achieve higher abstraction in circuit descriptions. In order to shorten the hardware development time via HLS, we present an adjustment of the Iterative and Incremental Design (IID) methodology, frequently used in software development. In particular, our methodology is relevant for the development of applications with unusual complexity: the method was applied here to the development of large modular arithmetic, commonly used for cryptography applications (e.g., Elliptic Curves). Rapid feedback on circuit characteristics is used to evaluate deep architectural changes in short time, greatly reducing the time-to-market with respect to hand-made designs. In addition, our approach is highly flexible, since the same generic high-level description can be used to produce an entire set of circuits, each with different area/performance trade-offs. Thanks to the proposed approach, any change to the initial specification (e.g., the curve used) is also very fast, while it may require a large effort in the case of hand-made designs. Simon Pontié, Alban Bourge, Adrien Prost-Boucle, Paolo Maistri, Olivier Muller, Régis Leveugle, Frédéric Rousseau 0001 |
DSD | 5 |
| 2016 | Demonstration of a context-switch method for heterogeneous reconfigurable systemsabstractNowadays, FPGAs are integrated in high-performance computing systems, servers, or even used as accelerators in System-on-Chip (SoC) platforms. Since the execution is performed in hardware, FPGA gives much higher performance and lower energy consumption compared to most microprocessor-based systems. However, the room to improve FPGA performance still exists, e.g. when it is used by multiple users. In multi-user approaches, FPGA resources are shared between several users. Therefore, one must be able to interrupt a running circuit at any given time and continue the task at will. An image of the state of the running circuit (context) is saved during interruption and restored when the execution is continued. The ability to extract and restore the context is known as context-switch. Arief Wicaksana, Alban Bourge, Olivier Muller, Frédéric Rousseau 0001 |
FPL | 3 |
| 2016 | On-board non-regression test of HLS tools targeting FPGAabstractHigh-Level Synthesis (HLS) has opened an opportunity for software programmers to target FPGA more rapidly. When developing HLS tools, tests are desirable to ensure their function, reliability and performance. When modifications are applied to a tool, Non-Regression Test (NRT) asserts that the changes have intended effect while Regression Test (RT) verifies that the tool still performs correctly without unwanted behaviour. Arief Wicaksana, Adrien Prost-Boucle, Olivier Muller, Frédéric Rousseau 0001, Arif Sasongko |
RSP | 3 |
| 2016 | Generating Efficient Context-Switch Capable Circuits through Autonomous Design FlowabstractCommercial off-the-shelf (COTS) Field-Programmable Gate Arrays (FPGAs) are becoming increasingly powerful. In addition to their huge hardware resources, they are also integrated into complete systems on chips (SOCs), e.g., in the latest Xilinx Zynq or Altera Stratix platforms. However, cooperation between FPGAs and their surroundings, and the flexibility of hardware task management could still be improved. For instance, mechanisms have yet to be automated to allow multi-user approaches. A reconfigurable resource can be shared between applications or users only if it has a context-switch ability allowing applications to be paused and resumed in response to system demands. Here, we present a high-level synthesis (HLS) design flow producing a context-switch-capable circuit. The design flow manipulates the intermediate representation of an HLS tool to build the context extraction mechanism and to optimize performance for the circuit produced. The method is based on efficient checkpoint selection and insertion of a powerful scan-chain into the initial circuit. This scan-chain can extract flip-flops or memory content. Experiments with the system produced show that it has a low hardware overhead for many benchmark applications, and that the hardware added has a negligible impact on application performance. Comparisons with current standard methods highlight the efficiency of our contributions. Alban Bourge, Olivier Muller, Frédéric Rousseau 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2015 | Automatic High-Level Hardware Checkpoint Selection for Reconfigurable SystemsabstractModern FPGAs provide great computational power and flexibility but there is still room for improving their performances. For example multi-user approaches are particularly underdeveloped as they require specific mechanisms still to be automated. Sharing an FPGA resource between applications or users requires a context switch ability. The latter enables pausing and resuming applications at system demand. This paper presents a method that automatically selects a good execution point, called hardware checkpoint, to perform a context switch on an FPGA. The method relies on a static analysis of the finite state machine of a circuit to select the checkpoint states. The obtained selection ensures that the context switch mechanism respects a given latency and tries to minimize the mechanism costs. The method takes advantage of its integration in an open-source HLS tool and preliminary results highlight its efficiency. Alban Bourge, Olivier Muller, Frédéric Rousseau 0001 |
FCCM | 2 |
| 2015 | A Novel Method for Enabling FPGA Context-Switch (Abstract Only)abstractModern FPGAs provide great computational power, flexible resources and a versatile environment. Managing to obtain the best of these three worlds is rather complicated given the actual design flows. Our work focus on enabling task multiplexing, as part of a more flexible FPGA usage. Task multiplexing in FPGAs raises indeed a lot of questions. Multiplexing the usage of a reconfigurable fabric is leading to a better utilization of its surface because it offers to share its resources not only in space (number of slices allocated to a task) but also in time (tasks are allowed in time slots). The base mechanism known as context-switch consists in removing a task after its allowed time slot has passed. The first step toward efficiently multiplex tasks in a reconfigurable fabric is to decide when this removal will have the least possible impact on the system. This poster presents our preliminary results concerning what we consider as necessary in order to enable such a feature. Our work focus on finding automatically the best instants of the task execution in order to effectively remove a running task from the FPGA, taking into account the time needed to extract a relevant context necessary to restart it later. This instant selection is performed at a high level of abstraction, enabling us to make choices with an accurate knowledge of the task nature and specificities. The second part of this poster presents the entire mechanism which makes use of the previously selected slots in order to switch between tasks. Alban Bourge, Olivier Muller, Frédéric Rousseau 0001 |
FPGA | 2 |
| 2014 | Fast and standalone Design Space Exploration for High-Level Synthesis under resource constraints
Adrien Prost-Boucle, Olivier Muller, Frédéric Rousseau 0001 |
J. Syst. Archit. | 2 |
| 2013 | A Fast and Autonomous HLS Methodology for Hardware Accelerator Generation under Resource ConstraintsabstractThis paper presents a new methodology for hardware accelerator generation, in the context of High Level Synthesis (HLS) for Field Programmable Gate Array (FPGA) components. The very high computing capacity available in the latest FPGA makes them choice targets in High-Performance Computing (HPC) as well as embedded systems. For a much wider adoption of FPGA as general-purpose computing devices, the proposed HLS design flow leverages the users from all issues related to circuit structure fine-tuning. The HLS methodology is autonomous and produces RTL descriptions quickly, under only global resource and frequency constraints. This is achieved by performing incremental transformations of the input design description. The low complexity of the Design Space Exploration (DSE) algorithm and its good usage of all internal circuit structure constraints, make this HLS methodology very fast and able to generate pertinent solutions. Moreover, the generated circuit is designed to fit into the targeted FPGA or a given partition of it. Such a methodology leads to autonomous, fast and transparent DSE, all these issues known to limit the use of HLS and FPGA. Results on several benchmarks highlight the capabilities of our DSE methodology. The results show a high generation speed-up compared to other existing HLS approaches, while preserving correct performance of the generated circuits. Adrien Prost-Boucle, Olivier Muller, Frédéric Rousseau 0001 |
DSD | 2 |
| 2012 | HCM: An abstraction layer for seamless programming of DPR FPGAabstractWell-known for its efficient computing capabilities, FPGA-based architectures also have the potential for high flexibility with dynamic reconfiguration features. Yet, writing applications on these architectures is laborious, poorly portable and hardly scalable to multi-user and/or multi-FPGA systems, mainly because of a mixture of application related code and flexibility management code. In this paper, we propose a new abstraction layer, called Hardware Component Manager (HCM), which clearly separates the allocation of a hardware function from the control of a reconfiguration procedure, and guarantees the security of coexisting configurations. The implementation of this HCM layer on realistic simulation platforms demonstrates its ability to ease the management of FPGA flexibility while preserving performance and ensuring hardware function protection. HCM implementation and its simulation environment are open-source in the hope of reuse by the community. Olivier Muller, Pierre-Henri Horrein, Frédéric Pétrot |
FPL | 2 |
| 2009 | From Parallelism Levels to a Multi-ASIP Architecture for Turbo DecodingabstractEmerging digital communication applications and the underlying architectures encounter drastically increasing performance and flexibility requirements. In this paper, we present a novel flexible multiprocessor platform for high throughput turbo decoding. The proposed platform enables exploiting all parallelism levels of turbo decoding applications to fulfill performance requirements. In order to fulfill flexibility requirements, the platform is structured around configurable application-specific instruction-set processors (ASIP) combined with an efficient memory and communication interconnect scheme. The designed ASIP has an single instruction multiple data (SIMD) architecture with a specialized and extensible instruction-set and 6-stages pipeline control. The attached memories and communication interfaces enable its integration in multiprocessor architectures. These multiprocessor architectures benefit from the recent shuffled decoding technique introduced in the turbo-decoding field to achieve higher throughput. The major characteristics of the proposed platform are its flexibility and scalability which make it reusable for all simple and double binary turbo codes of existing and emerging standards. Results obtained for double binary WiMAX turbo codes demonstrate around 250 Mb/s throughput using 16-ASIP multiprocessor architecture. Olivier Muller, Amer Baghdadi, Michel Jézéquel |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2007 | Butterfly and benes-based on-chip communication networks for multiprocessor turbo decodingabstractSeveral research activities have recently emerged aiming to propose multiprocessor implementations in order to achieve flexible and high throughput parallel iterative decoding. Besides application algorithm optimizations and application-specific instruction-set processor design, the on-chip communication network constitutes a major issue in this application domain. In this paper, the authors propose to use multistage interconnection networks as on-chip communication networks for parallel turbo decoding. Adapted benes and butterfly networks are proposed with detailed hardware implementation of network interfaces, routers, and topologies. In addition, appropriate packet format and routing for interleaved/deinterleaved extrinsic information exchanges are proposed. The flexibility of these on-chip communication networks enables their use for all turbo code standards and constitutes a promising feature for their reuse for any similar interleaved/deinterleaved iterative communication profile Hazem Moussa, Olivier Muller, Amer Baghdadi, Michel Jézéquel |
DATE | 2 |
| 2006 | ASIP-based multiprocessor SoC design for simple and double binary turbo decodingabstractThis paper presents a new multiprocessor platform for high throughput turbo decoding. The proposed platform is based on a new configurable ASIP combined with an efficient memory and communication interconnect scheme. This application-specific instruction-set processor has an SIMD architecture with a specialized and extensible instruction-set and 5-stages pipeline control. The attached memories and communication interfaces enable the design of efficient multiprocessor architectures. These multiprocessor architectures benefit from the recent shuffling technique introduced in the turbo-decoding field to reduce communication latency. The major characteristics of the proposed platform are its flexibility and scalability which make it reusable for various standards and operating modes. Results obtained for double binary DVB-RCS turbo codes demonstrate a 100 Mbit/s throughput using 16-ASIP multiprocessor architecture Olivier Muller, Amer Baghdadi, Michel Jézéquel |
DATE | 1 |
| 2006 | On the Parallelism of Convolutional Turbo Decoding and Interleaving InterferenceabstractIn forward error correction, convolutional turbo codes were introduced to increase error correction capability approaching the Shannon bound. Decoding of these codes, however, is an iterative process requiring high computation rate and latency. Thus, in order to achieve high throughput and to reduce latency, crucial in emerging digital communication applications, parallel implementations become mandatory. This paper explores and analyses existing parallelism techniques in convolutional turbo decoding with the BCJR algorithm. For component-decoder parallelism, we illustrate the influence of interleaving scheme and we propose new interleaving rules allowing to maximize parallelism efficiency. Olivier Muller, Amer Baghdadi, Michel Jézéquel |
GLOBECOM | 1 |