VLDB 2026 Research / reviewers in the wild / expert
Frédéric Pétrot
dblp:41/1007
· DBLP profile ↗
80ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0003-0624-7373ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 68 · 3 first-author · 10 since 2021Software engineering, systems software and programming languages · 21 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Network Folding for Resource-Efficient Implementation of Stream-Dataflow Deep Neural Network Inference on FPGAsabstractDeep Neural Networks (DNNs) have achieved state-of-the-art accuracy across various domains, often surpassing human performance. However, this accuracy necessitates significant computational and storage overhead, complicating their deployment on edge devices. To address these challenges, research has increasingly focused on optimizing hardware design constraints, including power efficiency, silicon area, and system scalability. Van-Quan Pham, Adrien Prost-Boucle, Olivier Muller, Frédéric Pétrot |
CF | 4 |
| 2025 | Experimental Software and Hardware Evaluation of Ad-Hoc Constant Division RoutinesabstractDividing by a constant is an operation that is often needed in algorithms, be they implemented in software or in hardware. Given the fact that general purpose processors have had fast hardware multipliers for decades, the solution for software is to multiply by a compile time computed reciprocal and perform some adjustment. However, for hardware implementations, in particular with relatively small and exotic bit sizes, shift-and-add solutions might be worth looking at. In this paper, we report our study on implementing constant unsigned division based on Li's work. We found that a few of his algorithms are wrong and propose corrections that need a bit more computations. For software, we show that the approach can be useful only for low-end microcontrollers. For hardware, our FPGA and ASIC synthesis outline that it has good scalability, although being not very efficient for small dividends. As delay and area are very dependent on the value of the divisors, this approach appears as yet another possibility to choose from when looking on how to divide by a given constant. Frédéric Pétrot |
ARITH | 1 |
| 2025 | Architecting Value Prediction around In-Order ExecutionabstractIn the search for performance, in-order execution cannot expect to prevail as older long latency instructions prevent younger ones from issuing. Although stall-on-use processors allow independent instructions to issue in the shadow of a cache miss, the compiler cannot always find enough independent work to keep pipeline resources busy. In this paper, we study how both value prediction based on address prediction and direct value prediction can be built into an in-order pipeline to unlock significant performance. We further show that the in-order execution property provides advantages in that the pipeline may speculate aggressively without suffering from any recovery penalty. Finally, we combine this data speculation infrastructure with a reworked cache hierarchy that relies on a fast first level cache that can be written speculatively. We show that such an in-order pipeline can reach a performance level that is comparable to an equally - although moderately - wide out-of-order processor, without requiring support for partial out-of-order execution such as out-of-order memory hazard handling or full-fledged register renaming. Overall, we increase the performance of a 32 -entry scoreboard, 4-issue in-order processor based on a scaled up Open Hardware Group CVA6 by 38.4% (geomean), achieving $\mathbf{8 6. 7 \%}$ and $\mathbf{4 6. 3 \%}$ of the gains brought by comparable out-of-order processors featuring 32/16-entry and 64/32-entry Reorder Buffer and scheduler, respectively. Pierre Ravenel, Arthur Perais, Benoît Dupont de Dinechin, Frédéric Pétrot |
HPCA | 4 |
| 2025 | Address/Data Instruction Steering in Clustered General Purpose ProcessorsabstractAlthough they differentiate between integer and floating-point datum, modern Instruction Set Architectures and their implementations do not differentiate integer datum used to address memory from integer datum used in purely arithmetic and logical computations. This is a perfectly reasonable choice as addresses are, in fact, integral quantities. However, in many cases, there is already a fundamental difference between addresses and integer data: Their width. As computer systems moved from 16 to 32, then to 64-bit pointers, with a potential future where 128-bit might be used for specific systems, the data width required to compute a given output with a given algorithm has remained the same, e.g., an ASCII character is still represented on a byte. This work aims to leverage this dichotomy to revisit hardware clustering, a well-known microarchitectural technique used to mitigate the cost of scaling processor backend structures by dividing the backend into several mostly independent execution clusters. We show that by treating instructions as manipulating addresses or data and steering them to a “data” or an “address” cluster accordingly, reasonable cluster load balancing can be achieved without the need for complex steering policies that can lead to performance on par with the baseline with limited hardware overhead. Moreover, we highlight two possible optimizations stemming from this distribution. First, the registers of the “address” cluster can easily be compressed thanks to address spatial and temporal locality. Second, if a processor requires a large address space but only processes narrow data (e.g., 32-bit data with 64-bit pointers or 64-bit data with 128-bit pointers), the “data” cluster datapath can be kept narrower than the “address” cluster datapath. Chandana S. Deshpande, Arthur Perais, Frédéric Pétrot |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | Page size exploration for RISC-V systems: the case for HPCabstractThe page size used for virtual to physical address translation has globally not changed since the late 1960’s: the IBM 360, circa 1964, already had 4 KiB pages. This 4 KiB page size has proven to be incredibly robust given the changes in processor architectures, workloads behavior, memory size, and access patterns. However, with 64-bit registers, 57-bit virtual addresses, and increasingly bigger physical memories, we have to ask ourselves whether 4 KiB is still an adequate page size for modern workloads on modern machines. Inherently, the page size has an influence on (a) the miss rate of the translation lookaside buffer, the cache that contains the recently used virtual to physical translations, and (b) the memory allocated by the system versus the memory actually used by a process. The page size also constraints some microarchitectural choices, such as cache design, which impacts the overall performance and energy efficiency. We focus more particularly on High Performance Computing (HPC) applications because they are extremely demanding in terms of memory, and are indicative of future general-purpose needs.In this paper, we empirically study the evolution of the miss rate and memory occupancy with respect to the page size, and conclude that a page size of 32 KiB is better suited for current HPC systems. We also propose a page table scheme for RISC-V-based HPC systems based on our observations and discuss its benefits. Eduardo Tomasi, César Fuguet Tortolero, Christian Fabre, Frédéric Pétrot |
RSP | 4 |
| 2023 | Quantization Modes for Neural Network Inference: ASIC Implementation Trade-offsabstractAs deep neural networks migrate close to the sensors, accuracy cannot be the single target anymore: inference tasks must also be highly energy efficient. For embedded devices, the power budget for one inference is typically in the range of a few tens of µW to single-digit mW. We have three levers of action for that: computational workload, number of values to memorize-be they network parameters or intermediate activation results-, and implementation strategy. Given the fact that Application Specific Integrated Circuits are two orders of magnitude more power-efficient than processors for a given technology node, the latter issue is solved using ad-hoc hardware implementations. For the two first former issues, we detail and compare in this work different existing quantization approaches, since reducing the number of bits of the weights and activations reduces computation complexity and storage needs. In addition, we also propose two new modes specifically aiming at low silicon footprint and power optimized hardware implementations that still provide an accuracy in par with existing works. We report the area/power and accuracy trade-offs all theses quantization modes provide when targeting low to ultra-low power devices. The evaluation is done using STMicroelectronics 40nm technology. It shows that the best results vary depending on the dataset and network architecture, which calls for application and quantization aware network architecture search. Nathan Bain, Roberto Guizzetti, Emilien Taly, Ali Oudrhiri, Bruno Paille, Pascal Urard, Frédéric Pétrot |
IJCNN | 7 |
| 2022 | Low-precision logarithmic arithmetic for neural network acceleratorsabstractResource requirements for hardware acceleration of neural networks inference is notoriously high, both in terms of computation and storage. One way to mitigate this issue is to quantize parameters and activations. This is usually done by scaling and centering the distributions of weights and activations, on a kernel per kernel basis, so that a low-precision binary integer representation can be used. This work studies low-precision logarithmic number system (LNS) as an efficient alternative. Firstly, LNS has more dynamic than fixed-point for the same number of bits. Thus, when quantizing MNIST and CIFAR reference networks without retraining, the smallest format size achieving top-1 accuracy comparable to floating-point is 1 to 3 bits smaller with LNS than with fixed-point. In addition, it is shown that the zero bit of classical LNS is not needed in this context, and that the sign bit can be saved for activations. The proposed LNS neuron is detailed and its implementation on FPGA is shown to be smaller and faster than a fixed-point one for comparable accuracy. Secondly, low-precision LNS enables efficient inference architectures where 1 / multiplications reduce to additions; 2/ the weighted inputs are converted to classical linear domain, but the tables needed for this conversion remain very small thanks to the low precision; and 3/ the conversion of the output activation back to LNS can be merged with an arbitrary activation function. Maxime Christ, Florent de Dinechin, Frédéric Pétrot |
ASAP | 3 |
| 2022 | Fast simulation of future 128-bit architecturesabstractWhether 128-bit architectures will some day hit the market or not is an open question. There is however a trend towards that direction: virtual addresses grew from 34 to 48 bits in 1999 and then to 57 bits in 2019. The impact of a virtually infinite addressable space on software is hard to predict, but it will most likely be major. Simulation tools are therefore needed to support research and experimentation for tooling and software. In this paper, we present the implementation of the 128-bit extension of the RISC-V architecture in the QEMU functional simulator and report first performance evaluations. On our limited set of programs, simulation is slowed down by a factor of at worst 5 compared to 64-bit simulation, making the tool still usable for executing large software codes. Fabien Portas, Frédéric Pétrot |
DATE | 2 |
| 2022 | A Case for Second-Level Software Cache Coherency on Many-Core AcceleratorsabstractCache and cache-coherence are major aspects of today's high performance computing. A cache stores data as cache-lines of fixed size, and coherence between caches is guaranteed by the cache-coherence protocol which operates on fixed size coherency-blocks. In such systems cache-lines and coherency-blocks are usually the same size and are relatively small, typically 64 bytes. This size choice is a trade-off selected for general-purpose computing: it minimizes false-sharing while keeping cache-maintenance traffic low. False-sharing is considered an unnecessary cache-coherence traffic and it decreases performances. However, for dedicated accelerator this trade-off may not be appropriate: hardware in charge of cache-coherence is expensive and not well exploited by most accelerator applications as by construction these applications minimize false-sharing. This paper investigates the possibility of an alternative trade-off of cache-coherency and cache-maintenance block size for many-core accelerators, by decoupling coherency-block and cache-lines sizes. Interests, advantages and difficulties are presented and discussed in this paper. Then we also discuss needs of software and hardware modifications in prototypes and the capability of such prototypes to evaluate different coherence-block sizes. Arthur Vianès, Frédéric Pétrot, Frédéric Rousseau 0001 |
RSP | 2 |
| 2021 | Arbitrary and Variable Precision Floating-Point Arithmetic Support in Dynamic Binary TranslationabstractFloating-point hardware support has more or less been settled 35 years ago by the adoption of the IEEE 754 standard. However, many scientific applications require higher accuracy than what can be represented on 64 bits, and to that end make use of dedicated arbitrary precision software libraries. To reach a good performance/accuracy trade-off, developers use variable precision, requiring e.g. more accuracy as the computation progresses. Hardware accelerators for this kind of computations do not exist yet, and independently of the actual quality of the underlying arithmetic computations, defining the right instruction set architecture, memory representations, etc, for them is a challenging task. We investigate in this paper the support for arbitrary and variable precision arithmetic in a dynamic binary translator, to help gain an insight of what such an accelerator could provide as an interface to compilers, and thus programmers. We detail our design and present an implementation in QEMU using the MPRF library for the RISC-V processor1. Marie Badaroux, Frédéric Pétrot |
ASP-DAC | 2 |
| 2021 | Simulation of Ideally Switched Circuits in SystemCabstractModeling and simulation of power systems at low levels of abstraction is supported by specialized tools such as SPICE and MATLAB. But when power systems are part of larger systems including digital hardware and software, low-level models become over-detailed; at the system level, models must be simple and execute fast. We present an extension to SystemC that relies on efficient modeling, simulation, and synchronization strategies for Ideally Switched Circuits. Our solution enables designers to specify circuits and to jointly simulate them with other SystemC hardware and software models. We test our extension with three power converter case studies and show a simulation speed-up between 1.2 and 2.7 times while preserving accuracy when compared to the reference tool. This work demonstrates the suitability of SystemC for the simulation of heterogeneous models to meet system-level goals such as validation, verification, and integration. Breytner Fernández-Mesa, Liliana Andrade, Frédéric Pétrot |
ASP-DAC | 3 |
| 2021 | Seamless Compiler Integration of Variable Precision Floating-Point ArithmeticabstractFloating-Point (FP) units in processors are generally limited to supporting a subset of formats defined by the IEEE 754 standard. As a result, high-efficiency languages and optimizing compilers for high-performance computing only support IEEE standard types and applications needing higher precision involve cumbersome memory management and calls to external libraries, resulting in code bloat and making the intent of the program unclear. We present an extension of the C type system that can represent generic FP operations and formats, supporting both static precision and dynamically variable precision. We design and implement a compilation flow bridging the abstraction gap between this type system and low-level FP instructions or software libraries. The effectiveness of our solution is demonstrated through an LLVM-based implementation, leveraging aggressive optimizations in LLVM including the Polly loop nest optimizer, which targets two backend code generators: one for the ISA of a variable precision FP arithmetic coprocessor, and one for the MPFR multi-precision floating-point library. Our optimizing compilation flow targeting MPFR outperforms the Boost programming interface for the MPFR library by a factor of 1.80 × and 1.67 × in sequential execution of the Poly Bench and RAJAPerf suites, respectively, and by a factor of 7.62 x on an 8-core (and 16-thread) machine for RAJAPerf in OpenMP. Tiago T. Jost, Yves Durand, Christian Fabre, Albert Cohen 0001, Frédéric Pétrot |
CGO | 5 |
| 2021 | To Pin or Not to Pin: Asserting the Scalability of QEMU Parallel ImplementationabstractDue to its speed in cross-executing sequential code, dynamic binary translation is the unchallenged technology for full system-level simulation. Among the translators, QEMU has become the de facto solution. It introduced parallel host execution of the target cores a few years ago for the ARM instruction set architecture and this support is now also available, among others, for RISC-V. Given the popularity of these instruction sets in multi and many-core systems, assessing the scalability of their parallel implementation makes sense. In this paper, we use a subset of the PARSEC benchmark to measure the execution time of QEMU’s parallel implementation, to which we added the ability to pin a target processor to a host core or hardware thread. We report the results of a wealth of experiments we performed on a 16-core/32-thread x86-64 SMP machine. They show that the support of parallelism in QEMU scales well, and that, somewhat counter intuitively, pinning does not improve performance. Marie Badaroux, Saverio Miroddi, Frédéric Pétrot |
DSD | 3 |
| 2021 | Synchronization of Continuous Time and Discrete Events Simulation in SystemCabstractModeling and simulation of cyber-physical devices is difficult because of their heterogeneity: they mix discrete-event (DE) hardware and software with continuous time (CT) physical processes. DEs simulation progresses by discrete timesteps while CT simulation does so in a time continuum; despite different time representations, data must be transmitted between both domains at precise points in time. Synchronizing these domains in an efficient and accurate manner is the core problem of CT and DEs simulation. As a design tool, SystemC AMS introduces a synchronization strategy that is based on fixed timesteps that proves useful for a large set of use cases. But this strategy generates inaccuracies that cannot be overcome without penalizing simulation speed. In this article, we propose a new CT and DEs synchronization algorithm on top of the SystemC framework and prove its causality, completeness, and liveness. We also propose an adaptive algorithm to adjust the synchronization step to provide near to optimum simulation speed. We test our solution with three demanding case studies that include nonlinear equations, Zeno behavior, and high sensitivity to accuracy errors. Results show that our algorithm circumvents these challenges, attains high accuracy with respect to established tools, and improves simulation speed. This work aims at enlarging the modeling and simulation capabilities of SystemC as a heterogeneous design tool. Breytner Fernández-Mesa, Liliana Andrade, Frédéric Pétrot |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | VP Float: First Class Treatment for Variable Precision Floating Point ArithmeticabstractOptimizing compilers for high performance computing only support IEEE~754 floating-point (FP) types and applications needing higher precision involve cumbersome memory management and calls to external libraries. We introduce an extension of the C type system to represent variable-precision FP arithmetic, supporting both static and dynamically variable precision. We design and implement a compilation flow bridging the abstraction gap between this type system and hardware FP instructions or software libraries. We demonstrate the effectiveness of our solution by enabling the full range of LLVM optimizations and leveraging two backend code generators: one for the ISA of a variable precision FP arithmetic coprocessor, and one for the MPFR multi-precision FP library. Both targets support the static and dynamically adaptable precision of our type system. On the PolyBench suite, our optimizing compilation flow targeting MPFR is shown to outperform the Boost programming interface for the MPFR library. Tiago T. Jost, Yves Durand, Christian Fabre, Albert Cohen 0001, Frédéric Pétrot |
PACT | 5 |
| 2020 | Accurate and Efficient Continuous Time and Discrete Events Simulation in SystemCabstractThe AMS extensions of SystemC emerged to aid the virtual prototyping of continuous time and discrete event heterogeneous systems. Although useful for a large set of use cases, synchronization of both domains through a fixed timestep generates inaccuracies that cannot be overcome without penalizing simulation speed. We propose a direct, optimistic, and causal synchronization algorithm on top of the SystemC kernel that explicitly handles the rich set of interactions that occur in the domain interface. We test our algorithm with a complex nonlinear automotive use case and show that it breaks the described accuracy and efficiency trade-off. Our work enlarges the applicability range of SystemC AMS based design frameworks. Breytner Fernández-Mesa, Liliana Andrade, Frédéric Pétrot |
DATE | 3 |
| 2020 | Low Power Tiny Binary Neural Network with improved accuracy in Human Recognition SystemsabstractHuman Activity Recognition requires very high accuracy to be effectively employed into practical applications, ranging from elderly care to microsurgical devices. The highest accuracies are achieved by Deep Learning models, but these are not easily deployable in handheld or wearable devices with very constrained resources. We therefore present a new HAR system suitable for a compact FPGA implementation. A new Binarized Neural Network (BNN) architecture achieves the classification based on data from a single tri-axial accelerometer. From our experiments, the effect of gravity and the unknown orientation of the sensor cause a degradation of the accuracy. In order to compensate for these issues, we propose a HW-friendly algorithm to pre-process the raw acceleration signal. Moreover, the very low power and hardware friendly BNN has been trained and validated on the PAMAP2 dataset, for which the pre-processing operations increase the accuracy from 51% to 99% in the best case. Aiming for a low-power design, we designed both a custom circuit to perform the pre-processing operations and a hardware accelerator for the BNN. The design on FPGA features a power dissipation of 72 mW and occupies 6788 LUTs. Antonio De Vita, Danilo Pau, Luigi Di Benedetto, Alfredo Rubino, Frédéric Pétrot, Gian Domenico Licciardo |
DSD | 5 |
| 2020 | (System)Verilog to Chisel Translation for Faster Hardware DesignabstractBringing agility to hardware developments has been a long-running goal for hardware communities struggling with limitations of current hardware description languages such as (System)Verilog or VHDL. The numerous recent Hardware Construction Languages such as Chisel are providing enhanced ways to design complex hardware architectures with notable academic and industrial successes. While the latter environments are now mature and perfectly suited for brand new projects, migrating partially or entirely existing Verilog code-base proves to be a challenging and very time-consuming process. Successful migrations need to be able to leverage finely tuned existing hardware descriptions as a basis to build complex systems through simple iterations. This article introduces sv2chisel, an open-source automated (System)Verilog to Chisel translator as entry point for this iterative migration processes. Our tool achieved the proper translation, with on-par resource usage of a real-world production FPGA design at OVHcloud as well as two independent open-source Verilog projects: a MIPS core and the size-optimized 32-bit RISC-V core PicoRV32. Jean Bruant, Pierre-Henri Horrein, Olivier Muller, Tristan Groleat, Frédéric Pétrot |
RSP | 5 |
| 2019 | Multi-Triggered Embedded Software Code Generation for Electrical Metering and Protection ApplicationsabstractCode generation can be an effective way to improve quality of a product and reduce development time. But applied naively to multi-triggered applications such as electrical metering and protection algorithms, performance can be drastically degraded. In this paper we present a solution to model a multi-triggered application in Mathworks Simulink for efficient code generation using a buffering mechanism specified in the model, with little overhead in usage of resources. The performance of the presented solution is compared with different modeling approaches and with a reference manually-coded implementation. Louis Bonicel, Roland Bohrer, Benoit Leprettre, Frédéric Rousseau 0001, Frédéric Pétrot |
RSP | 5 |
| 2019 | Loop aware CFG matching strategy for accurate performance estimation in IR-level native simulation
Omayma Matoussi, Frédéric Pétrot |
Integr. | 2 |
| 2019 | Efficient Decompression of Binary Encoded Balanced Ternary SequencesabstractA balanced ternary digit, known as a trit, takes its values in {-1, 0, 1}. It can be encoded in binary as {11, 00, 01} for the direct use in digital circuits. In this brief, we study the decompression of a sequence of bits into a sequence of binary encoded balanced ternary digits. We first show that it is useless, in practice, to compress sequences of more than five ternary values. We then provide two mappings, one to map 5 bits to 3 trits and one to map 8 bits to 5 trits. Both mappings were obtained by human analysis and lead to Boolean implementations that compare quite favorably with others obtained by tweaking assignment or encoding optimization tools. However, mappings that lead to better implementations may be feasible. Olivier Muller, Adrien Prost-Boucle, Alban Bourge, Frédéric Pétrot |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | Dynamic Coherent Cluster: A Scalable Sharing Set Management ApproachabstractThe most widely used programming models expect hardware to guarantee coherent shared memory accesses. However, with the increasing number of integrated cores on chip, resource and performance efficient scalable cache coherence protocols are needed. To address the scalability issues due to the size of the sharing set, we propose to encode, on a fixed size bit-vector, a rectangular cluster whose goal is to cover most of the sharers. The cluster size is fixed but its height, width and position are determined for each cache block and can change during execution. We use a fixed size linked list for the first few outliers, and resort to broadcast when the list overflows. We compare our solution to snoop, directory-based full bit-vector, and Ackwise. It leads to similar mean latency and 10% less traffic than Ackwise, and only a few percent more than the complete sharing set on these metrics. More importantly, it generates ten times less broadcasts than Ackwise while using similar hardware resources for a 64 cores architecture. Julie Dumas, Eric Guthmuller, Frédéric Pétrot |
ASAP | 3 |
| 2018 | A mapping approach between IR and binary CFGs dealing with aggressive compiler optimizations for performance estimationabstractIn this work, we define a mapping approach between the compiler intermediate representation and the binary control flow graph for the purpose of performance estimation in native simulation. Our approach handles aggressive compiler optimizations such as loop unrolling without having to introduce any modification to the compiler. Our mapping approach experimentally leads to a good accuracy (0.59% error) while keeping a 25x speedup for native simulation compared to instruction set simulation. Omayma Matoussi, Frédéric Pétrot |
ASP-DAC | 2 |
| 2018 | Overview of the state of the art in embedded machine learningabstractNowadays, the main challenges in embedded machine learning are related to artificial neural networks. Inspired by the biological neural networks, artificial neural networks are able to solve complex problems, by performing a tremendous amount of relatively simple parallel computations. Embedding such networks in autonomous devices raises the issues of energy efficiency, resource usage and accuracy. The aim of this paper is to provide a comprehensive analysis of the efforts made in recent years to implement artificial neural network architectures suitable for embedded applications. Liliana Andrade, Adrien Prost-Boucle, Frédéric Pétrot |
DATE | 3 |
| 2018 | Message-Oriented Devices on FPGAsabstractEmbedded systems increasingly include an FPGA for performance or power efficiency. Fortunately, FPGA makers provide efficient tools to develop and assemble multiple Intellectual Properties (IPs) as devices on FPGAs. Unfortunately, the integration challenge does not stop there, hardware devices are only usable if they have available software device drivers executing on the processor subsystem. To better approach this end-to-end integration challenge, across both hardware and software, we argue that both sides need to evolve. We propose to take a step towards message-based interfaces for hardware devices integrated on an FPGA. Our goal is to deliver to the FPGA market the plug-and-play value of the USB stack with essentially no performance overhead, negligible power increase, and a reasonable surface cost. We have implemented our proposal on the Xilinx Zynq SoC, combining ARM cores and an FPGA, demonstrating the feasibility of the approach. Thomas Baumela, Olivier Gruber, Olivier Muller, Frédéric Pétrot |
RSP | 4 |
| 2018 | High-Efficiency Convolutional Ternary Neural Networks with Custom Adder Trees and Weight CompressionabstractAlthough performing inference with artificial neural networks (ANN) was until quite recently considered as essentially compute intensive, the emergence of deep neural networks coupled with the evolution of the integration technology transformed inference into a memory bound problem. This ascertainment being established, many works have lately focused on minimizing memory accesses, either by enforcing and exploiting sparsity on weights or by using few bits for representing activations and weights, to be able to use ANNs inference in embedded devices. In this work, we detail an architecture dedicated to inference using ternary {−1, 0, 1} weights and activations. This architecture is configurable at design time to provide throughput vs. power trade-offs to choose from. It is also generic in the sense that it uses information drawn for the target technologies (memory geometries and cost, number of available cuts, etc.) to adapt at best to the FPGA resources. This allows to achieve up to 5.2k frames per second per Watt for classification on a VC709 board using approximately half of the resources of the FPGA. Adrien Prost-Boucle, Alban Bourge, Frédéric Pétrot |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2017 | Loop aware IR-level annotation framework for performance estimation in native simulationabstractNative simulation is an interesting virtual prototyping candidate to speed-up architecture exploration and early software developments. It however does not provide out-of-the box non-functional information needed for software performance estimation. Annotating software with information is complex as highlevel codes and binary codes have different structures due to compiler optimizations. This work proposes an annotation framework at compiler IR-level that focuses on loop structures, and reflects optimizations through a mapping scheme between the binary and the high-level IR. Experiments on instruction count show in average around 2% of error. Omayma Matoussi, Frédéric Pétrot |
ASP-DAC | 2 |
| 2017 | Modeling instruction cache and instruction buffer for performance estimation of VLIW architectures using native simulationabstractIn this work, we propose an icache performance estimation approach that focuses on a component necessary to handle the instruction parallelism in a very long instruction word (VLIW) processor: the instruction buffer (IB). Our annotation approach is founded on an intermediate level native-simulation framework. It is evaluated with reference to a cycle accurate instruction set simulator leading to an average cycle count error of 9.3% and an average speedup of 10. Omayma Matoussi, Frédéric Pétrot |
DATE | 2 |
| 2017 | A Distributed NUCA Architecture Using an Efficient NoC Multicasting SupportabstractExploiting at best every bit of memory on chip is a must for finding the best trade-off between cost and performance when designing Many-Core MPSoCs. In this paper, we propose a new memory hierarchy organization that maximizes the usage of the available memory at each cache level while avoiding data redundancy. We also aim to reduce data access time by avoiding data migration. Our scheme is based on Non-Uniform Cache Architectures (NUCA), and makes use of a novel NoC multicast messaging support. It requires to logically partition the network into imagined layers in which each layer is bound to one of the cache levels. We assess the efficiency of our proposal by evaluating performance metrics, using the Gem5 simulator and the Splash2 benchmark suite. Our results show an improvement of 24% in L2 miss rate over baseline S_NUCA organization and more than 25% in average network latency. We report a maximum improvement of 40% in execution time. Also, we show that our approach outperforms R_NUCA. Hela Belhadj Amor, Abbas Sheibanyrad, Frédéric Pétrot |
DSD | 3 |
| 2017 | Optimizing Memory Access Performance Using Hardware Assisted Virtualization in Retargetable Dynamic Binary TranslationabstractDynamic Binary Translation is one of the most efficient strategies for the simulation of System-on-Chips, with recent studies showing that a large part of the simulation time is spent in realizing memory accesses. Indeed, the simulation of each load and store instructions requires a software emulation of the hardware Memory Management Unit (MMU). In this work, we propose to realize memory accesses in hardware, taking advantage of the hardware-assisted virtualization capabilities that are now available in modern processors. To do so, we have to setup and maintain shadow page tables, like any regular hypervisor would do, running the entire simulator on a virtual CPU. Now, each load and store instructions can be translated to just a couple of load and store instruction, executing at regular speed, therefore avoiding entirely the overhead of the software emulation of hardware MMU. The goal of this paper is to explain how it can be done. To demonstrate our idea, we have implemented our approach in the QEMU retargetable DBT engine, speeding up the simulation by as much as 40%. Antoine Faravelon, Olivier Gruber, Frédéric Pétrot |
DSD | 3 |
| 2017 | Scalable high-performance architecture for convolutional ternary neural networks on FPGAabstractThanks to their excellent performances on typical artificial intelligence problems, deep neural networks have drawn a lot of interest lately. However, this comes at the cost of large computational needs and high power consumption. Benefiting from high precision at acceptable hardware cost on these difficult problems is a challenge. To address it, we advocate the use of ternary neural networks (TNN) that, when properly trained, can reach results close to the state of the art using floatingpoint arithmetic. We present a highly versatile FPGA friendly architecture for TNN in which we can vary both the number of bits of the input data and the level of parallelism at synthesis time, allowing to trade throughput for hardware resources and power consumption. To demonstrate the efficiency of our proposal, we implement high-complexity convolutional neural networks on the Xilinx Virtex-7 VC709 FPGA board. While reaching a better accuracy than comparable designs, we can target either high throughput or low power. We measure a throughput up to 27 000 fps at ≈7W or up to 8.36 TMAC/s at ≈13 W. Adrien Prost-Boucle, Alban Bourge, Frédéric Pétrot, Hande Alemdar, Nicholas H. M. Caldwell, Vincent Leroy 0001 |
FPL | 3 |
| 2017 | Ternary neural networks for resource-efficient AI applicationsabstractThe computation and storage requirements for Deep Neural Networks (DNNs) are usually high. This issue limits their deployability on ubiquitous computing devices such as smart phones, wearables and autonomous drones. In this paper, we propose ternary neural networks (TNNs) in order to make deep learning more resource-efficient. We train these TNNs using a teacher-student approach based on a novel, layer-wise greedy methodology. Thanks to our two-stage training procedure, the teacher network is still able to use state-of-the-art methods such as dropout and batch normalization to increase accuracy and reduce training time. Using only ternary weights and activations, the student ternary network learns to mimic the behavior of its teacher network without using any multiplication. Unlike its {-1,1} binary counterparts, a ternary neural network inherently prunes the smaller weights by setting them to zero during training. This makes them sparser and thus more energy-efficient. We design a purpose-built hardware architecture for TNNs and implement it on FPGA and ASIC. We evaluate TNNs on several benchmark datasets and demonstrate up to 3.1 × better energy efficiency with respect to the state of the art while also improving accuracy. Hande Alemdar, Vincent Leroy 0001, Adrien Prost-Boucle, Frédéric Pétrot |
IJCNN | 4 |
| 2017 | Dynamic Binary Translation of VLIW Codes on Scalar ArchitecturesabstractMany of the recently announced integrated manycore architectures targeting specific applications embed several, if not many, very long instruction word (VLIW) processors. To start developing software while the hardware is still being designed, virtual prototypes of the full system are commonly used. Fast processor simulation is thus a requirement. To that aim, this paper introduces a strategy to perform dynamic binary translation (DBT) of VLIW codes on scalar architectures. We propose a high level simulation algorithm which takes into account VLIW oddities, such as explicit instruction parallelism, instructions with non unit register update latency, and delayed slots in branches. We present the implementation details of this algorithm within a DBT environment, as it raises many corner cases that are irrelevant in scalar DBT. Our experiments confirm that our solution is functionally correct, and show speedups of 1 and 2 orders of magnitude compared to raw instruction interpretation, even though no optimizations were performed on the code during and after translation. Luc Michel, Frédéric Pétrot |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | Detecting Software Cache Coherence Violations in MPSoC Using Traces Captured on Virtual PlatformsabstractSoftware cache coherence schemes tend to be the solution of choice in dedicated multi/many core systems on chip, as they make the hardware much simpler and predictable. However, despite the developers’ effort, it is hard to make sure that all preventive measurements are taken to ensure coherence. In this work, we propose a method to identify the potential cache coherence violations using traces obtained from virtual platforms. These traces contain causality relations among events, which allow first to simplify the analysis, and second to avoid relying on timestamps. Our method identifies potential violations that may occur during a given execution for write-through and write-back cache policies. Therefore, it is independent of the software coherence protocol. We conducted experiments on parallel applications running on a lightweight SMP operating system, and we were able to detect coherence issues that we could then solve. Marcos Aurélio Pinto Cunha, Omayma Matoussi, Frédéric Pétrot |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2017 | Efficient and Versatile FPGA Acceleration of Support Counting for Stream Mining of Sequences and Frequent ItemsetsabstractStream processing has become extremely popular for analyzing huge volumes of data for a variety of applications, including IoT, social networks, retail, and software logs analysis. Streams of data are produced continuously and are mined to extract patterns characterizing the data. A class of data mining algorithm, called generate-and-test , produces a set of candidate patterns that are then evaluated over data. The main challenges of these algorithms are to achieve high throughput, low latency, and reduced power consumption. In this article, we present a novel power-efficient, fast, and versatile hardware architecture whose objective is to monitor a set of target patterns to maintain their frequency over a stream of data. This accelerator can be used to accelerate data-mining algorithms, including itemsets and sequences mining. The massive fine-grain reconfiguration capability of field-programmable gate array (FPGA) technologies is ideal to implement the high number of pattern-detection units needed for these intensive data-mining applications. We have thus designed and implemented an IP that features high-density FPGA occupation and high working frequency. We provide detailed description of the IP internal micro-architecture and its actual implementation and optimization for the targeted FPGA resources. We validate our architecture by developing a co-designed implementation of the Apriori Frequent Itemset Mining (FIM) algorithm, and perform numerous experiments against existing hardware and software solutions. We demonstrate that FIM hardware acceleration is particularly efficient for large and low-density datasets (i.e., long-tailed datasets). Our IP reaches a data throughput of 250 million items/s and monitors up to 11.6k patterns simultaneously, on a prototyping board that overall consumes 24W in the worst case. Furthermore, our hardware accelerator remains generic and can be integrated to other generate and test algorithms. Adrien Prost-Boucle, Frédéric Pétrot, Vincent Leroy 0001, Hande Alemdar |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2016 | Simulation driven insertion of data prefetching instructions for early software-on-SoC optimizationabstractA System-On-Chip (SoC) is a combination of hardware and software which interact to perform a set of functions, usually for some specific application domain. However, with the increasing complexity of both hardware and software, estimating and optimizing the performance of software on SoC during the design process is becoming a difficult objective. The aim of this paper is to show that observability provided by cycle accurate (CA) simulations using virtual platforms available early in the design flow is helping to obtain estimations and to design a software strategy which optimizes software-on-SoC performance, based on some particular metrics. We apply this approach to a case study which tackles the memory wall problem and contributes to its resolution by using software data prefetching technique. We exploit profiling capabilities provided by simulation models, by analyzing the collected data to identify high memory accesses latencies and finally inserting data prefetching instructions in a suitable and non-intuitive way for increasing system execution performance. Bubble sort is used as experimental methodology and the Inverse Discrete Cosine Transform (JPEG IDCT) routine typical from an industrial decoder is used as initial case study. Based on platform simulation observability, we reach an improvement of more than 25% on the overall execution time. Perrin N. Ntafam, Eric Paire, Alain Clouard, Frédéric Pétrot |
RSP | 4 |
| 2015 | Fast and accurate branch predictor simulation
Antoine Faravelon, Nicolas Fournel, Frédéric Pétrot |
DATE | 3 |
| 2014 | Scalability bottlenecks discovery in MPSoC platforms using data mining on simulation tracesabstractNowadays, a challenge faced by many developers is the profiling of parallel applications so that they can scale over more and more cores. This is especially critical for embedded systems powered by Multi-Processor System-on-Chip (MPSoC), where ever demanding applications have to run smoothly on numerous cores, each with modest power budget. The reasons for the lack of scalability of parallel applications are numerous, and it can be time consuming for a developer to pinpoint the correct one. In this paper, we propose a fully automatic method which detects the instructions of the code which lead to a lack of scalability. The method is based on data mining techniques exploiting low level execution traces produced by MPSoC simulators. Our experiments show the accuracy of the proposed technique on five different kinds of applications, and how the information reported can be exploited by application developers. Sofiane Lagraa, Alexandre Termier, Frédéric Pétrot |
DATE | 3 |
| 2014 | Device driver generation targeting multiple operating systems using a model-driven methodologyabstractWe present a new device driver generation approach capable of automatically generating a large portion of device drivers code, and this for different operating systems (OSes). This approach is based on a model-driven methodology, where a tiny language is utilized to model the device features and abstract low-level complexities of a driver. The approach can handle different driver architectures. We demonstrate the genericity of the approach by applying it to a fairly mature device class that has standardized interfaces, and also to a brand-new device that has significant functionality differences. The code was generated for two OSes, one targeting the embedded space and the other a full featured one. Guillaume Godet-Bar, Frédéric Rousseau 0001, Frédéric Pétrot |
RSP | 4 |
| 2014 | Assignment of Vertical-Links to Routers in Vertically-Partially-Connected 3-D-NoCsabstractThis paper addresses elevator assignment in vertically-partially-connected 3-D-networks-on-chip (NoCs). Elevators are vertical links between dies. Because of yield issues, Through-Silicon-Via (TSV) cost, and heterogeneity in dimension, topology, and technology of different dies, vertically-partially-connected topologies seem unavoidable in the emerging 3-D-NoCs as opposed to fully-connected topologies. In such partially-connected topologies, as there are fewer elevators than routers, the assignment of elevators to routers becomes a new 3-D-specific optimization problem. An improper assignment can lead to dramatic network performance degradation. This paper proposes an elevator assignment method for best-effort wormhole 3-D-NoCs to improve the average network performance. Experimental results show an improvement of about 90% even compared to a greedy (intrinsically good) initial assignment. Sahar Foroutan, Abbas Sheibanyrad, Frédéric Pétrot |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2013 | Native simulation of complex VLIW instruction sets using static binary translation and Hardware-Assisted VirtualizationabstractWe introduce a static binary translation flow in native simulation context for cross-compiled VLIW executables. This approach is interesting in situations where either the source code is not available or the target platform is not supported by any retargetable compilation framework, which is usually the case for VLIW processors. The generated simulators execute on a Hardware-Assisted Virtualization (HAV) based native platform. We have implemented this approach for a TI C6x series processor and our simulation results show a speed-up of around two orders of magnitude compared to the cycle accurate simulators. Mian Muhammad Hamayun, Frédéric Pétrot, Nicolas Fournel |
ASP-DAC | 2 |
| 2013 | Data mining MPSoC simulation traces to identify concurrent memory access patternsabstractDue to a growing need for flexibility, massively parallel Multiprocessor SoC (MPSoC) architectures are currently being developed. This leads to the need for parallel software, but poses the problem of the efficient deployment of the software on these architectures. To address this problem, the execution of the parallel program with software traces enabled on the platform and the visualization of these traces to detect irregular timing behavior is the rule. This is error prone as it relies on software logs and human analysis, and requires an existing platform. To overcome these issues and automate the process, we propose the conjoint use of a virtual platform logging at hardware level the memory accesses and of a data-mining approach to automatically report unexpected instructions timings, and the context of occurrence of these instructions. We demonstrate the approach on a multiprocessor platform running a video decoding application. Sofiane Lagraa, Alexandre Termier, Frédéric Pétrot |
DATE | 3 |
| 2013 | A system-level overview and comparison of three High-Speed Serial Links: USB 3.0, PCI Express 2.0 and LLI 1.0abstractHigh Speed Serial Links (HSSL) are found in almost all today's System-on-Chip (SoC) connecting different components: the main chip and its I/Os, chip to chip (the main chip to a companion chip, memory sharing between two chips), etc... A variety of standards exist, each of which is used for a specific application, and many parameters affect their performance. In this paper we make a comparison of three high-speed protocols, the USB 3.0, PCI Express 2.0 (PCIe) and LLI 1.0. We analyze their different parameters, mainly the data exchange protocol, errors management, the Bit Error Rate (BER), data efficiency and the quality of service (QoS) for each of the protocols. We also show the relation between these parameters and how improving one parameter could result in a degradation of another, and based on this analysis, we finish by concluding the reason why USB is used for I/Os, PCIe is used for data hungry devices and LLI for memory sharing. Julien Saade, Frédéric Pétrot, Andre Picco, Joel Huloux, Abdelaziz Goulahsen |
DDECS | 2 |
| 2013 | A General Framework for Average-Case Performance Analysis of Shared ResourcesabstractContemporary embedded systems are based on complex heterogeneous multi-core platforms to cater to the increasing number of applications, some of which have (soft) real-time requirements. To reduce cost, resources are shared using diverse arbitration mechanisms, such as Time-Division Multiplexing (TDM), Static-Priority (SP), and Round-Robin (RR), depending on application and resource requirements. However, resource sharing results in interference between sharing applications making it difficult to estimate if the average latency is sufficient to satisfy their real-time requirements. Existing work proposes isolated models that either fail to address the diversity of arbitration mechanisms or cannot capture the dynamic arrival and service processes of applications and resources in multi-core platforms. This paper addresses this problem by proposing a general framework for average-case performance analysis of shared resources in multi-core platforms. The two main contributions are: 1) a general model for resource sharing based on queuing theory that can be used with different arbiters and that captures architectural features of the shared resource, such as pipelining and arbitration delay, and 2) three arbiter models for TDM, SP, and RR, respectively that assume general distributions (G/G/1) and fits within the framework. Sahar Foroutan, Benny Akesson, Kees Goossens, Frédéric Pétrot |
DSD | 4 |
| 2013 | Automated generation of efficient instruction decoders for instruction set simulatorsabstractFast Instruction Set Simulators (ISS) are a critical part of MPSoC design flows. The complexity of developing these ISS combined with the ability to extend instruction sets tend to make automated generation of ISS a need. One important part of every ISS is its instruction decoder, but as the encoding of instruction sets becomes less orthogonal because of the incremental addition of instructions, the generation of a decoder is not anymore an obvious task. In this paper, we present two automated decoder generation strategies that are able to handle non-orthogonal instruction encodings. The first one builds a decision tree that does not consider the instruction's occurrences while the second considers these frequencies. In both cases, we use binary decision diagrams to represent the instructions encodings and the complex conditions due to the non-orthogonality of the encodings in order to generate the decoders. Our experiments on the MIPS and ARM (including VFP and Neon extensions) instruction sets show that both algorithms produce efficient decoders, and that it is beneficial to consider instruction frequencies. Nicolas Fournel, Luc Michel, Frédéric Pétrot |
ICCAD | 3 |
| 2013 | Elevator-First: A Deadlock-Free Distributed Routing Algorithm for Vertically Partially Connected 3D-NoCsabstractIn this paper, we propose a distributed routing algorithm for vertically partially connected regular 2D topologies of different shapes and sizes (e.g., 2D mesh, torus, ring). The topologies that are the target of this algorithm are of practical interest in the 3D integration of heterogeneous dies using Through-Silicon-Vias (TSVs). Indeed, TSV-based 3D integration allows to envision the stacking of dies with different functions and technologies, using as an interconnect backbone a 3D-NoC. Intrinsically, 3D topologies have better performances, but yield and active area (and thus the cost) are function of the number of TSVs; therefore, the designs tend to use only a subset of available TSVs between two dies. The definition of blockage free and low implementation cost distributed deterministic routing on this kind of topology is thus of theoretical and practical interests. We formally prove that independently of the shape and dimensions of the planar topologies and of the number and placement of the TSVs, the proposed routing algorithm using two virtual channels in the plane is deadlock and livelock free. We also experimentally show that the performance of this algorithm is still acceptable when the number of vertical connections decreases. Florentine Dubois, Abbas Sheibanyrad, Frédéric Pétrot, Maryam Bahmani |
IEEE Trans. Computers | 3 |
| 2013 | An Iterative Computational Technique for Performance Evaluation of Networks-on-ChipabstractThe trend toward integrated many-core architectures makes the network-on-chip (NoC) technology, the on-chip communication infrastructure of choice. However, and as opposed to a simple bus, due to its distributed and complex nature in terms of topology, wire size, routing algorithm, and so on, the timing behavior and thus performance of the infrastructure is difficult to predict. Therefore, one of the important phases in the NoC design flow is performance evaluation, which is to extract performance metrics to verify whether a specific instance from the NoC design space satisfies the requirements of the entire system. In this sense, reducing the time to obtain the NoC performance and consequently speeding-up the design space exploration is one of the keys that can considerably reduce the design-flow time and cost. In an effort toward this direction, we propose in this paper a novel analytical performance evaluation method that can be used in the earliest stages of the design flow, before using time-consuming simulations. The analytical method is used to evaluate the performance of a general purpose NoC and we show that it can predict the router latency, end-to-end per-flow latency, and network saturation point with an accuracy comparable to a cycle-accurate simulation. To systematically analyze the accuracy of our method compared to the corresponding simulation model, we present also an innovative accuracy analysis method. Sahar Foroutan, Yvain Thonnart, Frédéric Pétrot |
IEEE Trans. Computers | 3 |
| 2012 | Cost-efficient buffer sizing in shared-memory 3D-MPSoCs using wide I/O interfacesabstractThis paper addresses link-buffer capacity allocation in the design process of best-effort 3DNoCs holding hotspot memory ports. We show that in 3DSoCs with integrated wide I/O DRAMs, the congestion spreading is different from SoCs with external DRAMs: the bottlenecks are not anymore the external memory ports but the network links that become saturated and retropropagate the congestion. The distribution of bottleneck links is directly affected by the traffic directed to the hot memory ports. Using an analytical performance evaluation method, we determine network link buffer capacities according to the given workload composed of regular and hotspot traffics. Sahar Foroutan, Abbas Sheibanyrad, Frédéric Pétrot |
DAC | 3 |
| 2012 | Multi-device Driver Synthesis Flow for Heterogeneous Hierarchical SystemsabstractHeterogeneous hierarchical architectures result from the interconnection of several heterogeneous MPSoCs through an efficient communication infrastructure. Each communication between two processing units requires the use of hardware devices managed by a software driver responsible to initialize, handle and complete the communication. This paper describes a multi-device driver synthesis flow based on available communication paths of the architecture. A small set of generic driver templates is available in a driver library. The flow is in charge to select one correct driver template from the library, and then to configure and specialize it in order to produce the source code. The effectiveness of our approach is illustrated by a significant example. Alexandre Chagoya-Garzon, Frédéric Rousseau 0001, Frédéric Pétrot |
DSD | 3 |
| 2012 | HCM: An abstraction layer for seamless programming of DPR FPGAabstractWell-known for its efficient computing capabilities, FPGA-based architectures also have the potential for high flexibility with dynamic reconfiguration features. Yet, writing applications on these architectures is laborious, poorly portable and hardly scalable to multi-user and/or multi-FPGA systems, mainly because of a mixture of application related code and flexibility management code. In this paper, we propose a new abstraction layer, called Hardware Component Manager (HCM), which clearly separates the allocation of a hardware function from the control of a reconfiguration procedure, and guarantees the security of coexisting configurations. The implementation of this HCM layer on realistic simulation platforms demonstrates its ability to ease the management of FPGA flexibility while preserving performance and ensuring hardware function protection. HCM implementation and its simulation environment are open-source in the hope of reuse by the community. Olivier Muller, Pierre-Henri Horrein, Frédéric Pétrot |
FPL | 4 |
| 2012 | Accurate on-chip router area modeling with Kriging methodologyabstractNetworks-on-chips (NoCs) have emerged as an effective interconnection solution for modern MPSoCs. However, NoCs are characterized by a wide range of parameters and early performance estimations have become keys. We propose an approach to build static cost models (e.g. area) of NoC components. The modeling relies on Kriging theory, which catches the complex interactions between parameters on the basis of few low-level results. Experimental results show that the produced model has a good level of accuracy and a predictable behavior. Florentine Dubois, Valerio Catalano, Marcello Coppola, Frédéric Pétrot |
ICCAD | 4 |
| 2012 | Automatic congestion detection in MPSoC programs using data mining on simulation tracesabstractThe efficient deployment of parallel software, specifically legacy one, on Multiprocessor systems on chip (MPSoC) is a challenging task. In this paper, we introduce the use of a data-mining approach on traces of a functionally correct program to automatically identify recurring congestion points and their sources. Each memory transaction, i.e. instruction fetch, data load and data store, occurring in the system is logged, thanks to the use of a virtual platform of the system. The resulting trace is analyzed to discover memory access patterns that are occurring frequently and that feature high latencies. These patterns are sorted by order of decreasing occurrence and estimated congestion level, allowing the easy identification of the sources of inefficiency. We have simulated a MPSoC with 16 processors running multiple applications, and have been able to automatically detect congestion on resources and their sources in the parallel program using this technique by analyzing gigabytes of traces. Sofiane Lagraa, Alexandre Termier, Frédéric Pétrot |
RSP | 3 |
| 2012 | Native Simulation of MPSoC Using Hardware-Assisted VirtualizationabstractIntegration of multiple heterogeneous processors into a single system-on-a-chip is a clear trend in embedded devices. Designing and verifying these devices requires high-speed and easy-to-build simulation platforms. Among the software simulation approaches, native simulation is a good candidate since the embedded software is executed natively on the host machine, and no instruction set simulator development effort is necessary. However, existing native simulation approaches are such that the simulated software shares the memory space of the modeled hardware modules and the host operating system, making impractical the support of legacy code running on the target platform. To overcome this issue seldom mentioned in the literature, we propose the addition of a transparent address space translation layer to separate the target address space from the host simulator one. For this, we exploit the hardware-assisted virtualization technology now available on most general-purpose processors. Experiments show that this solution does not degrade the native simulation speed, while keeping the ability to accomplish software performance evaluation. Hao Shen 0009, Mian Muhammad Hamayun, Frédéric Pétrot |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2011 | Handling dynamic frequency changes in statically scheduled cycle-accurate simulationabstractAlthough high level simulation models are being increasingly used for digital electronic system validation, cycle accuracy is still required in some cases, such as hardware protocol validation or accurate power/energy estimation. Cycle-accurate simulation is however slow and acceleration approaches make the assumption of a single constant clock, which is not true anymore with the generalization of dynamic voltage and frequency scaling techniques. Fast cycle-accurate simulators supporting several clocks whose frequencies can change at run time are thus needed. This paper presents two algorithms we designed for this purpose and details their properties and implementations. Marius Gligor, Frédéric Pétrot |
ASP-DAC | 2 |
| 2011 | Speeding-up SIMD instructions dynamic binary translation in embedded processor simulationabstractThis paper presents a strategy to speed-up the simulation of processors having SIMD extensions using dynamic binary translation. The idea is simple: benefit from the SIMD instructions of the host processor that is running the simulation. The realization is unfortunately not easy, as the nature of all but the simplest SIMD instructions is very different from a manufacturer to an other. To solve this issue, we propose an approach based on a simple 3-addresses intermediate SIMD instruction set on which and from which mapping most existing instructions at translation time is easy. To still support complex instructions, we use a form of threaded code. We detail our generic solution and demonstrate its applicability and effectiveness using a parametrized synthetic benchmark making use of the ARMv7 NEON extensions executed on a Pentium with MMX/SSE extensions. Luc Michel, Nicolas Fournel, Frédéric Pétrot |
DATE | 3 |
| 2011 | An Environment for (re)configuration and Execution Managenment of Flexible Radio PlatformsabstractThis paper presents the Flexible Radio Kernel (FRK), a configuration and execution management environment for hybrid hardware/software flexible radio platform. The aim of FRK is to manage platform reconfiguration for multi-mode, multi-standard operation, with different levels of abstraction. A high level framework is described, to manage multiple MAC layers, and to enable MAC cooperation algorithms for cognitive radio. A low-level environment is also available to manage platform reconfiguration for radio operations. Radio can be implemented using hardware or software elements. Configuration state is hidden to the high-level layers, offering pseudo concurrency (time sharing) properties. This study presents a global view of FRK, with details on some specific parts of the environment. A practical study with algorithmic description is presented. Pierre-Henri Horrein, Christine Hennebert, Frédéric Pétrot |
DSD | 3 |
| 2011 | Spidergon STNoC design flowabstractIn this demonstration we present an enhanced version of the usual Spidergon STNoC design flow. In addition, we show the automatic generation of a simulation platform that can be used to perform early architecture exploration. Florentine Dubois, José Cano 0001, Marcello Coppola, José Flich, Frédéric Pétrot |
NOCS | 5 |
| 2010 | A flexible hybrid simulation platform targeting multiple configurable processors SoCabstractMultiple Configurable Processors System-on-Chip (MCPSoC) platforms have both performance and power advantages for embedded applications. Unfortunately, at early design stages, because of the processor configuration, I/O device changes and MCPSoC architecture modifications, designers waste much time on the Operating System (OS) porting work with general Instruction Set Simulator (ISS) based SoC simulation platforms. In this paper, we propose a hybrid simulation platform which uses general ISS and implements the Hardware Abstraction Layer (HAL) Application Programming Interfaces (APIs) and I/O device driver APIs with the SystemC modules on host machines directly. This hybrid simulation platform can shorten the application validation process by avoiding assembly code and hard-coded address modifications of traditional OS porting work. We show the advantages of our new hybrid simulation platform with a video decoding case study in the end. Hao Shen 0009, Frédéric Pétrot |
ASP-DAC | 2 |
| 2010 | Lightweight Transactional Memory systems for NoCs based architectures: Design, implementation and comparison of two policies
Quentin L. Meunier, Frédéric Pétrot |
J. Parallel Distributed Comput. | 2 |
| 2010 | Hardware/software support for adaptive work-stealing in on-chip multiprocessor
Quentin L. Meunier, Frédéric Pétrot, Jean-Louis Roch |
J. Syst. Archit. | 2 |
| 2009 | A System Framework for the Design of Embedded Software Targeting Heterogeneous Multi-core SoCsabstractEmbedded appliances designers rely on heterogeneous multi-core system-on-chips (HMC-SoC) to provide the computing power required by modern applications. Due to the inherent complexity of this kind of platform, the development of specific system architectures is not considered as an option to provide low-level services to an application. Hence, the software is built either from scratch - when the softwarepsilas requirements are not too high - or over a general-purpose operating system, leading to performance and memory usage trade-offs. Our contribution is a component-based system framework that provides high-level system services for embedded software applications with few impacts on the memory usage and final performances, thanks to strong interfaces that enable the reuse of existing software elements and facilitate the support of multiple hardware platforms. The efficiency of our approach is demonstrated on an existing MC-SoC. Xavier Guerin, Frédéric Pétrot |
ASAP | 2 |
| 2009 | Automatic instrumentation of embedded software for high level hardware/software co-simulationabstractWe propose an automatic instrumentation method for embedded software annotation to enable performance modeling in high level hardware/software co-simulation environments. The proposed ldquocross-annotationrdquo technique consists of extending a retargetable compiler infrastructure to allow the automatic instrumentation of embedded software at the basic block level. Thus, target and annotated native binaries are guaranteed to have isomorphic control flow graphs (CFG). The proposed method takes into account the processor-specific optimizations at the compiler level and proves to be accurate with low simulation overhead. Aimen Bouchhima, Patrice Gerin, Frédéric Pétrot |
ASP-DAC | 3 |
| 2009 | Novel task migration framework on configurable heterogeneous MPSoC platformsabstractHeterogeneous MPSoC architectures can provide higher performance and flexibility with less power consumption and lower cost than homogeneous ones. However, as processor instruction sets of general heterogeneous MPSoCs are not identical, tasks migration between two heterogeneous processors is not possible. To enable this function, we propose to build one specific heterogeneous MPSoC platform in which all heterogeneous processors are based on the same core instruction set for the operating system realization. Different extended instructions can be added for different processors to improve the system performance. Tasks can be migrated from one processor to another only if the target processor has all instructions which can meet the execution requirement of this task. This paper concentrates on the infrastructure that is necessary to support the scheduling and migration of tasks between the processors. By using the motion-JPEG case study, we confirm that our task migration framework can achieve higher processor usage rate and more flexibility. Hao Shen 0009, Frédéric Pétrot |
ASP-DAC | 2 |
| 2009 | Extending IP-XACT to support an MDE based approach for SoC designabstractWe are interested in the problem of improving ipreuse in SoC design. This paper presents an MDE based approach based on a proposed IP-XACT standard extension. This approach combines the benefits of using MDE techniques in SoC design such as abstraction levels definition and model transformation for code generation, and the benefits of the IP-XACT standard such as a unique exchange format of packaged IPs (Intellectual Property) with reuse capabilities. Amin El Mrabti, Frédéric Pétrot, Aimen Bouchhima |
DATE | 2 |
| 2009 | Adaptive Dynamic Voltage and Frequency Scaling Algorithm for Symmetric Multiprocessor ArchitectureabstractSymmetric multiprocessor architectures (SMP) are gaining popularity for system on chip (SoC) applications as they provide an interesting power/performance/flexibility tradeoff. We address the problem of power consumption in such architectures by designing a new adaptive Dynamic Voltage Frequency Scaling (DVFS) algorithm for non real time operating system (non-RTOS) running on a SMP-SoC. We show the effectiveness of our algorithm on a cycle accurate simulation environment running a real-life multi-threaded application. The experimental results show a potential gain of up to 55%. Marius Gligor, Nicolas Fournel, Frédéric Pétrot |
DSD | 3 |
| 2009 | A MPSoC Prototyping Platform for Flexible Radio ApplicationsabstractFull-fledged software radio platforms are complex and expensive systems, focused on signal processing, and not very suitable for easy development and large scale experimentation. We propose a Multi-Processor System-on-Chip (MPSoC) prototyping platform targeting the support for flexible radio. This platform is fully customizable at every layer of the wireless networking stack, making it easy to prototype new protocols from the radio to the application layers. Our goal was threefold: design an efficient but cheap platform supporting flexible radio, provide support for a full system on the platform so that it can run autonomously, use "standard" components as much as possible and a modular design to ensure fast and simple development and testing to network developers. We rely on a highly modular Field-Programmable Gate Array (FPGA) based architecture. The practical results achieved so far show the effectiveness of the proposed solution in term of flexibility and cost. Damien Hedde, Pierre-Henri Horrein, Frédéric Pétrot, Robin Rolland, Franck Rousseau |
DSD | 3 |
| 2009 | Abstract Description of System Application and Hardware Architecture for Hardware/Software Code GenerationabstractThe deployment of a system application over a hardware architecture is a costly phase in the design process. This cost increases when dealing with complex applications in terms of computation requirements and exchange of data and for advanced architectures with complex and configurable communication infrastructures. The usage of abstract models for application, architecture and mapping is a key element for automatic hardware/software code generation and for the final deployment. In this paper, we present languages for abstract modeling of application, architecture, meta-mapping and mapping and we introduce a code generation flow. The use of those models allows the extraction and exploitation of architectural and application information for specific code generation to a target platform. A case study of modeling and deploying a complex 4G telecommunication application on a heterogeneous and multi core platform is presented. Amin El Mrabti, Hamed Sheibanyrad, Frédéric Rousseau 0001, Frédéric Pétrot, Romain Lemaire, Jérôme Martin |
DSD | 4 |
| 2008 | Efficient Implementation of Native Software Simulation for MPSoCabstractEfficient and precise simulation models at a high abstraction level are required in order to perform early design validations and architecture explorations of multi- processor system-on-chip (MPSoC) platforms. Although native software simulation approaches provide interesting capabilities, they quickly become unsuitable when complex hardware architecture have to be considered. In this paper, we present a SystemC-based MPSoC platform implementation that allows native software simulation while keeping details of the underlying hardware model. The key contribution of this work is a realistic memory mapping modelling that makes possible the simulation of operating systems and software applications on complex hardware models with multiple processors and DMA devices. This method also allows the reuse of different software components for the target processor(s). Experimental results show the efficiency of the proposed method to validate software on complex hardware architectures. Patrice Gerin, Xavier Guerin, Frédéric Pétrot |
DATE | 3 |
| 2008 | Comparison of memory write policies for NoC based Multicore Cache Coherent SystemsabstractThe following study shows a direct comparison of memory write policies in Shared Memory Multicore Systems. Although there are much work and many studies about this issue, our work takes into account the difficulties related to on chip communication using network-like interconnects. Our study is based on cycle approximate bit accurate simulations (CABA) of platforms with up to 64 processors, modelling accurately all the aspects of multi-threaded program execution and memory accesses. Our main results show that write-through caches perform well compared to write-back ones, with a slightly simpler implementation and comparable traffic. Pierre Guironnet de Massas, Frédéric Pétrot |
DATE | 2 |
| 2008 | Large Scale On-Chip Networks : An Accurate Multi-FPGA Emulation PlatformabstractInterconnect validation is an important early step toward global SoC (system-on-chip) validation. Fast performances evaluation and design space exploration for NoCs (networks-on-chip) are therefore becoming critical issues. A significant speed up of the global validation process for NoC-centric SoCs could be achieved by prototyping such systems on reconfigurable devices (FPGA). However, as SoC complexity increases with the technology scaling, existing general purpose prototyping platforms are far from being suited for large systems. In this paper we present a study for a scalable multi-FPGA platform, designed for NoCs emulation and debugging. This platform allows the integration of complete systems as well as a near cycle-accurate performance estimation. Abdellah-Medjadji Kouadri-Mostefaoui, Benaoumeur Senouci, Frédéric Pétrot |
DSD | 3 |
| 2007 | Scalable Multi-FPGA Platform for Networks-On-Chip EmulationabstractInterconnect validation is an important early step toward global SoC (system-on-chip) validation. Fast performances evaluation and design space exploration for NoCs (networks-on-chip) are therefore becoming critical issues. A significant speedup of the global validation process for NoC-centric SoCs could be achieved by prototyping such systems on reconfigurable devices (FPGA). However, as SoC complexity increases with the technology scaling, existing general purpose prototyping platforms are far from being suited for large systems. In this paper we present a study for a scalable multi-FPGA platform, designed for NoCs emulation and debugging. This platform allows the integration of complete systems as well as a near cycle-accurate performance estimation. Abdellah-Medjadji Kouadri-Mostefaoui, Benaoumeur Senouci, Frédéric Pétrot |
ASAP | 3 |
| 2006 | Programming models and HW-SW interfaces abstraction for multi-processor SoCabstractFor the design of classic computers the Parallel programming concept is used to abstract HW/SW interfaces during high level specification of application software. The software is then adapted to an existing multiprocessor platforms using a low level software layers that implement the programming model. Unlike classic computers, the design of heterogeneous MPSoC includes also building the processors and other kind of hardware components required to execute the software. In this case, the programming model hides both hardware and software refinements. This paper deals with parallel programming models to abstract both hardware and software Interfaces in the case of heterogeneous MPSoC design. Different abstraction levels will be needed. For the long term, the use of higher level programming models will open new vistas for optimization and architecture exploration like CPU/RTOS tradeoffs. Ahmed Amine Jerraya, Aimen Bouchhima, Frédéric Pétrot |
DAC | 3 |
| 2006 | On Cache Coherency and Memory Consistency Issues in NoC Based Shared Memory Multiprocessor SoC ArchitecturesabstractThe concept of network on chip (NoC) is a recent breakthrough in the system on chip (SoC) design area. A lot of work has been done to define efficient NoC architectures and implementations. In this paper, our goal is twofold. Firstly, we want to outline that the use of a NoC based sharedmemory multiprocessor SoC challenges the application integrator because of the underlying assumptions of software, namely cache coherency and memory consistency. These problems are well known in general purpose shared memory multiprocessors. However, when designing a SoC, we benefit on the one hand from the knowledge of the applications, the much simpler usage of virtual memory, lower interconnect latencies and very high bandwidth at lost cost, but on the other hand we suffer from more tight design constraints (yield, power, predictable performances, ...). Secondly, we define simple and yet attractive solutions -in term of design time and hardware cost- to both problems in the context of application specific multiprocessor SoCs. Frédéric Pétrot, Alain Greiner, Pascal Gomez |
DSD | 1 |
| 2005 | A unified HW/SW interface model to remove discontinuities between HW and SW designabstractOne major challenge in System-on-Chip (SoC) design is the definition and design of interfaces between hardware and software. Traditional ASIC designer and software designer model HW/SW interface twice. Using two separate models introduces a discontinuity between hardware and software. This paper introduces a unified HW/SW component model to describe different parts of HW/SW interface at different abstraction levels. The benefits of using the proposed model are two fold: first, it provides a single model to present system design from abstract specification to mixed HW/SW implementation and second, it enables full system simulation at different abstraction level during refinement flow. Aimen Bouchhima, Frédéric Pétrot, Wander O. Cesário, Ahmed Amine Jerraya |
EMSOFT | 3 |
| 2005 | Platform-based design from parallel C specificationsabstractThis paper presents Disydent, a framework dedicated to system-on-a-chip (SoC) platform-based design for shared memory multiple instructions multiple data (MIMD) architectures. We define a platform-based design problem as a triplet (system, application, constraints) where the system is both an operating system (OS) and a hardware (HW) template that can be enhanced with dedicated coprocessors. Our contributions are: 1) the definition of a complete flow for platform-based design, from application to integration including all necessary intermediate steps and 2) a set of tightly bound operational tools to implement the flow. Disydent is based on four tools. The distributed process network (DPN) is a C library for describing Kahn process network (KPN)-based applications. The asynchronous serial interface mode register 0 (ASIM0) is a multiprocessor target platform running a microkernel. This platform can be enhanced with coprocessors generated by the user-guided high-level synthesis (UGH) tool. Cycle accurate system simulator (CASS) is a high-performance cycle-accurate simulator. The main steps of the design flow are KPN modeling, functional validation, design space exploration, high-level synthesis, and temporal validation. The design flow starts by modeling the application as a KPN. This initial description is done in C using the DPN library. The functional validation is performed by running the initial description directly on the host. Without modifying the initial description, the user can simulate a HW/software (SW) partitioning by indicating the number of processors and the processes that are to be migrated to HW. This simulation is done at the cycle-accurate level for the whole system, except for the migrated processes for which the user must provide estimated time models. The description of the processes that are selected for HW implementation must be translated into a subset of C and then synthesized. This new description is still compatible with the DPN library, so it can be used for functional validation. The temporal validation is done at the cycle-accurate level using the initial description for the SW processes and cycle-accurate models automatically generated from the C subset description for the HW processes. Disydent's strength relies on its formal KPN model that ensures a behavior that is independent of the overall system scheduling, its fast cycle-accurate validation that is several orders of magnitude faster than classical event-driven simulators, and its single description of a process that is used as input of DPN, CASS, and UGH. Ivan Augé, Frédéric Pétrot, François Donnet, Pascal Gomez |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2004 | Modular On-chip Multiprocessor for Routing Applications
Saifeddine Berrayana, Etienne Faure, Daniela Genius, Frédéric Pétrot |
Euro-Par | 4 |
| 2003 | Lightweight Implementation of the POSIX Threads API for an On-Chip MIPS Multiprocessor with VCI Interconnect
Frédéric Pétrot, Pascal Gomez |
DATE | 1 |
| 2000 | COSY communication IP'sabstractThe Fsprit/OMI-COSY project defines transaction-levels to set-up the exchange of IP's in separating function from architecture and body-behavior from proprietary interfaces. These transaction-levels are supported by the “COSY COMMUNICATION IPs” that are presented in this paper. They implement onto Systems-On-Chip the extended Kahn Process Network that is defined in COSY for modeling signal processing applications. We present a generic implementation and performance model of these system-level communications and we illustrate specific implementations. They set system communications across software and hardware boundaries, and achieve bus independence through the Virtual Component Interface of the VSI Alliance. Finally, we describe the COSY-VCC flow that supports communication refinement from specification, to performance estimation, to implementation. Jean-Yves Brunel, Wido Kruijtzer, H. J. H. N. Kenter, Frédéric Pétrot, L. Pasquier, Erwin A. de Kock, W. J. M. Smits |
DAC | 4 |
| 2000 | A generic programmable arbiter with default master grantabstractThis paper details the design and implementation of a centralized bus arbiter implementing programmable fixed priorities arbitration. The arbiter also handles default master grant to the master with highest priority. The arbitration algorithm is computed using a tree of specialized comparators to fully exploit hardware parallelism. The design is implemented as a generic VHDL model whose parameter is the number of masters. After synthesis and place & route, a 16 masters arbiter has a critical path delay of 7.5 ns in 0.5 /spl mu/m technology. Frédéric Pétrot, Denis Hommais |
ISCAS | 1 |
| 1995 | A High Performance Modular Embedded ROM ArchitectureabstractWe describe a CMOS Read Only Memory architecture designed for high performances and low power consumption using domino logic. Short read delays are achieved using hierarchical evaluation of the read busses, at the price of some more material. Partial block evaluation allows power consumption to be greatly reduced for blocks with an important number of words, turning into an advantage this material increase. This architecture is well suited for memories embedded within synchronous systems due to it excellent speed/power performance. The architecture implementation is done as a parameterized generator, using a tiler and leaf cell approach. The leaf cells are designed using symbolic layout, providing a high degree of process independence. The tiler is written using the general purpose C language to ensure software portability. Marcello Duhalde, Alain Greiner, Frédéric Pétrot |
ISCAS | 3 |